MedRAG Documentation
Welcome to the MedRAG documentation. MedRAG is a medical knowledge retrieval service — a REST API and MCP server for RAG over healthcare datasets.
What is MedRAG?
MedRAG provides:
- Vector search over medical documents with domain-specific chunking
- Multi-tenant architecture with per-tenant isolation
- MCP integration for connecting to LLM agents
- PII anonymization built-in for sensitive healthcare data
- Encryption at rest for all stored content
Quick Links
- Quickstart — Get up and running in 5 minutes
- API Reference — All endpoints documented
- Concepts — Knowledge bases, documents, chunks, embeddings
- MCP Integration — Connect MedRAG to AI agents
Quickstart
Get from zero to your first query in under 5 minutes.
1. Sign Up
Create a free account at the MedRAG portal.
Your free tier includes:
- 1 knowledge base
- 1,000 queries/month
- 10 MB storage (~5k chunks)
2. Create a Knowledge Base
curl -X POST https://your-instance/v1/knowledge-bases \
-H "Authorization: Bearer $MEDRAG_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "guidelines", "dimensions": 1024}'
3. Ingest a Document
curl -X POST https://your-instance/v1/documents \
-H "Authorization: Bearer $MEDRAG_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"knowledge_base_id": "<kb-id>",
"title": "Hypertension Guidelines",
"content": "Blood pressure targets for adults..."
}'
4. Query
curl -X POST https://your-instance/v1/query \
-H "Authorization: Bearer $MEDRAG_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query": "blood pressure targets", "top_k": 5}'
Next Steps
- Authentication — API key management and OAuth
- Concepts — How MedRAG organizes your data
- MCP Integration — Connect your AI agent (Claude, Cursor) in one line
- API Reference — Full endpoint documentation
Authentication
MedRAG uses two authentication mechanisms, each serving a different access pattern:
| Interface | Auth Method | Secrets in Config? |
|---|---|---|
| REST API | API Keys | Yes (in env vars / headers) |
| MCP Endpoint | OAuth access tokens (Client Credentials today; browser login coming) | Client secret, held by your agent or CI system |
REST API: API Keys
Every request to the REST API must include a valid API key in the Authorization header.
API Key Format
medrag_sk_<random>
Usage
curl -H "Authorization: Bearer medrag_sk_..." \
https://api.medrag.eu/v1/query \
-d '{"query": "blood pressure targets"}'
Managing API Keys
From the web portal:
- Create — Generate a new key with a specific role (read, write, admin)
- Revoke — Immediately invalidate a key
- List — View all keys with creation dates and last-used timestamps
Roles
| Role | Capabilities |
|---|---|
read | Query knowledge bases, list resources |
write | All of read + ingest documents, manage KBs |
admin | All of write + manage keys, billing, settings |
MCP Endpoint: OAuth Access Tokens
The MCP endpoint (https://medrag.eu/mcp/sse) does not accept API keys. It accepts short-lived OAuth access tokens issued by MedRAG’s authorization server.
For Backend Agents and CI/CD (available now)
Use OAuth Client Credentials for non-interactive access. OAuth applications are currently registered by MedRAG support on request; you receive a client_id and a client_secret:
curl -X POST https://medrag.eu/oauth/token \
-d "grant_type=client_credentials" \
-d "client_id=medrag_client_xxxx" \
-d "client_secret=$SECRET" \
-d "scope=mcp:read"
See MCP Integration → Enterprise for full details.
For Desktop AI Clients (Claude, Cursor, VS Code) — coming soon
Interactive sign-in from desktop AI clients (browser login, tokens managed by the client) is not available yet. Adding the MedRAG URL to a desktop client’s MCP configuration will not work until it ships. Use the REST API with an API key in the meantime.
OAuth Scopes
| Scope | Capabilities |
|---|---|
mcp:read | Query, list knowledge bases |
mcp:write | Query, list, ingest documents, manage KBs |
mcp:admin | Full access including key management |
Web Portal Account
You sign in to the web portal with your account’s email address and password. The portal is where you manage API keys, billing and settings.
Forgotten Password
- On the login page, choose Forgot password? and enter your account’s email address.
- If the address belongs to an account, a reset link is emailed to it. The page shows the same message either way, so it can’t be used to find out which addresses are registered. The link expires after 1 hour and works once.
- Open the link and choose a new password (at least 8 characters).
Resetting your password signs you out of the portal everywhere, so any portal session someone else may hold is signed out too. API keys and OAuth clients are not affected. Revoke them separately if you think they are compromised.
You can request at most 3 reset emails per hour. If the email doesn’t arrive, check your spam folder before requesting another.
Security Best Practices
API Keys (REST)
- Never commit API keys to version control
- Store keys in environment variables or a secrets manager
- Rotate keys periodically
- Revoke any key that may have been exposed
- Use the minimum required role for each key
OAuth (MCP)
- Access tokens expire in 15 minutes; request a new one before expiry
- Store the
client_secretin your CI system’s secrets manager and never commit it to version control - Contact support to revoke the OAuth application if the secret may have been exposed
- Scope your OAuth applications to the minimum permissions needed
- Enterprise: use sub-tenant scoping to enforce data isolation
Leaked Key Protection
MedRAG is enrolled in the GitHub Secret Scanning Partner Program. If you accidentally push a medrag_sk_* key to a public GitHub repository:
- GitHub detects the key instantly
- MedRAG automatically revokes it
- You receive an email notification with the repository where it was found
This protects you even if you don’t notice the leak yourself. This protection covers API keys only; keep OAuth client secrets out of version control as well.
API Reference
Base URL: https://your-instance (all REST endpoints live under /v1).
Unless noted, endpoints require authentication via Authorization: Bearer <api-key>. See
Authentication for roles. A machine-readable OpenAPI spec (all public endpoints except the OAuth, MCP and internal admin ones) is served at
GET /v1/openapi.json.
All request and response bodies are JSON unless stated otherwise. Errors use the format described in the Error Reference.
Roles. Each endpoint lists the minimum role: read < write < admin. Endpoints marked “any key”
accept any valid API key.
Health
No authentication.
| Endpoint | Description |
|---|---|
GET /health/live | 200 if the process is alive |
GET /health/ready | 200 if the database is reachable, otherwise 503 |
GET /health/index | 200 if the vector indexes exist, otherwise 503. Body: {"idx_chunks_embedding": "ok|missing", "idx_platform_chunks_embedding": "ok|missing"} |
Knowledge Bases
Create Knowledge Base
POST /v1/knowledge-bases
Role: write.
{
"name": "guidelines",
"dimensions": 1024
}
| Field | Type | Notes |
|---|---|---|
name | string | Required, 1–200 characters |
dimensions | integer | Optional, 1–4096, default 1024. Immutable after creation |
Response 201: {"id": "<kb-uuid>"}
Errors: 403 knowledge_base_limit_reached (plan limit), 422 (validation).
List Knowledge Bases
GET /v1/knowledge-bases
Any key. Returns an array of knowledge bases. Sub-tenant keys only see their own.
Update Knowledge Base
PATCH /v1/knowledge-bases/{kb_id}
Role: write. Body may contain name and/or translation_config (see
Translation Settings). Including dimensions is rejected with 422.
Response: 204.
List Documents in a Knowledge Base
GET /v1/knowledge-bases/{kb_id}/documents
Any key. Returns an array of documents. 404 if the knowledge base is not accessible to your key.
There is no endpoint to delete a knowledge base through the API; delete its documents individually or use the web portal / account deletion.
Documents
Ingest Document (JSON)
POST /v1/documents
Role: write. Rate limited per tenant.
{
"knowledge_base_id": "uuid",
"title": "Document Title",
"content": "Full text content...",
"anonymize": true,
"tags": {"specialty": "cardiology"}
}
| Field | Type | Notes |
|---|---|---|
knowledge_base_id | string (uuid) | Required |
title | string | Required, 1–500 characters |
content | string | Required, max 10 MB and the plan’s max document size |
anonymize | boolean | Optional. Defaults to the tenant’s anonymization setting |
tags | object | Optional metadata, filterable at query time via filters |
Response 202: {"job_id": "<uuid>"}. Ingestion is asynchronous; poll Get Job.
The size of content counts against your storage quota (and the sub-tenant’s, if the knowledge base belongs to one); deleting the document frees it.
Errors: 400 invalid knowledge_base_id, 404 not_found, 413 (document_too_large, storage quota exceeded, or sub_tenant_storage_quota_exceeded), 402 document_limit_reached,
402 chunk_quota_exceeded, 422. Enterprise sub-tenants may also see 403 sub_tenant_document_limit_reached,
429 sub_tenant_chunk_quota_exceeded and 429 tenant_pool_quota_exceeded.
Upload Document (File)
POST /v1/documents/upload
Content-Type: multipart/form-data
Role: write. Rate limited per tenant. Maximum request size 55 MB.
| Part | Notes |
|---|---|
file | Required. Content type text/plain or application/pdf |
metadata | Required JSON string. Must include knowledge_base_id |
curl -X POST https://your-instance/v1/documents/upload \
-H "Authorization: Bearer $MEDRAG_API_KEY" \
-F 'file=@guideline.pdf;type=application/pdf' \
-F 'metadata={"knowledge_base_id": "<kb-id>"}'
Response 202: {"job_id": "<uuid>"}.
Errors: 400 (missing file, invalid metadata), 404 not_found, 413 (document_too_large, storage quota exceeded, or sub_tenant_storage_quota_exceeded with current and limit when the knowledge base’s sub-tenant is over its storage quota), 402 (document/chunk limits), 422 (unsupported file type, pdf_extraction_failed,
parse_timeout, content_too_large).
Get Document
GET /v1/documents/{id}
Any key. 404 not_found if it does not exist.
Download Original File
GET /v1/documents/{id}/download
Any key. Returns the stored original file (application/pdf or application/octet-stream). 404 if the
document has no stored file.
Delete Document
DELETE /v1/documents/{id}
Role: write. Response: 204, or 404 not_found.
Get Job
GET /v1/jobs/{id}
Any key. Returns the status of an asynchronous job. job_type is ingestion, transcription or report_generation;
status progresses from queued through processing to completed or failed (with a sanitized error):
{
"id": "uuid",
"tenant_id": "uuid",
"document_id": "uuid",
"knowledge_base_id": "uuid",
"job_type": "ingestion",
"status": "completed",
"error": null,
"total_chunks": 42,
"processed_chunks": 42,
"created_at": "2026-10-04T10:00:00Z",
"completed_at": "2026-10-04T10:00:03Z"
}
Query
Semantic Search
POST /v1/query
Role: read. Rate limited per tenant (see Rate Limits). Optionally send an
X-User-Id header (max 256 characters) to attribute the query to an end user for per-user limits and
preferences.
{
"query": "blood pressure targets",
"top_k": 5,
"knowledge_base_ids": ["uuid"],
"filters": {"tags.specialty": "cardiology", "created_after": "2025-01-01"},
"metadata_filter": {},
"source_language": "auto",
"translate_response": false
}
| Field | Type | Notes |
|---|---|---|
query | string | Required, 1–4096 characters |
top_k | integer | Optional, 1–100, default 10 |
knowledge_base_ids | string[] | Optional. Restrict to these knowledge bases |
filters | object | Optional. Keys: tags.<key>, created_after, created_before, document_title (string values) |
metadata_filter | object | Optional. Metadata filter for platform corpora results |
source_language | string | Optional. auto (default) or one of en, nl, fr, de, es, it, bg, ro |
translate_response | boolean | Optional. Translate result content back to the query language. Not available on the Free plan |
Response 200:
{
"results": [
{
"content": "Blood pressure targets for adults...",
"score": 0.87,
"source": "tenant",
"document_id": "uuid",
"document_title": "Hypertension Guidelines",
"chunk_id": "uuid",
"knowledge_base_id": "uuid",
"chunk_index": 3,
"document_tags": {"specialty": "cardiology"}
}
]
}
Results from your own documents have "source": "tenant" and the document fields above. Results from
platform corpora carry source_record_id and metadata instead. When
translation is in effect, a metadata object is added at the top level with translation_config_source
and translation_latency_ms, and translated results include the original text in content_original.
Response headers: X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset, Retry-After.
Errors: 400 unsupported_language, 402 quota_exhausted (Free plan), 402 payment_required (suspended
account), 403 (response translation on Free plan), 422, 429 rate_limit_exceeded /
per_user_rate_limit_exceeded, 503 translation_unavailable.
Platform Corpora
List Corpora
GET /v1/corpora
Any key.
{
"corpora": [
{
"slug": "pubmed",
"name": "PubMed Abstracts",
"license_info": "…",
"requires_dua": true,
"accepted": false,
"accessible": true
}
]
}
accessible reflects your plan; accepted whether your tenant has accepted the corpus’s Data Use Agreement.
Accept Data Use Agreement
POST /v1/corpora/{slug}/accept-dua
Role: admin. Response 200: {"status": "ok"}. 404 if the corpus does not exist, 403 if it is not
available on your plan.
Translation
Role: read for translate endpoints. Supported languages: en, nl, fr, de, es, it, bg, ro.
Translate
POST /v1/translate
{"text": "Hypertension", "source": "en", "target": "nl"}
Response: {"translated_text": "…", "source": "en", "target": "nl", "latency_ms": 12.3}
Translate Batch
POST /v1/translate/batch
{"texts": ["Hypertension", "Diabetes"], "source": "en", "target": "nl"}
1–100 texts per request. Response: {"translations": ["…", "…"], "source": "en", "target": "nl", "latency_ms": 20.1}
Detect Language
POST /v1/detect
Any key. Body: {"text": "…"}. Response: {"language": "nl", "confidence": 0.973}
Translation Health and Languages
| Endpoint | Description |
|---|---|
GET /v1/translation/health | {"status": "healthy|degraded|disabled", "models_loaded": n, "models_total": n, "models": [{"source", "target", "status"}]} |
GET /v1/translation/languages | Array of supported {"source", "target"} pairs |
Translation errors: 400 (text_too_long, invalid_batch_size, unsupported_language_pair),
503 translation_unavailable, 500 translation_failed.
Translation Settings
Translation configuration object (all fields optional):
{
"query_translation": "auto",
"response_translation": "on_request",
"preferred_language": "nl",
"fallback_behaviour": "passthrough"
}
| Field | Values |
|---|---|
query_translation | auto, never |
response_translation | always, on_request, never |
fallback_behaviour | passthrough, error |
| Endpoint | Role | Description |
|---|---|---|
GET /v1/tenant/settings | admin | Returns {"translation_config": {…}} |
PATCH /v1/tenant/settings | admin | Body {"translation_config": {…}}. The Free plan cannot set response_translation to always (403 free_tier_restriction) |
PATCH /v1/sub-tenants/{id}/users/{user_id}/preferences | admin | Per-user preferences, same body. Enterprise only (403 enterprise_only) |
API Keys
Create API Key
POST /v1/auth/keys
Role: admin.
{"name": "ci-key", "role": "read"}
role must be read, write or admin. Response 201: {"key": "medrag_sk_…", "id": "<uuid>"}. The
key is shown only once.
Revoke API Key
DELETE /v1/auth/keys/{id}
Role: admin. Response: 204.
Usage and Billing
Get Current Usage
GET /v1/usage
Any key.
{
"chunks": {"used": 120, "limit": 5000, "percent": 2.4},
"documents": {"used": 3, "limit": 50, "percent": 6.0},
"storage_bytes": {"used": 10240, "limit": 52428800, "percent": 0.0},
"knowledge_bases": {"used": 1, "limit": 1},
"queries_this_period": {"used": 10, "limit": 1000, "percent": 1.0},
"query_rate_limit_per_min": 10
}
Billing Usage
GET /v1/billing/usage
Role: admin. Usage over the last 30 days plus current-period counters:
{
"plan_type": "standard",
"included_queries": 50000,
"total_requests": 1234,
"total_input_tokens": 0,
"total_output_tokens": 0,
"total_cost_microcents": 0,
"query_count": 1234,
"ingestion_chunks": 560,
"overage_queries": 0,
"overage_cost_microcents": 0
}
Billing Events
GET /v1/billing/events
Role: admin. Currently always returns an empty array.
Plan Changes
Role: admin for all. Plans: free, payg, standard, enterprise. These return 503 billing_not_available
when billing is not enabled on the instance.
| Endpoint | Body | Response |
|---|---|---|
POST /v1/billing/upgrade | {"plan": "standard"} | {"status": "upgraded", "plan": "standard", "prorated_queries": n} |
POST /v1/billing/downgrade | {"plan": "free"} | {"status": "scheduled", "plan": "free"} — applied at the end of the period |
POST /v1/billing/cancel | — | {"status": "cancellation_scheduled"} |
DELETE /v1/billing/cancel | — | {"status": "cancellation_retracted"}, or 404 no_pending_cancellation |
POST /v1/billing/interval | {"billing_interval": "monthly|annual"} | {"status": …, "billing_interval": …} |
POST /v1/billing/renew | — | Renews an annual plan; 409 not_annual_plan otherwise |
Errors: 422 invalid_plan, 409 (already_on_plan, use_downgrade_endpoint, billing_details_required,
cancellation_pending, already_on_free), 422 exceeds_target_plan_limits (downgrade blocked because current
usage exceeds the target plan).
Data Processing Agreement
| Endpoint | Role | Description |
|---|---|---|
GET /v1/dpa/status | any key | {"compliant": bool, "current_version": n, "accepted_version": n|null} |
POST /v1/dpa/accept | admin | Body {"version": n}. Optional X-Accepted-By header records who accepted. Response {"status": "accepted", "dpa_version": n} |
When your tenant has not accepted the current DPA version, API responses carry the header X-DPA-Action-Required: DPA re-acceptance required.
Account
Role: admin.
| Endpoint | Description |
|---|---|
POST /v1/account/export | Exports all account data. Returns download_url (valid for expires_in_seconds: 86400) plus account, knowledge_bases, documents, usage_history |
DELETE /v1/account | Deletes the account and all data. Requires header X-Confirm-Delete: DELETE_MY_ACCOUNT. Returns 202 |
Enterprise: Sub-Tenants
Requires the Enterprise plan and an admin key that is not itself scoped to a sub-tenant
(403 enterprise_plan_required, 403 sub_tenant_keys_cannot_manage_sub_tenants).
Sub-Tenants
| Endpoint | Description |
|---|---|
POST /v1/sub-tenants | Create. Returns 201 (or 200 with the existing record if external_id already exists) |
GET /v1/sub-tenants | List |
GET /v1/sub-tenants/{id} | Get one |
PATCH /v1/sub-tenants/{id} | Update name and/or limits |
DELETE /v1/sub-tenants/{id} | Delete. Returns 204 |
Create body:
{
"name": "Clinic A",
"external_id": "clinic-a",
"max_total_chunks": 100000,
"max_knowledge_bases": 5,
"max_documents": 1000,
"max_document_size_bytes": 26214400,
"query_rate_limit_per_min": 100,
"included_queries_per_period": 10000,
"storage_quota_bytes": 1073741824
}
Only name is required (1–200 characters); every other field is optional. Sub-tenant limits are enforced inside the parent tenant’s plan limits.
Sub-Tenant API Keys
| Endpoint | Description |
|---|---|
POST /v1/sub-tenants/{id}/api-keys | Body {"role": "read|write"}. Returns 201 {"key", "role", "sub_tenant_id"} |
GET /v1/sub-tenants/{id}/api-keys | List keys |
DELETE /v1/sub-tenants/{id}/api-keys/{key_id} | Revoke. Returns 204 |
Keys created here only see knowledge bases and documents of that sub-tenant.
Sub-Tenant Usage
| Endpoint | Description |
|---|---|
GET /v1/sub-tenants/{id}/usage | chunks, documents, queries_this_period, storage_bytes (each {used, limit}; storage counts uploaded files and is freed when a document is deleted) and active_users |
GET /v1/usage/breakdown | {"sub_tenants": [{sub_tenant_id, name, chunks, queries_this_period, active_users}], "total_chunks": n} |
GET /v1/usage/users | Array of {sub_tenant_id, name, active_users} for the current month |
Add-on: Transcription and Reports
Requires the corresponding add-on on your plan (otherwise 403 {"error": "addon_not_enabled", "addon": "transcription|report_generation"})
and the feature to be enabled on the instance (503 scribe_api_not_available). Role: write. Both endpoints are asynchronous: they return a job_id and
deliver the result to your callback_url.
Transcribe Audio
POST /v1/transcribe
Content-Type: multipart/form-data
Maximum request size 25 MB. Parts: audio (required), language, callback_url, callback_secret.
Response 202: {"job_id": "<uuid>"}.
Errors: 400 (unsupported_language, audio_too_long with max_seconds, invalid_callback_url),
503 transcription_unavailable.
Generate Report
POST /v1/reports/generate
{
"transcript": "…",
"callback_url": "https://example.com/hook",
"callback_secret": "…"
}
Response 202: {"job_id": "<uuid>"}.
Errors: 400 (empty transcript, transcript_too_long with max_chars, invalid_callback_url),
503 report_generation_unavailable.
OAuth and MCP
OAuth 2.1 endpoints (/oauth/token, /oauth/revoke, /oauth/register, /.well-known/oauth-*) and the MCP
transport (/mcp/sse, /mcp/message) are documented in MCP Integration and
Authentication.
Concepts
Understanding how MedRAG organizes and retrieves medical knowledge.
Tenants
MedRAG is multi-tenant. Each tenant has:
- Isolated knowledge bases and documents
- Separate API keys
- Individual billing plans and usage limits
- Optional sub-tenants for delegated access within an organization
Knowledge Bases
A knowledge base is a collection of documents grouped by purpose. Examples:
- Clinical guidelines (NICE, AHA, ESC)
- Drug information databases
- Patient education materials
Each knowledge base has a configured embedding dimension (e.g., 1024 for the self-hosted BGE-M3 model).
Documents
A document is a single piece of content ingested into a knowledge base. During ingestion:
- PII anonymization is applied (if enabled for the tenant)
- The content is split into chunks
- Each chunk is embedded and stored
Chunks
Chunks are the atomic units of retrieval. MedRAG splits documents by word count with overlap:
- Default chunk size: 200 words
- Overlap: 1/5 of chunk size (40 words) to preserve context across boundaries
- Each chunk retains its position index and parent document reference
Platform Corpora
In addition to tenant-uploaded documents, MedRAG provides pre-loaded medical datasets (e.g., PubMed abstracts). Access to platform corpora requires accepting a Data Use Agreement (DUA). Platform corpora use corpus-specific chunking strategies optimized for their format.
Embeddings
Each chunk is converted into a vector embedding for semantic search. Embeddings are generated by self-hosted models running on MedRAG’s own EU infrastructure; your content is never sent to a third-party AI provider.
Queries
A query is a natural-language question or search term. MedRAG:
- Embeds the query using the same model as the knowledge base
- Performs approximate nearest neighbor (ANN) search using HNSW indexes
- Returns the top-k most relevant chunks with their source documents
Queries support optional filters:
- Knowledge base IDs — restrict search to specific knowledge bases
- Metadata filters — filter results by document metadata (tags, titles)
- Sub-tenant scope — restrict to knowledge bases accessible by a sub-tenant
PII Anonymization
MedRAG can automatically detect and redact personally identifiable information during document ingestion. This is configurable per tenant and runs before chunking to ensure no PII leaks into stored chunks or embeddings.
Rate Limits & Pricing
All prices exclude VAT.
Plans
| Feature | Free | PAYG | Standard | Enterprise |
|---|---|---|---|---|
| Monthly fee | €0 | €0 | €49 | €199 |
| Included queries/month | 1,000 | 0 | 50,000 | 500,000 |
| Knowledge bases | 1 | 5 | 20 | Unlimited |
| Documents | 50 | 1,000 | 1,000 | 10,000 |
| Max document size | 5 MB | 25 MB | 25 MB | 100 MB |
| Storage | 50 MB | 50 MB | 1 GB | Custom |
| Max chunks | 5,000 | 100,000 | 100,000 | 1,000,000 |
| Rate limit (queries/min) | 10 | 100 | 100 | 500 |
| Overage (per query) | — | €0.004 | €0.002 | €0.001 |
| Ingestion cost | — | €0.001/chunk | Free | Free |
Rate Limiting
API requests are rate-limited per tenant based on your plan:
- Query rate limit: Applied to
POST /v1/query(per-minute, by plan: 10 Free, 100 PAYG and Standard, 500 Enterprise) - Burst limits: 20 requests/second per API key on
/v1/query, 10 requests/second per API key on/v1/documents - Global rate limit: 100 requests/second per IP across all endpoints
When rate-limited, you’ll receive a 429 Too Many Requests response with a Retry-After header.
Response Headers
Successful query responses include rate limit headers:
X-RateLimit-Limit: 10
X-RateLimit-Remaining: 7
X-RateLimit-Reset: 1719936000
Overage
- Free tier: Hard cap — once the 1,000 monthly queries are used, queries are rejected with
402 quota_exhausted - PAYG tier: All queries charged at the €0.004 rate (no included queries); indexed chunks are charged at €0.001 each
- Standard tier: Overage charged at €0.002 per query beyond included amount
- Enterprise: Overage charged at €0.001 per query beyond included amount (billed by bank transfer; contact admin@medrag.eu)
Changing plan
A downgrade or cancellation is refused (422 exceeds_target_plan_limits) when your current usage is above the target plan’s limits — for example, more knowledge bases than it allows, or more stored data than its storage quota. Delete the excess and try again.
Error Reference
MedRAG returns JSON errors with an error field:
{
"error": "error_code",
"message": "Optional human-readable description"
}
Error Codes
| Code | HTTP Status | Description | Resolution |
|---|---|---|---|
missing_auth | 401 | No Authorization header | Send Authorization: Bearer <api-key> |
invalid_api_key | 401 | Malformed or unknown API key | Check the key and header format |
forbidden | 403 | API key role is too low for this endpoint | Use a key with a higher role |
not_found | 404 | Resource doesn’t exist or isn’t accessible to your key | Check the resource ID |
payment_required | 402 | Account suspended for non-payment | Settle the outstanding invoice |
quota_exhausted | 402 | Free plan query quota used up | Upgrade your plan or wait for reset |
document_limit_reached | 402 | Plan’s document limit reached | Delete documents or upgrade |
chunk_quota_exceeded | 402 | Plan’s chunk limit reached | Delete documents or upgrade |
knowledge_base_limit_reached | 403 | Plan’s knowledge base limit reached | Upgrade your plan |
enterprise_plan_required | 403 | Endpoint needs the Enterprise plan | Upgrade your plan |
addon_not_enabled | 403 | Transcription/report add-on not enabled | Enable the add-on |
document_too_large | 413 | Document exceeds the plan’s size limit (max_bytes returned) | Reduce document size or upgrade |
sub_tenant_storage_quota_exceeded | 413 | Sub-tenant storage quota reached (current, limit returned) | Raise the sub-tenant quota or delete documents |
validation_error / message | 422 | Invalid request body | Check required fields and formats |
rate_limit_exceeded | 429 | Too many requests | Wait for Retry-After seconds |
per_user_rate_limit_exceeded | 429 | Per-user limit hit (X-User-Id) | Wait for Retry-After seconds |
sub_tenant_chunk_quota_exceeded | 429 | Sub-tenant chunk quota reached | Raise the sub-tenant limit |
tenant_pool_quota_exceeded | 429 | Tenant-wide chunk pool exhausted | Upgrade or delete documents |
internal_error | 500 | Server error | Retry; contact support if persistent |
translation_unavailable / *_unavailable | 503 | Feature not available on this instance right now | Retry later |
Validation errors (422) and some others return a plain-language message in error
(for example "name must be 1-200 characters") rather than a code. Not every error includes a
message field.
Rate Limit Errors
When rate-limited, the response includes:
{
"error": "rate_limit_exceeded",
"retry_after_seconds": 5,
"docs_url": "https://docs.medrag.eu/usage-limits"
}
Quota Errors
When a plan quota is exhausted:
{"error": "quota_exhausted"}
Document and chunk limits additionally return current, limit and docs_url.
MCP Integration
MedRAG provides a remote MCP (Model Context Protocol) endpoint that lets LLM agents query your medical knowledge bases directly. Your agent connects over HTTP — no local installation.
Availability
| Client | Status |
|---|---|
| Backend agents, CI/CD, server-to-server (OAuth Client Credentials) | Available — see below |
| Desktop AI clients (Claude Desktop, Cursor, VS Code) with browser sign-in | Coming soon — not available yet |
The MCP endpoint is https://medrag.eu/mcp/sse (with https://medrag.eu/mcp/message for tool calls). It does not accept API keys; it accepts OAuth access tokens only.
How Authentication Works
MedRAG’s authorization server issues short-lived access tokens (15 minutes):
- Your OAuth application is registered by MedRAG support, which gives you a
client_idand aclient_secret - Your agent requests a token from
https://medrag.eu/oauth/tokenusing theclient_credentialsgrant - Your agent sends the token as
Authorization: Bearer <token>to the MCP endpoint - Without a valid token the endpoint answers
401with aWWW-Authenticateheader
Interactive sign-in for desktop AI clients (browser login with your MedRAG account, tokens managed by the client) is in development. Until it ships, adding the MedRAG URL to a desktop client’s MCP configuration will not work.
The client_secret is a credential: keep it in your CI system’s secrets manager, never in version control.
Enterprise: Backend Agents & CI/CD
Enterprise customers running automated agents (CI/CD pipelines, server-to-server integrations) that cannot open a browser can use OAuth Client Credentials:
1. Register an OAuth Application
Coming soon. Self-service OAuth application management (create/view/revoke) will be available in the MedRAG web portal under Settings → OAuth Applications.
In the meantime, contact support to register an OAuth application for your tenant. You’ll receive a
client_idandclient_secret.
Store the secret securely (e.g., in your CI system’s secrets manager).
2. Obtain an Access Token
Your agent requests a short-lived token:
curl -X POST https://medrag.eu/oauth/token \
-d "grant_type=client_credentials" \
-d "client_id=medrag_client_xxxx" \
-d "client_secret=$MEDRAG_CLIENT_SECRET" \
-d "scope=mcp:read"
Response:
{
"access_token": "eyJ...",
"token_type": "Bearer",
"expires_in": 900,
"scope": "mcp:read"
}
The token expires in 15 minutes. Your agent should request a new one before expiry.
3. Connect to the MCP Endpoint
# SSE connection with Bearer token
curl -H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Accept: text/event-stream" \
https://medrag.eu/mcp/sse
Or send individual tool calls:
curl -X POST https://medrag.eu/mcp/message \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/call",
"params": {
"name": "medrag_query",
"arguments": {"query": "blood pressure targets", "top_k": 5}
},
"session_id": "<session_id_from_sse>"
}'
Sub-Tenant Scoping (Enterprise)
If your OAuth application is scoped to a specific sub-tenant, the access token will only grant access to that sub-tenant’s knowledge bases. This ensures data isolation across your hospitals or organisations.
Available Tools
medrag_query
Semantic search across your knowledge bases.
| Parameter | Type | Required | Description |
|---|---|---|---|
query | string | yes | Search query text |
top_k | integer | no | Number of results (default: 10) |
knowledge_base_ids | string[] | no | Limit search to specific KBs |
medrag_list_knowledge_bases
List available knowledge bases. Returns ID, name, and document count for each.
No parameters required.
How It Works
- Your AI client discovers MedRAG’s tools via MCP
tools/listat connection time - When the user asks a question that could benefit from medical knowledge retrieval, the agent decides to call
medrag_query - MedRAG returns scored chunks with source document references
- The agent uses the retrieved information to ground its answer
The agent decides when to call MedRAG based on the tool description and user’s question — just like any other MCP tool.
Security
- No secrets in config: The only thing in your MCP config is a URL
- Short-lived tokens: Access tokens expire in 15 minutes; refresh is automatic
- Scope control: Request only the permissions your agent needs
- Tenant isolation: Each token is scoped to your tenant; cross-tenant access is impossible
- Sub-tenant isolation (Enterprise): Tokens can be further restricted to specific sub-tenants