MCP servers for Physical AI Oncology Clinical Trials — provides authorization, FHIR clinical data, DICOM imaging, audit ledger, and provenance tracking for robotic systems in oncology trials
This oncology trials MCP server provides 18 tools across 5 functional domains (authz, DICOM, FHIR, ledger, provenance). Most tools have descriptions (15/18 complete), and input schemas are present with typed parameters. However, several critical gaps limit the score: (1) Output schemas are NOT documented, tools describe what they return conceptually but not in machine-readable schema form. (2) Parameter descriptions vary widely in quality; some are detailed (authz tools, provenance tools) while others are sparse (DICOM query modality filter has no explanation of valid values). (3) Error handling guidance is minimal, tools reference schema files but do not explain recovery paths. (4) Tool naming is generally clear (verb_noun pattern), but some tools ('dicom_query', 'fhir_search') could benefit from more specific naming to distinguish query behavior. (5) The descriptions are serviceable (averaging ~110 chars) but lack LLM-optimized depth about WHEN to call each tool vs. alternatives.
Evaluate an authorization request against the default policy. Returns a schema-valid authz-decision per schemas/authz-decision.schema.json. Fields align to the canonical schema: allowed, effect, role, server, tool, matching_rules (structured), evaluated_at, and optionally deny_reason.
Issue a bearer token for the given role. Returns token metadata per spec/security.md.
Immediately revoke a token.
Validate a previously issued token.
Query DICOM studies with role-based modality restrictions and UID validation. Supports querying at STUDY level with optional filters by modality, patient_id, and study_uid.
Retrieve DICOM study metadata by study UID with UID validation.
Output schemas not documented. Tools describe return values conceptually (e.g., 'Returns a schema-valid authz-decision') but do not expose machine-readable JSON Schema definitions for return types. This forces LLMs to infer output structure from descriptions alone, increasing errors in downstream tool chaining.
Parameter constraints not consistently expressed as enums or min/max bounds. dicom_query 'modality' parameter says 'Optional DICOM modality filter (e.g., 'CT', 'MR')' but does not declare an enum. fhir_search 'limit' says capped at 100 but does not specify min bound. This invites LLMs to pass invalid values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 56 | - | v1 |
Look up a patient by pseudonym with de-identification.
Read a FHIR resource by resource type and ID with mandatory HIPAA Safe Harbor de-identification and HMAC-SHA256 pseudonymization.
Search FHIR resources by resource type with optional parameters and de-identification. Returns FHIR Bundle with de-identified resources (max 100 results).
Query the status of a clinical trial study by study ID.
Append a new record to the hash-chained audit ledger. Returns a schema-valid audit-record per schemas/audit-record.schema.json with fields: audit_id, timestamp, server, tool, caller, parameters, result_summary, previous_hash, hash.
Export all audit records in the specified format.
Query audit records with optional filters by server, tool, caller and result limit.
Verify the integrity of the audit chain. Checks hash continuity and previous_hash alignment across all records.
Query provenance records in the backward direction (backward lineage) from a given record ID up to a specified depth.
Query provenance records in the forward direction (forward lineage) from a given record ID up to a specified depth.
Record a provenance event in the DAG-based data lineage tracking system. Captures source ID, action, actor, tool call, source type, origin server, and optional parent record IDs.
Verify the integrity of provenance records and DAG structure. Optionally verify a specific record by ID.
Error handling descriptions missing. Tools do not explain recovery paths (e.g., 'If token validation fails, try authz_issue_token first' or 'If DICOM query returns no results, check modality constraints'). Errors will be raw API responses, leaving LLMs without guidance.
Sparse descriptions on several tools. fhir_study_status ('Query the status of a clinical trial study by study ID'), ledger_verify ('Verify the integrity of the audit chain'), and dicom_query lack guidance on WHEN to call them vs. alternatives or what prerequisites are required. Descriptions average 50 - 70 chars for these tools (below the 194-char baseline for A-grade tools).
No tool annotation hints (readOnlyHint, destructiveHint, idempotentHint). While tools are correctly marked as READ_ONLY or WRITE in metadata, input schemas lack explicit io.modelcontextprotocol.toolAnnotations to signal idempotency and safety. This forces agents to guess whether retrying is safe.
Generic parameter names without type suffix. Several tools accept ambiguous parameters: 'role', 'server', 'tool', 'caller' are generic identifiers that could be IDs, names, or objects. No suffix (role_name, caller_id) to disambiguate. LLMs may pass wrong types.