AI-powered MCP server for FHIR document search and conversation with healthcare AI tools for querying patient data, reranking documents, and managing embeddings
Atlas MCP exhibits significant definition quality gaps typical of early-stage healthcare tools. While 21 tools are registered with clear names and basic descriptions, many lack proper input schemas, parameter descriptions are inconsistent, and output structures are not documented. The server relies heavily on external API responses without schema validation or transformation. Critical issues: (1) Tool definitions appear to be inferred from LangChain @tool decorators rather than explicit MCP registration, we can see function signatures but not MCP tool registration metadata; (2) Input schemas visible in code show type hints but many parameters lack descriptions in the schema layer; (3) Output schemas are not documented, tools return dict/Any without schema specification; (4) Error handling is present but inconsistent; (5) Several tools accept generic 'object' parameters (filter_metadata) without documenting structure. The naming convention is good (verb_noun), but the lack of explicit schema documentation and undescribed parameters prevent a higher score.
Clear agent session history for a given session ID.
Check agent health status and configuration.
Query the HC-AI agent with a natural language question. The agent uses a multi-agent workflow with query classification, tool-based retrieval (hybrid search + reranking), validation, and response synthesis.
Batch rerank multiple queries in parallel.
Get error logs with optional filtering.
Get ingestion queue status.
Get database connection and queue statistics.
Output schemas are not documented. Tools return Dict[str, Any] without specifying what fields callers should expect. The LLM cannot plan downstream operations or extract specific values without trial-and-error.
Health check tools (agent_health, embeddings_health, db_stats, db_queue) have empty input schemas with no description of what they return. This violates the 'every parameter must have a description' and 'document output schema' rules.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Check embeddings service health and configuration.
Return recent FDA drug recalls by name.
Return current FDA drug shortages by name.
Return adverse event summaries for a drug from FAERS.
Search WHO GHO indicators by name.
Ingest a clinical note for chunking and embedding.
Lookup RxNorm RxCUI by drug name.
Rerank documents by relevance to a query.
Rerank documents with optional full FHIR bundle context.
Search ClinicalTrials.gov for studies matching a query.
Search openFDA drug labels by generic or brand name.
Search ICD-10-CM codes by term.
Search PubMed articles and return summaries.
Validate whether an ICD-10-CM code exists by searching for exact code match.
Generic object parameter 'filter_metadata' in rerank/rerank_with_context tools lacks documentation of expected structure. LLMs cannot construct valid filter_metadata without examples or a schema describing required/optional fields.
Tool definitions appear inferred from LangChain @tool decorators rather than explicit MCP tool registration. We can see function signatures but not MCP ToolDefinition objects with formal schema registration. This suggests tools may not be properly exposed via MCP's ListTools.
No error recovery guidance. Tools like search_fda_drugs catch exceptions and return FDAResponse with error messages, but do not tell the LLM what to do next (retry? try a different tool? ask the user?). Error responses must guide the agent's next action.
No pagination strategy documented. Tools like search_pubmed, search_clinical_trials accept 'max_results' but do not document how to retrieve the next page, whether a cursor or offset is provided, or what the total count is. Large result sets could blow the context window.
Agent tools (agent_query, agent_clear_session) accept session_id and patient_id but do not document session lifecycle, when to create new sessions, or how sessions map to conversation context. This is critical for multi-turn interactions.
The 'ingest' tool accepts resource_json as a string. No documentation on expected JSON structure, validation, or what happens if the JSON is malformed. LLMs will guess and pass invalid JSON.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present. The agent_clear_session tool is clearly destructive but has no marker. rerank tools are read-only but unmarked. This prevents clients from enforcing safety policies.