A Model Context Protocol server for medical information retrieval, lab test analysis, and disease pattern matching using Qdrant vector store and MongoDB
Medical MCP server shows foundational implementation but falls short of production quality across multiple dimensions. Tool descriptions are present but lack depth and specificity. Parameter schemas are partially defined but inconsistent across tools. Critical issue: four tools (analyze_lab_tests, explain_lab_tests, medical_get_overview, medical_get_sections) have incomplete schema definitions with nested 'object' types lacking field specifications. Output schemas are not documented. Error handling exists but lacks actionable recovery guidance. The server operates within a specialized medical domain but does not employ domain-specific patterns like structured clinical output validation or differential diagnosis ranking transparency.
Analyzes patient lab test results against disease patterns to identify potential diagnoses with scoring and contradiction detection
Provides detailed explanations for patient lab test results including medical context and interpretation
Retrieves overview information for specified disease IDs with optional query-based semantic ranking
Retrieves detailed section content for a specified disease with optional query-based semantic ranking
Normalizes medical queries into ranked disease candidates using semantic search and optional reranking with ICD code extraction
Health check tool that verifies connectivity to the medical Qdrant store
Nested object parameters lack field specifications. analyze_lab_tests 'tests' parameter is defined as array of object type with no field schema. This forces the LLM to guess at test data structure (name/value pairs? ICD codes? lab-specific formats?).
Output schemas are not documented in tool definitions. Code returns dict structures with fields like 'disease_id', 'matched_patterns', 'total_score', 'normalized_score' but the schema is never declared. LLMs cannot plan downstream tool calls or extract correct fields.
Error responses lack actionable recovery guidance. Code returns {"success": false, "error": str(e)} with raw exception messages. LLMs receive no signal on whether to retry, escalate to user, or call a different tool.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Parameter descriptions lack constraint details. 'List of lab test dictionaries' does not specify required fields (test name format? units?). 'Optional list of test categories to filter by' does not enumerate valid categories. LLMs must hallucinate valid values.
Compound naming pattern: 'medical_normalize_query', 'medical_get_overview', 'medical_get_sections' use prefix 'medical_' redundantly given server context. Within MCP, all tools are medical. Prefixing wastes characters and adds noise. Rename to normalize_query, get_overview, get_sections for clarity.
Pagination support is missing. tools returning multiple results (analyze_lab_tests returns top_k=10 results, medical_normalize_query returns top_k=5) lack limit/offset/cursor parameters and total_count in response. Large result sets will not be handled.
Response stripping incomplete. analyze_lab_tests returns 'expected_patterns', 'redundant_data', 'missing_data', fields likely internal to the analyzer that add context for LLM reasoning but are not filtered. Strip to core fields: disease_id, canonical_name, total_score, normalized_score, matched_details, contradictions.