MCP server for clinical diagnosis providing symptom extraction from clinical text, diagnosis with treatment recommendations, and relevant PubMed literature discovery
Single tool 'clinical_diagnosis' has a reasonable description (~180 chars), but lacks critical schema documentation, parameter validation guidance, and error handling patterns. The input schema is minimal (only 'clinical_text: string'), with no constraints, ranges, or format guidance. No output schema is documented. Tool definition is directly visible in source but quality gaps are substantial. This server implements blocking clinical NLP operations with no retry or timeout guidance, and exposes medical data without output field documentation.
Extract symptoms from clinical text, provide a possible diagnosis with recommended treatment, and surface relevant PubMed literature.
Missing output schema documentation. Tool returns formatted markdown report but does not document the structure, fields, or what LLM should expect. LLMs cannot reliably parse freeform text responses.
Input parameter 'clinical_text' lacks validation constraints or format guidance. No guidance on min/max length, character restrictions, required elements, or expected medical note format. LLMs may pass invalid or oversized inputs.
No error handling or recovery guidance. Tool can fail silently (no active symptoms, API timeouts, OpenAI/PubMed failures) with no actionable error messages. LLM cannot self-correct or retry intelligently.
No timeout or rate-limit specification. Tool spawns async threads to OpenAI and PubMed without explicit timeout. A hung external API blocks the entire call indefinitely.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Tool name 'clinical_diagnosis' is a noun phrase, not a verb_noun action. Better naming: 'analyze_clinical_note', 'extract_diagnosis_from_note', or 'generate_clinical_report'. This violates the production baseline where 90% of A+ tools start with an action verb.
No documentation of returned symptom metadata structure (e.g. what fields are in each symptom object, canonical_name vs display_name distinction). LLM cannot reason about output without explicit schema.