A medical symptom analysis MCP server that extracts symptoms and details from user input, matches them to possible diseases, and provides medical advice using LLM and web search integration.
SymptomAI MCP suffers from critical gaps across all definition quality dimensions. While the 4 tools are explicitly registered via @mcp.tool() decorators in server.py, their schemas, descriptions, and parameter documentation fall significantly short of production standards. Tool names lack clear action verbs (extract_symptoms, extract_details, match_diseases, get_advice are acceptable but generic). Descriptions exist but are minimal (10-40 chars) and fail to explain WHEN to use each tool or clarify dependencies between them. Input schemas are present but severely underdeveloped: parameters are typed as simple strings without enums, min/max bounds, or format guidance. Most critically, no output schemas are documented, LLMs cannot plan downstream tool chains when return types are unspecified. Error handling is absent: tools return raw text or exceptions with no recovery guidance. The sequential dependency chain (symptoms → details → diseases → advice) is implicit in client.py but entirely undocumented in tool definitions, forcing agents to infer the workflow rather than having it declared. This is a proof-of-concept implementation that requires substantial hardening before production use.
Extract relevant details from user free text input.
Extract symptoms from user free text input.
Provide medical advice for diseases.
Match symptoms and details to possible diseases, return explanations.
No output schemas documented. All 4 tools return unstructured text (str) with no indication of structure, fields, or format. LLMs cannot parse responses or plan multi-tool chains without knowing return types.
Tool descriptions are 10-40 characters, well below the 34-char baseline minimum for clarity. None explain WHEN to call each tool, prerequisites, or how tools chain together. Example: 'Extract symptoms from user free text input.' lacks context on what 'extraction' means for medical diagnosis.
All 4 tools accept only string parameters with minimal descriptions. No enums, ranges, or format constraints. match_diseases accepts 'symptoms' and 'details' as comma-separated strings but never specifies the format, field count, or examples of valid input.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 30 | - | v1 |
No error handling or recovery guidance. Tools invoke external LLM and search APIs (Google Generative AI, Tavily) but do not handle failures, timeouts, or rate limits. When match_diseases returns 'No diseases identified' or get_advice returns 'No advice available', LLMs have no guidance on what to do next.
Tool workflow dependencies are implicit, not declared. The sequential flow (user_input → extract_symptoms → extract_details → match_diseases → get_advice) is hard-coded in client.py but tools do not declare that match_diseases depends on extract_symptoms/extract_details output. LLMs must infer the workflow from calling patterns.
API keys (GOOGLE_API_KEY, TAVILY_API_KEY) are hardcoded in server.py as placeholder strings. While the rubric emphasizes never exposing secrets in tool parameters, server-side injection via environment variables is also not properly implemented, keys are hardcarded as '###' strings.
match_diseases and get_advice leverage external search and LLM APIs but provide no visibility into search quality, result truncation, or API failures. safe_extract_text() silently degrades malformed results, potentially returning empty strings without signaling to the LLM that retrieval failed.
Parameter descriptions are generic and do not explain format or constraints. 'user_input' is described as 'Free text input from the user describing their symptoms' but does not specify length limits, language, character encoding, or what makes input valid vs invalid.