This server shows significant quality gaps across definition quality. While all 8 tools are present with basic schemas visible in the code (api_server.py and mcp_server.py), there are critical deficiencies: (1) Tool descriptions are extremely brief (under 50 chars for most tools), providing minimal context for LLM selection. (2) Parameter descriptions, while present, are often generic and lack detail about constraints, formats, or expected values. (3) Return types are not documented in the tool definitions themselves, the agent cannot infer what fields will be returned. (4) No error handling guidance is visible, tools that query a database may fail, but provide no recovery hints. (5) Several tools accept string IDs without validation constraints or format guidance (hpo_id format should be specified as 'HP:XXXXXXX'). The search_hpo_for_symptom tool appears to use VoyageAI embeddings, but this dependency is not reflected in the tool description, leaving ambiguity about why this tool behaves differently from simple lookups. Overall, descriptions are functional but far below the 50-200 char baseline for A+ tools.
Tool descriptions are critically short (28-78 chars), far below the 50-200 char baseline for A+ tools. Descriptions do not explain WHEN to use each tool or how they differ from one another.
Parameter descriptions lack constraint details: hpo_id should document format (e.g., 'HPO identifier in format HP:XXXXXXX'), gene_id should specify 'NCBI gene identifier (integer or string)', disease_id should clarify format (OMIM or other). Current descriptions are too generic.
Expand all tool descriptions to 80-150 chars, following this template: '[WHAT the tool does] [WHEN to use it instead of similar tools] [Key input format] [Key output type]'. Example: 'Retrieve genes associated with an HPO phenotype term. Use when you have an HPO ID (HP:XXXXXXX) and need to find causative genes. Returns a list of Gene objects with ncbi_gene_id and gene_symbol.'
For each parameter, add format and constraint examples: hpo_id → 'HPO identifier in format HP:0000007'; gene_id → 'NCBI gene ID (e.g., 675 or as integer)'; disease_id → 'Disease ID in format OMIM:243400'; k → 'Integer between 1 and 50 (default 5)'. Replace vague descriptions with specific constraints.
Document return types in each tool definition. Add a 'Returns' field or similar: 'Returns: { genes: [{ ncbi_gene_id: string, gene_symbol: string }], hpo_id: string, hpo_name: string }'. This allows LLMs to know what fields are available for downstream operations.
Add error handling guidance to each tool description. Example: 'If HPO term not found, consider using search_hpo_for_symptom() with the disease symptom instead, then map backwards.'
Clarify search_hpo_for_symptom's unique role: expand description to 'Semantic search for HPO phenotype terms using vector similarity. Enter a symptom description (e.g., 'joint pain', 'hearing loss') and get the top k matching HPO terms with their IDs. Preferred when you have symptom text but not an HPO ID yet. Returns ranked list of HPOToGene objects with relevance scores.' Include a note that this tool uses ML embeddings, unlike direct ID lookups.
Score history
Overall score trend
↑ 42 points across a rubric change (v1 → v2)
42/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
42
2026-07-28+
v2
2026-03-09
F
0
-
v1
auth
source verified
57/100
Search HPO terms for a specific English symptom (optimized for workflow step 2).
No output schemas documented. Tools return structured objects (Gene, Disease, HPOToGene, etc. as seen in api_server.py), but the MCP tool definitions do not document these return types. LLMs cannot infer what fields to extract or plan downstream tool calls.
No error handling guidance visible. Tools query sqlite3 and may fail (404s, connection errors). No error descriptions or recovery hints (e.g., 'If HPO not found, try search_hpo_for_symptom() first'). Agents have no guidance on how to recover from failures.
search_hpo_for_symptom uses VoyageAI embeddings (ML-based semantic search) but this is not mentioned in the tool description. LLMs may not understand that this tool solves a different problem than direct lookup tools and might conflate it with simple get_hpo_* operations.
Tool naming lacks differentiation for related operations. Multiple 'get_*' tools with similar names (get_genes_by_hpo, get_hpo_by_gene, get_diseases_by_hpo, get_hpo_by_disease) may confuse LLMs about which tool to use. No clarifying descriptions to explain the semantic difference between 'get_X_by_Y' variants.
Add bounds to numeric parameter 'k': document as 'k: integer, range 1-100, default 5. Higher values return more candidates but cost more compute.'
For related tools (get_genes_by_hpo vs get_hpo_by_gene), add disambiguation in descriptions: 'get_genes_by_hpo: START with an HPO term, get GENES. Reverse operation: get_hpo_by_gene.' This helps LLMs pick the correct direction.
Consider grouping related lookup tools with shared input/output documentation. For example, a 'Phenotype-Gene Associations' section documenting both directions: hpo→gene and gene→hpo, with clear examples of when each is appropriate.
Add validation examples in error descriptions: 'HPO ID must start with HP: and contain 7 digits. Invalid: HP:123, hpo:0000007. Valid: HP:0000007.' This guides LLMs to format IDs correctly.
Document API limits and pagination: if results can be large, add guidance like 'For HPO terms with hundreds of associated genes, consider using limit/offset parameters if available, or filtering by gene importance scores.'