MCP server for bioinformatic analysis of viral metadata from the Sequence Read Archive, supporting hypothesis validation and anomaly detection
This server has three tools with basic schemas and descriptions, but significant quality gaps across naming, parameter documentation, error handling, and output specifications. All three tools are read-only operations, which is positive for safety, but the definitions lack the rigor expected for production use. Tool names follow verb_noun convention (good), but descriptions are minimal (averaging ~60 chars, well below the 194-char baseline for A+ tools). Parameter descriptions are present but sparse. No output schemas are documented. Error handling is absent, there's no guidance for the LLM on what to do if a query fails or returns no results. The codebase shows infrastructure (Neo4j, NCBI, LangGraph workflows) but tool definitions themselves are underdeveloped.
Fetch palm_ids from a given virus name.
Fetch similar viruses based on palm_ids and percent_identity.
Run metadata analysis based on input virus and hypothesis.
Tool descriptions are too brief (30 - 60 chars, well below the 194-char production baseline) and lack context on WHEN to use the tool and WHAT happens on failure.
No output schemas are documented. LLMs cannot determine what fields to expect from these tools, preventing downstream chaining and forcing LLMs to guess at result structure.
Parameter descriptions are minimal or absent. 'palm_ids' has a description, but the meaning of percent_identity thresholds and expected value ranges are not explained. Parameters lack constraints (min/max, valid ranges).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 47 | - | v1 |
No error handling or recovery guidance. If a species_name is not found, or if percent_identity filtering returns zero results, there is no documentation of what the tool returns or what the LLM should do next.
get_virus_metadata_analysis has vague naming. 'Run metadata analysis' does not clearly convey what the tool does. A more specific name like 'validate_virus_hypothesis' or 'analyze_virus_cofactors' would be clearer.
Default values for 'virus_species' and 'hypothesis' parameters are hardcoded (Papaya meleira virus, cancer cofactor). This invites LLMs to reuse these defaults literally in real calls instead of adapting to user context.
Schema definitions for input parameters are present but incomplete. Parameters lack minimum/maximum constraints, enum definitions, and format specifications. For example, 'percent_identity' should have min=0, max=100.