Biomedical CLI and MCP server — genes, variants, trials, articles, drugs, diseases, pathways, proteins, adverse events, PGx
BioMCP has 4 tools with basic descriptions and input schemas, but falls significantly short of production quality. Tool naming lacks clear action verbs (discover, search, variant, biomcp are vague or generic). Descriptions are present but minimal, averaging ~80 chars, below the 194-char baseline for A+ tools. Input schemas exist but lack rich documentation: parameters have types but minimal description text. Parameter descriptions are terse (e.g., 'Gene symbol, variant identifier...' for discover.identifier). No output schemas are documented, critical for an agent to know what fields to expect from biomedical lookups. Error handling guidance is absent. The 'biomcp' tool appears to be a catchall CLI descriptor, not a proper tool. Composition is weak: discover and search overlap conceptually, and the relationship between variant operations and discovery is unclear. No evidence of idempotency declarations, permission gates, or secret injection patterns.
BioMCP CLI and MCP server for biomedical data discovery and analysis across genes, variants, clinical trials, PubMed articles, drugs, diseases, pathways, proteins, adverse events, and pharmacogenomics. Integrates with external biomedical databases and knowledge sources.
Discover genes, variants, or other biomedical entities by identifier
Search for biomedical entities (diseases, articles, genes, drugs, variants, pathways, proteins, adverse events)
Perform operations on genetic variants including normalization and annotation
Tool 'biomcp' is not a proper tool, it reads as a CLI description (Cargo.toml metadata) rather than a distinct operation. The schema is missing entirely and the name does not start with an action verb. This appears to be metadata pollution in the tool registry.
Tool naming lacks clear action verbs. 'discover', 'variant' are not verb_noun forms. Should be 'discover_entity', 'search_genes', 'annotate_variant', etc. LLMs rely on verb_noun naming to infer intent from the name alone.
Output schemas are not documented. Agents cannot plan downstream calls or extract required fields without knowing what discover(), search(), or variant() return. This violates the 100% A+ baseline that all tools must have documented return types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Parameter descriptions are minimal and generic. E.g., discover.identifier reads 'Gene symbol, variant identifier, or biomedical entity identifier' but does not explain validation constraints, formats, or what constitutes a valid identifier. Compare to baseline: good param descriptions are 50-150 chars with explicit constraints.
The 'search' tool's entity_type parameter should be an enum (disease, article, gene, drug, variant, pathway, protein, adverse_event) but is documented as a free-form string. Free-form invites hallucinated values. Enums are self-documenting.
No error handling guidance. If discover() fails to find an identifier, does it return a 404, an empty result, or a structured error? Agents need recovery hints (e.g., 'Entity not found. Try search() with a partial query').
Potential tool composition confusion: both 'discover' and 'search' retrieve biomedical entities. The distinction is unclear. Is discover for exact matches and search for fuzzy? Should they be combined or clearly differentiated? Currently ambiguous naming invites agent misuse.
The 'variant' tool accepts an 'operation' parameter (normalize, annotate, etc.) as a free-form string. Should be an enum. Also, the description does not explain what normalize vs annotate returns or when to use each.
No documentation of pagination for search results. The 'limit' parameter exists but there is no mention of offset, page_size bounds, or how to handle results exceeding the limit. Large biomedical result sets will blow the context window without pagination guidance.
Tool descriptions are below the 194-char baseline (avg ~80 chars). They lack WHEN to use the tool, prerequisite context, or dependency hints. E.g., 'discover' does not mention that it requires an exact identifier or suggest calling search() first for fuzzy lookup.