Model Context Protocol server for the Ensembl genomics database. Provides AI-powered access to gene lookup, variant analysis (VEP), comparative genomics, regulatory features, coordinate mapping, and sequence retrieval across 300+ species.
The Ensembl MCP server demonstrates solid definition quality with 10 well-structured tools covering a comprehensive genomics domain. All tools have explicit JSON Schema input definitions with type declarations, descriptions, and reasonable parameter constraints. Tool names follow the verb_noun pattern (ensembl_*) consistently. However, there are systematic gaps: (1) output schemas are completely undocumented, no guidance on what each tool returns or how to chain results; (2) parameter descriptions, while present, are verbose and could be more concise for LLM parsing; (3) error handling is minimal, no recovery guidance in descriptions; (4) some parameter relationships are underdocumented (e.g., how raw/page_size interact); (5) descriptions occasionally drift into implementation details rather than user intent. The baseline for production tools is 194 chars avg description, these are 200-350 chars, signaling over-explanation. All 10 tools are directly visible in src/handlers/tools.ts with explicit registration, so no inference penalty applies.
Get homology, orthology, paralogy, gene trees, and cross-species alignments. Covers /homology, /genetree, /alignment, and /cafe endpoints.
Find genomic features (genes, transcripts, regulatory elements) that overlap with a genomic region or specific feature. Automatically handles assembly-specific format variations (GRCh38/hg38, chromosome naming conventions, coordinate systems). Covers /overlap/region and /overlap/id endpoints.
Look up genes, transcripts, variants by ID or symbol. Get cross-references and perform ID translation. Covers /lookup/* and /xrefs/* endpoints plus variant_recoder.
Map coordinates between genome assemblies and coordinate systems. Handles inter-assembly mapping, chromosome → scaffold, exon coordinates.
Get server metadata, data releases, species info, and system status. Covers /info/* endpoints and /archive/id for version tracking.
Ontology and taxonomy search. Get ontology terms, taxonomy information, and cross-references to standardized biological vocabularies.
Output schemas completely undocumented across all 10 tools
No error handling or recovery guidance in tool descriptions
Parameter descriptions are verbose (200-350 chars) when best practice is 50-100 chars, wastes tokens and buries key constraints
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 58 | - | v1 |
Get protein-level features, domains, and annotations for proteins and translations.
Get regulatory features, binding matrices, and regulatory annotations. Covers regulatory overlap endpoints and binding matrix data.
Retrieve DNA, RNA, or protein sequences for genes, transcripts, regions. Covers /sequence/id and /sequence/region endpoints.
Variant analysis, VEP (Variant Effect Predictor) annotation, phenotype associations, and population frequencies. Covers /variation, /vep, /ld, and /phenotype endpoints.
Implementation details (API endpoints like /overlap/region, /xrefs) leaked into tool descriptions instead of focusing on user intent
Parameter mutual exclusivity and relationships documented inconsistently or not at all
Tool names contain domain jargon ('ensembl_meta', 'ensembl_compara', 'ensembl_ontotax') that is opaque to non-genomics agents
No documented examples of parameter combinations or expected result structures
Batch parameter limits (200 identifiers) mentioned in parameter description but not enforced or validated visibly in schema