An MCP server for genetic variant analysis using Ensembl VEP and ClinVar, providing tools to look up biological effects of genetic variants and their clinical significance.
VepClin MCP defines 2 tools with strong, domain-specific descriptions (194 - 450 chars each) that explain WHAT, WHEN, and expected output. Both tools have proper input schemas with typed parameters and descriptions. Naming is clear and verb-driven (get_*). However, output schemas are documented only in prose within descriptions, not as formal JSON Schema. Error handling is absent, no guidance on invalid HGVS, missing genes, or API failures. No tool annotations (readOnlyHint, destructiveHint). Parameter descriptions are thorough but lack explicit constraints (e.g., HGVS format regex, enum for genome build). Both tools are read-only and stateless, which is correct for this domain.
Look up clinical significance and associated diseases for a genetic variant in ClinVar, given its gene symbol and short-form protein change (e.g. "BRAF", "V600E"). Call get_vep_consequence first and pass its gene_symbol and protein_change_short values as arguments here — not protein_change (the p.Val600Glu-style HGVS format), which will not match. Returns a dict with: - found: True if results are found or False - matches: List filled with dict of classifications and traits: - "variation_id": Clinvar variation ID - "title": Variant title (e.g. 'NM_004333.6(BRAF):c.1799T>A (p.Val600Glu)') - "germline_classification": (e.g. 'Conflicting classifications of pathogenicity') - "germline_traits": (e.g. ['Cardiovascular phenotype', 'Vascular malformation', 'RASopathy']) - "clinical_impact_classification": How clinically impactful (e.g. 'Tier I - Strong') - "clinical_impact_traits": (e.g. ['Ganglioglioma', 'Malignant glioma', 'Dysembryoplastic neuroepithelial tumor']) - "oncogenicity_classification": Tendency to cause tumors/cancer (e.g. 'Oncogenic') - "oncogenicity_traits": (e.g. ['Neoplasm']) If found is False, function returns {"found": False, "matches": []} Each match includes "position_verified": true if the genomic position matches what get_vep_consequence reported, false if it doesn't (indicating possible ambiguity or a different variant), or null if verification wasn't possible. Treat false results with caution and mention the discrepancy to the user.
Look up the biological effect of a genetic variant given in supported HGVS notation. Use genomic HGVS (e.g. "chr7:g.140753336A>T") or coding HGVS with a transcript accession (e.g. "NM_004333.6:c.1799T>A"). Bare coding HGVS like "c.1799T>A" is ambiguous and unsupported. Use this when a user asks what a specific DNA variant does. Genome build and transcript scope (MANE Select only vs. all transcripts) are set by the user via CLI commands, not by this tool's arguments. Returns a dict with: - gene_symbol: the affected gene (e.g. "BRAF") - consequence: the molecular consequence type (e.g. "missense_variant") - protein_change: the amino acid change in HGVS protein notation (e.g. "p.Val600Glu") ("None" if not applicable, like for synonymous variants) - protein_change_short: the same change in short form, (e.g. "V600E" — use this exact value when calling get_clinvar_summary) - impact: Ensembl's severity bucket for the consequence type (HIGH/MODERATE/LOW/MODIFIER) - sift_prediction / sift_score: SIFT's tolerance prediction and probability (lower score = more likely damaging; "None" if not applicable to this consequence type) - polyphen_prediction / polyphen_score: PolyPhen-2's damage prediction and probability (higher score = more likely damaging; "None" if not applicable) - build: genome build used for this lookup ("grch38" or "grch37") - input_hgvs: the original HGVS input string - hgvs_format: "genomic" for g. inputs or "coding" for transcript-qualified c. inputs - transcripts: only present if the user has selected "all transcripts" mode and the variant has more than one transcript consequence. A list of dicts, each with transcript_id, consequence, protein_change, impact, sift_prediction, and polyphen_prediction for that transcript. The top-level fields above always reflect the primary (MANE Select or first-returned) transcript, which is also what gets passed to get_clinvar_summary — do not use a value from the transcripts list for that purpose.
Output schemas documented only in prose descriptions, not as formal JSON Schema. LLMs cannot parse expected response structure from docstrings alone.
No error handling guidance. Tool descriptions do not explain what happens on invalid HGVS, missing genes, API timeouts, or network failures. LLMs have no recovery path.
No tool annotations. Both tools are read-only but lack readOnlyHint. This prevents clients from optimizing caching or safety policies.
Parameter 'variant' in get_vep_consequence lacks format constraint. Description mentions 'supported HGVS notation' but does not specify regex or enum. LLMs may pass invalid formats.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 83 | 2026-07-28+ | v2 |
Genome build ('grch38' or 'grch37') is set via CLI, not tool parameters. This couples tool behavior to server state, violating stateless request handling. Clients cannot control build per-call.