An MCP server for genetic research and DNA analysis, created for the NVIDIA Hackathon. Provides tools for querying genetic databases, analyzing patient DNA against disease risk profiles, and generating improved DNA sequences.
GeneMCP has significant definition quality gaps across all tools. While tool names follow verb_noun convention (chat_with_model, deep_research, assess_patient_genetic_risk_profile, dna_generator), descriptions are present but inconsistent in depth. Input schemas are partially visible but lack complete type information and parameter descriptions. Output schemas are not documented anywhere in the provided source. Error handling and recovery guidance are absent. The server uses FastMCP framework which provides basic registration, but the tools themselves need substantial schema and documentation improvements.
Assesses patient's genetic data against pre-compiled risk profiles for specified conditions. This tool will meticulously check every RSID provided in the condition_risk_profiles.
Chat with an AI model
Conducts deep research on genetic conditions and variants using GWAS Catalog, web search, and LLM analysis. Returns a compiled genetic research report with variant details, web summaries, and LLM-generated insights.
DNA generator using the EVO2-40B Model. Your goal is to generate a DNA sequence based on the input sequence. The input should ideally be bad genetics that require improvement. The model will generate a new sequence that is an improvement over the input.
Output schemas are completely undocumented. No tool documents what fields, types, or structure the LLM will receive in the response. chat_with_model returns a string, but deep_research, assess_patient_genetic_risk_profile, and dna_generator return complex objects with no documented schema.
Parameter descriptions are missing or incomplete. assess_patient_genetic_risk_profile has a 'condition_risk_profiles' parameter described as 'Dict[str, VariantInfo]' but the response structure of VariantInfo (rsid, risk_allele, odds_ratio, beta, pvalue, etc.) is only visible in the deep_research module, not in the tool definition itself. LLMs cannot infer the expected input structure.
No input validation or error guidance. Tools accept free-form strings (e.g., 'condition' in deep_research can be 'autism', 'diabetes', or anything else) with no enum constraint, format specification, or examples of valid inputs. Error responses are not defined. If a tool fails, there is no recovery guidance for the LLM.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 34 | - | v1 |
Parameter types are inconsistent or underspecified in schema. The 'model' parameter in chat_with_model lists valid options in the description ('deepseek', 'nemotron', 'palmyra') but does not use an enum constraint, forcing the LLM to either memorize the list or guess. The 'condition_risk_profiles' parameter is a complex object with no JSON Schema definition visible.
Tool descriptions lack composition guidance. No tool describes when to call it vs. when to call a related tool. For example, deep_research queries GWAS Catalog and performs LLM analysis, but there is no documentation on when an LLM should call deep_research vs. assess_patient_genetic_risk_profile (which is a follow-up analysis tool).
dna_generator has a poorly specified purpose. The description says 'Your goal is to generate a DNA sequence based on the input sequence. The model will generate a new sequence that is an improvement over the input.' This is vague and example-driven ('ACTGACTGACTGACTG' appears in the description, which LLMs may reuse literally). The description does not explain what 'improvement' means or what metric determines success.
No pagination or result limiting guidance. deep_research queries GWAS Catalog and web search, but the tool description does not state how many results are returned, whether there is pagination, or how large the response might be. For a genetic research tool returning potentially thousands of variants, lack of result-limiting guidance is a significant oversight.
No idempotency guarantees or confirmation steps for sensitive operations. dna_generator and assess_patient_genetic_risk_profile operate on genetic data but lack any confirmation, dry-run, or idempotency documentation. If an LLM retries a call due to a network hiccup, does it generate a duplicate report or reuse the cached result?