MCP server providing genomic data analysis tools for the 1000 Genomes Project data, including variant queries, sample metadata, kinship analysis, and population statistics
OneKGPd MCP Server demonstrates strong domain expertise in genomic querying with 11 well-named, specialized tools. All tools start with action verbs (get_, find_, compute_) and have clear descriptions (average ~120 chars, well within 10-1024 range). Input schemas are present and properly typed for all tools. However, the server lacks error handling guidance, output schema documentation, and tool annotations. The findVariants tool is exceptionally complex with 28 parameters, but they are all properly described with clear constraints (CSV values, numeric ranges, boolean flags). Missing are recovery hints, permission declarations, and explicit output field documentation. Pagination is properly handled (skip/limit/countOnly pattern), which is a strength. No security issues detected (no API keys in params). Overall quality is solid but not exceptional, strong naming and schema structure offset by missing error guidance and annotation patterns.
Calculate kinship coefficients and relatedness degree between two samples
Calculate kinship coefficients for a trio (parent1, parent2, child) to determine biological relationships
Identify de novo variants in a child that are absent in both parents
Find heterozygous dominant variants (present in one parent and child, absent in other parent)
Find homozygous recessive variants (present as homozygous in child, heterozygous in both parents, absent in unrelated samples)
Find samples that are homozygous reference at a given genomic position
Find samples carrying a specific variant (heterozygous or homozygous)
No output schema documentation visible in tool definitions. Tools like getDatasetInfo, getPopulationStats, and getSuperpopulationStats lack explicit return type specifications, forcing LLMs to guess at output structure.
Missing error handling guidance. No recovery hints, retryable error classification, or suggestions for alternative tools when a resource is not found. E.g., if a sample ID is invalid, what should the LLM do?
No tool annotations present. Missing readOnlyHint, destructiveHint, or idempotentHint attributes. While all tools are read-only (safe), explicit annotations would improve agent planning and safety.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2025-06-18+ | v2 |
Query variants in genomic regions with filtering by annotations, allele frequency, functional impact, and genotype
Retrieve metadata about the genomic dataset including total samples, variants, and sex distribution
Retrieve population statistics including allele frequencies and sample counts
Retrieve superpopulation statistics including sample counts and summary information
findVariants has 28 parameters. While well-described, this complexity could benefit from parameter grouping, composite filter objects, or a query builder pattern to reduce cognitive load on LLMs.
No explicit permission declarations or scope requirements documented. Tools should declare minimum required permissions (e.g., 'read:genomic-data') for audit and least-privilege agent configuration.
Parameter dependencies undocumented. In findVariants, parameters like 'selectHet' and 'selectHom' are interdependent (at least one should be true for meaningful results), but the description does not explain the relationship.