Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
AlphaGenome MCP server demonstrates solid definition quality with comprehensive parameter schemas and detailed descriptions. All 7 tools have explicit input schemas with full type definitions and descriptions. Tool names follow verb_noun conventions (predict_*, list_*, score_*, validate_*). Descriptions are substantive (150-300+ characters), providing context on what each tool does and when to use it. However, there are notable gaps: (1) Output schemas are not documented, we can see input schemas in pydantic models, but return types are not visible in the code excerpt; (2) No error handling guidance, tools lack recovery suggestions or error classification (retryable vs fatal); (3) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite all tools being READ_ONLY; (4) Parameter enums are described as lists in text rather than declared as JSON Schema enums, forcing LLMs to parse text rather than selecting from constrained options; (5) Dependencies between parameters (e.g., ontology_terms only valid with certain organisms) are described but not formally enforced. The code shows fastmcp usage and Pydantic BaseModel definitions, indicating proper schema structure, but actual MCP tool registration details are not visible in the excerpt provided.
Tools (7)
list_supported_organismsread only50/100
List supported organisms for AlphaGenome predictions
list_supported_outputsread only50/100
List available genomic output types for AlphaGenome predictions
Output schemas not documented. Input schemas are comprehensive Pydantic models, but return types for all 7 tools are not visible in the provided code. LLMs cannot plan downstream tool calls or extract result fields without knowing what fields are returned.
Enumerated parameters declared as prose instead of JSON Schema enums. 'requested_outputs' is described as a list of allowed strings (e.g., ['RNA_SEQ', 'CAGE', ...]), but these are not constrained in JSON Schema. LLMs may hallucinate invalid output types, requiring error recovery.
No tool annotations despite clear intent. All 7 tools are marked Risk: READ_ONLY, but the fastmcp implementation does not use the current MCP readOnlyHint annotation. Tools should declare their side effects explicitly so agents know which calls are safe to retry.
Recommendations
Document output schemas for all tools. For predict_sequence, predict_interval, predict_variant, and score_variant, explicitly describe return types (e.g., 'Returns a PredictionResult object with fields: predictions (dict of output_type→tracks), confidence_scores (dict), processing_time_ms (int)'). For list_supported_organisms and list_supported_outputs, document the list structure (e.g., 'Returns {organisms: [{name: str, reference_genome: str}]}'). This enables LLM planning.
Convert prose enum lists to JSON Schema enums. For predict_sequence, predict_interval, predict_variant, and score_variant, declare requested_outputs as 'enum': ['RNA_SEQ', 'CAGE', 'DNASE', 'ATAC', 'CHIP_HISTONE', 'CHIP_TF', 'SPLICE_SITES', 'SPLICE_SITE_USAGE', 'SPLICE_JUNCTIONS', 'CONTACT_MAPS', 'PROCAP'] in the schema, not just description text.
Add tool annotations using fastmcp's readOnlyHint. Explicitly mark all 7 tools as read-only so agents know they can safely retry without side effects: @mcp.tool(description=..., name=..., readOnlyHint=True).
Add error handling guidance to tool descriptions. For predict_sequence, add: 'Returns error if sequence contains non-ACGTN characters or is not between 2KB and 1MB after resizing. Call validate_sequence first to check compatibility.' For predict_interval, add: 'Returns error if chromosome is invalid (call list_supported_organisms for valid chromosomes), or if interval length is outside supported range.'
Clarify and formalize coordinate systems in predict_variant. Update description to emphasize: 'interval_start and interval_end are 0-based (inclusive/exclusive), but variant_position is 1-based. Constraint: interval_start <= variant_position-1 < interval_end.' Consider adding a separate helper tool predict_variant_with_sequence that accepts raw sequence instead of coordinates to reduce confusion.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
No error handling guidance. Tools lack descriptions of what errors can occur (e.g., invalid organism, unsupported sequence length, invalid chromosome) and how to recover. LLMs get raw errors with no guidance on next steps.
Confusing coordinate system mixing in predict_variant. The tool mixes 0-based coordinates (interval_start, interval_end) with 1-based variant_position. The description attempts to clarify this, but it's error-prone and should be enforced in schema constraints. Currently, 'interval_start < variant_position-1 < interval_end' is only documented in text, not validated.
Minimal descriptions for discovery tools. list_supported_organisms and list_supported_outputs have 50-char descriptions that do not explain when to call them or what insights they provide. Agents don't know if they should call these before other tools.
list_supported_organismslist_supported_outputs
Enhance discovery tool descriptions. For list_supported_organisms, change description to: 'List organisms and their reference genomes. Call this first if unsure whether to use HOMO_SAPIENS or MUS_MUSCULUS.' For list_supported_outputs, add: 'List available prediction output types and their descriptions. Call before predict_sequence if you need to know what genomic features can be predicted.'
Add schema constraints for strand, organism, and chromosome parameters. For strand, declare enum: ['POSITIVE', 'NEGATIVE']. For organism, declare enum: ['HOMO_SAPIENS', 'MUS_MUSCULUS']. For chromosome, consider adding a regex or reference to list_supported_organisms (e.g., description: 'Chromosome identifier. Call list_supported_organisms for valid values for your organism.').
Document the relationship between ontology_terms and available tissues. Add a reference to a discovery resource or expand the description to note: 'UBERON terms filter predictions to specific tissues. Not all tissues are available in all prediction modes. If ontology_terms is omitted, predictions are tissue-agnostic.'
Add examples to parameter descriptions, formatted as 'Example: ' rather than 'e.g.' to reduce hallucination risk. E.g., for chromosome: 'Chromosome identifier. Example: chr1 (not 'chromosome 1' or '1'). Valid values: chr1 - chr22, chrX, chrY, chrM.'
Consider adding a validate_variant helper tool. Before agents call predict_variant or score_variant, they could validate the variant specification (chromosome, interval, alleles) to catch errors early and provide actionable feedback.