The KEGG server provides 13 well-structured tools for bioinformatics research with comprehensive input validation logic visible in the source. All tools have descriptions and documented parameter types via TypeScript interfaces. However, critical gaps exist: (1) output schemas are not formally documented in the MCP protocol sense, only internal TypeScript interfaces exist; (2) parameter descriptions in the tool definitions lack guidance on expected formats, ranges, and when to use each tool relative to others; (3) no error handling guidance visible in tool definitions to help LLMs recover from common failures; (4) generic parameter names like 'format' and 'operation' lack enum constraints in the tool definitions themselves (though validation logic exists server-side). The validation logic (isValidSearchArgs, isValidPathwayInfoArgs, etc.) demonstrates defensive input handling, but these constraints are not exposed to the LLM at registration time, forcing the LLM to discover them empirically through errors.
Tools (13)
batch_queryread only50/100
Retrieve information for multiple entries in a single batch query
brite_searchread only50/100
Search KEGG BRITE hierarchical database for functional classifications
compound_inforead only50/100
Get information about a specific chemical compound from KEGG
disease_inforead only50/100
Get information about a specific disease from KEGG Disease database
drug_inforead only50/100
Get information about a specific drug from KEGG Drug database
enzyme_inforead only50/100
Get information about a specific enzyme from KEGG
gene_inforead only50/100
Get information about a specific gene from KEGG
module_inforead only50/100
Get information about a specific KEGG module (set of functionally related genes)
Output schemas not documented in MCP protocol definitions. Internal TypeScript interfaces (KEGGPathwayInfo, KEGGGeneInfo, KEGGCompoundInfo) exist but are not exposed to the LLM at tool registration time. LLMs cannot see what fields to expect in responses, forcing discovery through trial-and-error.
Parameter descriptions lack actionable constraints and format guidance. 'search_type', 'format', 'operation', 'hierarchy_type' are validated server-side with enums but descriptions do not enumerate valid values. LLMs rely on descriptions, not source code, to understand constraints.
searchpathway_infobatch_querybrite_search
Recommendations
Add formal output schemas to every tool definition. At tool registration time, include a documented 'outputSchema' field describing fields, types, and what they represent. Example: pathway_info should declare it returns {entry: string, name: string, description?: string, genes?: {id: string, name: string}[], ...}.
Convert enum constraints into formal schema enums. For 'format' in pathway_info, change from description text to: {type: 'string', enum: ['json', 'kgml', 'image', 'conf', 'aaseq', 'ntseq'], description: 'Output format...'} at registration time.
Expand parameter descriptions with actionable guidance. Example for search: 'Search query string. Example search patterns: "gene_name hsa" (specific organism), "pathway glycolysis" (pathway by name), "compound glucose" (compound by name). Use search_type to filter results.' Include range constraints: 'max_results (1-1000, default 100)'.
Add discovery guidance to all info tools. Example for disease_info: 'Get detailed disease information. Call search first with a disease name to find the disease_id if you only have a human-readable name.' Repeat for drug_info, enzyme_info, module_info.
Document parameter relationships explicitly. For pathway_list: 'If both pathway_id and organism_code are provided, pathway_id takes precedence. Use pathway_id for specific pathways, organism_code to list all pathways in an organism.' Similar clarity for include_genes and include_compounds flags.
Add error recovery guidance to tool descriptions. Example for gene_info: 'If gene_id is not found, try search() with the gene name and organism code. Common errors: invalid organism code (use organism_list to discover valid codes), gene_id format mismatch (should be organism:number, e.g., hsa:1234).'
Score history
Overall score trend
↑ 49 points across a rubric change (v1 → v2)
49/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
49
2026-07-28+
v2
2026-03-09
F
0
-
v1
read only
50/100
List organisms available in the KEGG database
pathway_inforead only50/100
Get detailed information about a specific KEGG pathway
pathway_listread only50/100
List pathways for specific organisms or all available pathways
reaction_inforead only50/100
Get information about a specific biochemical reaction from KEGG
searchread only50/100
Search KEGG database for genes, pathways, compounds, reactions, and other entries
No error handling guidance in tool descriptions. Tool definitions do not explain how LLMs should recover from common failures (invalid ID format, API timeout, rate limit). Validation logic exists but error messages are not documented.
Tool descriptions lack differentiation and discovery guidance. Multiple tools retrieve similar data (search, pathway_list, pathway_info, gene_info all relate to pathways/genes). Descriptions do not explain when to call one vs. another, forcing LLMs to reason empirically. Missing: 'Call pathway_list to discover available pathways; then call pathway_info for details.'
Parameters lack guidance on format and discovery. Tools like disease_info, drug_info, enzyme_info, module_info require IDs (disease_id, drug_id, enzyme_id, module_id) but descriptions do not explain format or how to find them if only a human-readable name is available. This forces extra lookup calls.
Parameter relationships undocumented. pathway_list accepts both pathway_id and organism_code; descriptions don't clarify if they're exclusive, cumulative, or if one overrides the other. Same for include_genes and include_compounds flags, no guidance on combined use.
pathway_list
Document batch_query operation types with expected outputs. 'operation types: "info" (returns entry metadata), "sequence" (returns amino acid and nucleotide sequences), "pathway" (returns pathway associations), "link" (returns cross-references). Choose based on what data you need.'
Explain BRITE hierarchy types in brite_search description. 'hierarchy_type: "br" (BRITE functional classifications), "ko" (KEGG Orthology for protein families), "jp" (Pathway information in Japanese). Use "br" for functional discovery, "ko" for orthology analysis.'
Add pagination and result-limiting guidance. Document if results are capped and how to page through large result sets. Example: 'Returns up to max_results entries. To retrieve all results, iterate by incrementing offset or using cursor-based pagination if supported.'
Include timeout and rate-limit expectations. Example: 'API timeout is 30 seconds. KEGG service has rate limits; if you receive a 503 error, wait 60 seconds before retrying. Batch queries are rate-limited to 1 request per 10 seconds.'
Create a discovery tool or enhance search documentation. Add guidance like: 'Use search() to discover entries by keyword. Use pathway_list() to list all known pathways. Use organism_list() to find valid organism codes before calling other tools.' This prevents exploratory dead-ends.
Document format implications for pathway_info. 'format: "json" (structured metadata, fastest), "kgml" (XML representation, large payload), "image" (PNG pathway diagram, slow), "aaseq"/"ntseq" (sequences, very large). Choose based on use case.'
Add enum declarations to batch_query operation parameter in tool registration schema: {type: 'string', enum: ['info', 'sequence', 'pathway', 'link'], description: '...'}. Do not rely solely on description text for LLM constraint discovery.