The server defines 20 tools with explicit schemas and descriptions. All tools follow verb_noun naming conventions (get_*, search_*) and have non-empty descriptions (10-100 chars). Input schemas are present for all tools with type definitions and required field declarations. However, descriptions are uniformly brief (10-50 chars), lacking context on WHEN to use each tool, what makes them distinct from similar tools, and actionable error guidance. Parameter descriptions are minimal ('Language (default: en)', 'Page number (default: 1)') without explaining ranges, constraints, or dependencies. Output schemas are entirely absent, no documentation of what fields are returned, their types, or how to chain results to downstream tools. Error handling is present (errorReporting=true in metadata) but not visible in the code provided. The schema definitions show good structuring (properties, required fields, enums for constrained values like lang and sortOrder), but lack documentation of output structures that would help LLMs plan multi-step workflows.
Output schemas completely undocumented. No visibility into what fields are returned by any tool, their types, or structure. LLMs cannot plan downstream tool calls or extract data intelligently.
Tool descriptions are uniformly 10-50 characters, insufficient to explain WHEN to use a tool, how it differs from similar tools, or what the LLM should expect. Baseline for A-grade tools is 50-200 chars with context about use cases and dependencies.
Recommendations
Document output schemas for all tools. Create a structured definition showing what fields each tool returns, their types, and an example. This enables LLMs to reason about downstream tool calls. Example: 'Returns array of Gene objects with fields: id (string), symbol (string), ncbi_id (string), description (string), lifespan_effect (string)'.
Expand tool descriptions to 100-200 characters. For each tool, explain: (1) what it does, (2) when to use it instead of similar tools, (3) what it returns at a high level. Example: 'Search genes by disease, age-related process, or protein class. Use this to find genes relevant to a biological question. Returns paginated list of matching genes with IDs and symbols.'
Enhance parameter descriptions with constraints and examples. For 'confidenceLevel', specify valid values (e.g., 'low, medium, high'). For 'page' and 'pageSize', add ranges: 'Page number (1-based, default 1)' and 'Results per page (1-100, default 20)'.
Add numeric constraints to JSON Schema for pagination. Set minValue=1, maxValue=100 for pageSize; minValue=1 for page. This prevents LLMs from passing absurd values and documents API limits formally.
Document tool chaining paths. E.g., 'search_genes returns gene_id; pass gene_id to get_gene_by_id for detailed info'. Include field name matches to prevent impedance mismatches.
Add error recovery guidance to descriptions. E.g., 'If search returns zero results, try broadening filters or calling get_genes_by_go_term to discover available GO terms'. This guides LLM retries and discovery.
Parameter descriptions lack actionable detail. Most describe only the parameter name or default value ('Language (default: en)', 'Page number (default: 1)') without explaining ranges, valid values, format constraints, or relationships between parameters. Baseline: parameter descriptions should average 72 chars with full context.
No numeric constraints on pagination parameters. page and pageSize have no minValue, maxValue, or default enforcement visible. Unbounded numeric parameters let LLMs pass absurd values (page=999999, pageSize=10000) that could break API or cause timeouts.
Missing tool-to-tool chaining documentation. Tools return IDs (implied by 'get_gene_by_id' pattern) but output schema is not visible. Cannot verify that search_genes returns gene IDs that downstream tools like get_gene_by_id accept, or whether field names match. Broken chains force discovery detours.
No error recovery guidance visible in tool descriptions. Descriptions do not explain what errors are possible, how to interpret them, or what LLM should do next. E.g., 'search_genes' does not document whether zero results is an error, what 'confidenceLevel' filter values are valid, or how to handle API failures.
Multiple tools serve similar discovery purposes (get_gene_suggestions, get_gene_symbols, get_model_organisms, get_diseases, etc.) with nearly identical signatures. No differentiation in descriptions explains when to use one vs. another. This increases LLM confusion and token waste on tool selection reasoning.
Differentiate discovery tools with explicit use-case guidance. E.g., 'get_gene_suggestions: use to autocomplete a gene name during free-text search' vs 'get_gene_symbols: use to retrieve the full symbol list for filtering/validation'. Prevents confusion.
Consider batching variants for high-frequency patterns. If an agent often needs multiple genes by symbol, offer get_genes_by_symbols (plural) accepting an array, to reduce sequential calls.
Add field-level examples in parameter descriptions to guide LLM input. E.g., 'byGeneSymbol: comma-separated list of exact symbol matches (e.g., "TP53,BRCA1")'. But avoid sample IDs, use domain examples.
Document pagination behavior: 'Returns up to pageSize results. Check if result length < pageSize to detect end of list, or return a next_cursor field for stateless pagination'.