A Model Context Protocol (MCP) server that enables Claude Desktop to search for clinical trials based on genetic mutations
This server has moderate structural quality but significant gaps in tool descriptions and parameter documentation. While all 11 tools are registered with names and basic descriptions, the descriptions are often generic or lack actionable guidance for LLM selection. Parameter schemas are minimal, most tools accept only 1-2 simple string parameters with limited constraint information. The tool set mixes operational concerns (cache management, metrics reporting, circuit breaker status) with domain logic (mutation matching), making tool composition unclear. Output schemas are not documented anywhere in the visible code. Error handling and recovery guidance are absent. The server relies on STDIO transport, which is a hard architectural limitation.
Returns comprehensive cache analytics and performance metrics.
Returns a formatted cache performance report.
Returns the status of all circuit breakers, including their current state (CLOSED, OPEN, HALF_OPEN), failure counts, and recovery timers.
Returns the comprehensive health status of the MCP server and its components.
Returns current metrics in JSON format. Includes counters (API calls, cache hits/misses, errors), gauges (current values like cache size, hit rates), and histograms (request durations, response sizes, etc.)
Returns current metrics in Prometheus format. Includes all counters, gauges, and histograms suitable for scraping by monitoring systems.
No output schemas documented. The code provides input parameter definitions but never declares what these tools return. LLMs cannot plan downstream operations or extract required fields without knowing the response structure.
Duplicate or confusingly similar tools. 'summarize_trials' and 'summarize_trials_async' appear to be the same tool with identical input and nearly identical descriptions ('Primary async function' vs 'Explicit async function'). This creates ambiguity for LLM selection and violates the single-responsibility principle.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 42 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Manually trigger cache invalidation for a specific pattern.
Async batch version for multiple mutations. Query clinical trials for multiple mutations concurrently and return a combined summary.
Primary async function for summarizing clinical trials. Query clinical trials for a specific mutation and return a summary. This function uses async/await with httpx for high-performance concurrent requests.
Explicit async function for summarizing clinical trials. Identical to summarize_trials but explicitly named for async usage.
Manually trigger cache warming for common mutations.
Operational tools mixed with domain tools. Tools like 'get_health_status', 'get_metrics_json', 'get_circuit_breaker_status' are infrastructure monitoring concerns, not clinical trial matching. This conflates two separate domains and makes tool composition unclear. The LLM must reason about when to call metrics vs domain tools.
Minimal parameter descriptions. Tools like 'invalidate_cache' accept a 'pattern' parameter with description 'Cache invalidation pattern (default: "*")', no constraint information, no examples of valid patterns, no guidance on when to use wildcards vs specific patterns. Descriptions under 50 chars lack LLM-actionable detail.
No error handling or recovery guidance. Tools provide no documentation of failure modes, retryability, or what the LLM should do if a call fails. No distinction between retryable errors (timeout) vs user-fixable (invalid mutation format) vs fatal (service down).
No tool annotations for readability hints. Tools like 'invalidate_cache' and 'warm_cache' modify state, but lack destructiveHint or idempotentHint annotations. The LLM cannot distinguish which tools are safe to retry.
Insufficient constraint documentation on input parameters. The 'mutation' parameter in summarize_trials accepts a string with description 'The genetic mutation to search for (e.g., "EGFR L858R")', but no regex pattern, length limits, or valid character set. An LLM may pass invalid formats.
No pagination or result limits documented. If 'summarize_multiple_trials' can accept many mutations, there is no documentation of result size limits, pagination support, or what happens if results exceed the context window.