An MCP server for searching, retrieving, and analyzing articles from PubMed
The server has 4 tools with action verb names (search_, get_, download_) which is good, but significant quality gaps across naming clarity, parameter validation, schema completeness, and error handling. Tool descriptions are present but generic (30-40 chars each), lacking LLM optimization. Parameter schemas exist but lack formal constraints (enums, patterns, min/max). Error handling returns generic error messages without recovery guidance. The 'download_pubmed_pdf' tool description does not declare it as a WRITE operation despite modifying local state. Output schemas are not documented. No tool supports pagination despite returning lists. Overall, the server falls in the D range (50-59) but scores slightly lower due to multiple compounding gaps.
Attempt to download the full text PDF for a PubMed article.
Fetch metadata for a PubMed article using its PMID.
Perform an advanced search for articles on PubMed.
Search for articles on PubMed using key words.
Parameter descriptions are minimal or missing. 'num_results' lacks range constraints (e.g., '1-100'); 'key_words' and 'term' lack guidance on format or what queries succeed. LLMs cannot infer valid input ranges or expected content from bare descriptions.
Tool descriptions are generic and do not explain WHEN to use each tool or how they differ. 'Search for articles on PubMed using key words' (32 chars) vs 'Perform an advanced search for articles on PubMed' (49 chars) do not tell an LLM which to choose for a given user query. Optimal range is 50-200 chars with explicit use-case guidance.
Output schemas are not documented in code or descriptions. LLMs cannot plan downstream calls or extract results without knowing what fields (title, authors, abstract, pmid, etc.) each tool returns. This violates the pattern:response-shaper baseline where 100% of A+ tools document return types.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Error handling returns generic catch-all errors ('An error occurred while searching: {str(e)}') without recovery guidance. Per pattern:recovery-guide, errors must tell the LLM what to do next (e.g., 'No results found. Try with a broader search term.' or 'Invalid date format. Use YYYY/MM/DD.').
The 'download_pubmed_pdf' tool modifies local state (writes files) but the description does not declare it as a destructive operation. Per pattern:command-tool, agents must know which calls have irreversible consequences. Schema declares Risk: WRITE but description should also warn that this creates side effects.
List-returning tools (search_pubmed_key_words, search_pubmed_advanced) lack pagination support (no limit, offset, or next_cursor parameters). Large result sets will blow context windows. Per pattern:paginated-result, tools returning lists should support pagination with total count or next_cursor.
Parameter 'num_results' has no min/max constraints in descriptions. An LLM could pass 10000 or 0, potentially breaking the tool or overwhelming the API. Per pattern:constrained-input, numeric parameters must declare valid ranges (e.g., 'integer, 1 - 100').
Date parameters in 'search_pubmed_advanced' (start_date, end_date) specify format 'YYYY/MM/DD' in the description but do not validate or reject malformed input. An LLM might pass '01/15/2024' or '2024-01-15' and fail silently. Per pattern:param-validation-rules, validation errors must be explicit and actionable.
The 'pmid' parameter accepts Union[str, int] but the description and code do not clarify whether leading zeros matter (e.g., is '000123' equivalent to '123'?). Per pattern:tool-description, parameter formats must be explicitly documented to prevent ambiguity.
The prompt 'deep_paper_analysis' is defined as a prompt (not a tool) and calls 'get_pubmed_metadata' internally. Per pattern:tool, prompts should not invoke other MCP tools, they should return a template/structure that the client renders. This blurs the separation between prompts and tools, risking confusion.