Enhanced PubMed MCP server with LLM-powered intelligent search and analysis
This MCP server implements 9 tools for PubMed research with reasonable descriptions and parameter schemas. However, several critical quality gaps prevent a higher score: (1) Tool names lack consistent verb-prefix patterns (e.g., 'analyze_paper_batch' mixes analysis + batch, 'get_mesh_tree' is clear but 'detect_open_access' should be 'get_open_access'); (2) Output schemas are NOT documented in the visible code, descriptions mention 'Returns structured metadata' or 'Returns structured analysis' but the actual response fields are not specified, violating the pattern:tool requirement; (3) Parameter descriptions vary widely in quality, some are good (search_pubmed's 'query' param) but others are vague (export_bibliography's 'format' just says 'ris' or 'bibtex' without explaining the difference); (4) Error handling is minimal, no recovery guidance or actionable error messages visible; (5) No evidence of pagination in search_pubmed despite max_results up to 100000, risking context explosion. The code shows reasonable structure (rate limiting, caching, proxy support) but tool definitions themselves need strengthening.
Analyze multiple papers using LLM-powered intelligent summarization and analysis. Requires OpenAI API key. Returns structured analysis with key findings, methodology summary, and relevance assessment.
Clear memory and/or file caches. Supports selective clearing by cache type and age threshold.
Detect open access availability for papers. Checks PubMed Central, Unpaywall, and other OA sources. Returns direct access URLs and license information.
Download full-text PDF from open access sources or via CrossRef API. Supports caching and automatic cleanup of expired PDFs.
Export paper metadata in various bibliography formats (RIS, BibTeX, etc.). Supports batch export with customizable formatting.
Retrieve caching statistics including hit rates, cache size, and performance metrics. Useful for monitoring and optimization.
Retrieve MeSH (Medical Subject Headings) tree structure for browsing and querying. Supports hierarchical navigation and keyword search.
Output schemas are not documented. Tool descriptions say 'Returns structured metadata' or 'Returns structured analysis' but do not specify the actual response field names, types, and structure. LLMs cannot plan downstream tool calls or extract data without knowing what fields to expect.
search_pubmed accepts max_results up to 100,000 with no documented pagination support (no offset, limit, next_cursor, or total_count). Returning thousands of results will blow the context window. The tool must support cursor-based or offset pagination with a reasonable default limit (20-50) and indicate the limit in the description.
Tool names lack consistent action-verb prefix patterns. 'analyze_paper_batch' combines two concerns (analyze + batch); 'detect_open_access' should be 'get_open_access' or 'check_open_access' for clarity. Inconsistent naming forces LLMs to reason about intent rather than reading the name directly.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Retrieve detailed metadata for a specific PubMed paper by PMID. Includes full abstract, authors, MeSH terms, citations, and open access information with intelligent caching.
Search PubMed database with optional LLM-powered query enhancement. Supports advanced search syntax and filters. Returns structured metadata with caching support.
Parameter descriptions vary in quality. 'abstract_mode' appears in multiple tools with the same constraint ('quick' or 'deep') but descriptions do not explain the functional difference or token/latency implications. Per the rubric, descriptions should be actionable for LLMs, e.g., 'Use quick (1500 chars) for fast retrieval; use deep (6000 chars) only if your model has >=120k context.'
Error handling is not visible in the tool definitions. No evidence of recovery guidance (e.g., 'If search returns no results, try broadening the query'), error classification (retryable vs. user-fixable), or actionable error messages. Agents need to know what to do next if a call fails.
analyze_paper_batch requires an OpenAI API key but does not document this dependency or state it as a prerequisite in the description. Parameters like 'model' are exposed (gpt-4-turbo) but the description does not explain when/why to use this tool or that it costs money. LLMs need dependency hints.
download_fulltext and export_bibliography are WRITE operations with reversible/destructible side effects (file caching, export). Tool descriptions do not mention this, and there is no confirmation step or dry-run capability. Per the rubric, irreversible operations should support a confirmation request pattern.
Enum values for parameters are not consistently enforced in schemas. 'sort' (relevance, date, journal), 'abstract_mode' (quick, deep), 'analysis_type' (summary, comparative, methodology, findings, custom), and 'format' (ris, bibtex) should be declared as enums with type:'string' and enum:[...] in the JSON Schema, not just listed in descriptions. This prevents LLMs from hallucinating invalid values.