MCP Server for PubMed Semantic Search. Exposes the PubMed search API as MCP tools so LLMs can query biomedical literature with semantic search over abstracts using embedding-based retrieval.
Server implements 3 tools with reasonable naming (verb-first: search_papers, get_paper, find_similar) and moderate descriptions. However, output schemas are completely undocumented, critical omission for LLM reasoning. Parameter descriptions are present but sparse. Error handling is absent. The tool definitions are directly visible in src/mcp/server.py, so not inferred; however, schemas lack return type documentation entirely, forcing LLMs to guess output structure. This is a fair-quality server with clear domain focus but missing production-grade documentation patterns.
Find papers that are semantically similar to a given paper. Useful for exploring related research or finding follow-up studies.
Retrieve the full details of a specific PubMed paper by its PMID. Returns title, authors, abstract, journal, publication date, and MeSH terms.
Semantic search over PubMed biomedical abstracts. Finds papers related to a natural language query about nutrition, exercise, psychology, behavioral science, or bioethics. Returns the most semantically similar papers with titles, authors, abstracts, and relevance scores.
Output schemas completely undocumented. No return type definitions anywhere in tool specs. LLMs cannot plan downstream calls or extract required fields (pmid for chaining, similarity scores for filtering). This violates pattern:tool and pattern:response-shaper.
search_papers returns 'relevance scores' but the field is actually called 'similarity' in the formatting code (_format_paper shows paper['similarity']). Parameter names must match response field names exactly to avoid forcing LLM field mapping logic. Mismatched naming (relevance vs similarity) risks errors.
No error handling documented or visible. No guidance on what happens when query is empty, pmid is invalid, API times out, or min_date/max_date are malformed. Pattern:recovery-guide requires errors to tell LLMs what to do next.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 59 | - | v1 |
Parameter descriptions lack constraint details. top_k says 'default: 5, max: 20' but no minimum stated (assume 1?). Date format is specified as YYYY-MM-DD but no validation or error handling if agent passes '2024-1-5'. Query parameter has no length limits or content restrictions.
No idempotency guarantees documented. If search_papers is called twice with identical inputs, will it return identical results? Agents retry on failures, non-idempotent tools risk silent duplicate side effects.
Tool composition lacks pagination/result limits for list-like tools. search_papers and find_similar return lists (top_k controls count) but no documented response structure, total count, or next_cursor for large result sets. Pattern:paginated-result requires documentation of how pagination works.