A MCP server for searching and downloading academic papers from multiple sources.
paper-search-mcp has 4 tools with reasonable naming and acceptable descriptions, but exhibits significant gaps in parameter documentation and schema completeness. Tool names follow verb_noun convention (search_papers, download_paper, read_paper, get_available_sources), which is positive. However, input parameters largely lack type information in visible schema declarations, and output schemas are not documented. Three of four tools accept free-form string parameters (sources, paper_id, source) without enum constraints, despite sources having a known, finite set of 20+ values. Error handling is minimal, no guidance on recovery actions, no categorization of error types (retryable vs. fatal), and no actionable error messages visible. The 'sources' parameter in search_papers should be an enum rather than a comma-separated string, per pattern:constrained-input. Description quality is moderate (194-char baseline acceptable, but many descriptions lack specificity on output structure, prerequisites, and when to use). No tool descriptions explain what fields are returned or how results should be used downstream.
Download a PDF paper from a specified source by paper ID. Returns the local path to the downloaded PDF file.
Get list of all available paper sources that can be searched.
Download and extract text content from an academic paper PDF. Returns the full text extracted from the paper.
Search for academic papers across one or more sources. Supports filtering, deduplication, and results from 20+ academic databases including arXiv, PubMed, bioRxiv, Semantic Scholar, CrossRef, OpenAlex, CORE, and others.
sources parameter in search_papers accepts free-form comma-separated string instead of enum. Should declare valid values (arxiv, pubmed, biorxiv, etc.) as enum constraint. Currently invites hallucinated source names.
No output schemas documented for any tool. Unclear what fields search_papers returns, what structure download_paper response has, or what format read_paper text content takes. LLMs cannot plan downstream calls without knowing response structure.
Error handling is absent. No visible error messages, recovery guidance, or categorization (retryable vs. fatal). If a download times out or source API fails, agent gets no instruction on what to do next.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 31 | - | v1 |
download_paper and read_paper accept 'source' as free-form string without enum constraint. Should declare valid sources to prevent invalid source names from being passed.
paper_id parameter across three tools (download_paper, read_paper, and implicitly in search_papers results) has no documented format or constraints. Unclear if it's a numeric ID, alphanumeric hash, or arbitrary string. No guidance on where to obtain it.
max_results parameter in search_papers has no documented bounds. Could agent pass max_results=1000000? What's the practical limit? Should enforce 1 - 100 or similar and state in description.
Tool descriptions do not explain output format or use cases. 'Download and extract text content from an academic paper PDF. Returns the full text extracted from the paper.' does not specify: Is it plaintext, markdown, or raw PDF text with OCR artifacts? Is it paginated? How large can it be? When should this be called vs download_paper?
search_papers description mentions 20+ databases but does not explain filtering/deduplication behavior, how results are ranked, or what metadata is included. Insufficient context for agent to decide if this tool meets its need.