A MCP server for searching and downloading academic papers from multiple sources.
Academic MCP has three well-intentioned tools for scholarly research, but exhibits significant gaps in definition quality. Tool naming is clear and action-oriented (paper_search, paper_download, paper_read), but descriptions lack depth beyond basic function statements. Input schemas are present and properly typed with Pydantic validation, but output schemas are not documented anywhere in the source code. Parameter descriptions are competent but could be more explicit about constraints and error conditions. The server follows reasonable patterns for constrained input (e.g., query length 1-500, max_results 1-100) but lacks actionable error guidance and security considerations around downloaded PDF handling.
Download PDF file for a specific academic paper from the search results.
Read and extract text content from an academic paper PDF.
Search academic papers from multiple sources. ## Available sources: arxiv, pubmed, pmc, biorxiv, medrxiv, google_scholar, iacr, semantic, crossref, sciencedirect, springer, ieee, scopus, acm, wos, jstor, researchgate, core ## Input Constraints: - query: 1-500 characters, required, cannot be empty - max_results: 1-100, default is 10 - year: Valid formats: '2019', '2016-2020', '2010-', '-2015' (only for semantic) - fetch_details: boolean (only for iacr) - kwargs: dict (only for crossref)
paper_download and paper_read have minimal descriptions (one sentence). Descriptions lack WHEN to use the tool, expected return structure, side effects (file I/O), and error cases. The 40-character baseline rubric penalizes descriptions under 20 chars, but these at 10-15 chars each fall below even that threshold.
No output schemas documented. The source code shows no return type annotations, response field documentation, or explanation of what paper_search returns (list of Paper objects?), what fields are included, pagination behavior, or whether results include download URLs. This violates the pattern:tool and pattern:response-shaper requirements that agents need to know what fields to expect.
paper_search accepts 18 different academic sources (arxiv, pubmed, pmc, etc.) but the description does not explain which sources return what data types, which are free vs. paywalled, which have rate limits, or which may fail silently. An LLM selecting 'springer' or 'ieee' needs to know if those APIs require authentication (suggested by line imports but not documented).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
paper_download and paper_read do not document error cases or recovery paths. What happens if a paper ID is invalid? What if the PDF is DRM-protected or too large? The rubric pattern:recovery-guide requires error responses to tell the LLM 'what to do next,' but the tool descriptions provide no guidance.
paper_download has 'WRITE' risk (downloads files to ACADEMIC_MCP_DOWNLOAD_PATH), but the description does not state this or warn about side effects, disk quota, file retention, or permissions. Agents need to know irreversible operations modify state. Missing risk disclosure violates pattern:command-tool.
No pagination support documented. paper_search returns max_results (1-100) but does not explain if results are paginated, whether there is a total count, whether a cursor or offset parameter exists for subsequent calls, or what happens if a query matches >100 papers. This violates pattern:paginated-result.
Parameters like 'paper_id' and 'searcher' in paper_download and paper_read lack explicit documentation of how paper_id varies by source. If paper_search returns an arxiv ID like '2301.12345' but a Semantic Scholar ID like '123456789', the tool description must clarify that paper_id is source-specific. Without this, agents misuse IDs across sources.
No permission or scope declaration. The tools access external APIs (arXiv, PubMed, Google Scholar, IEEE, Scopus, etc.) some of which may require authentication, API keys, or institutional access. The source code does not document which credentials are required, how they are injected, or what the blast radius is if an API key leaks. This violates pattern:scope-declaration and pattern:secret-injection.