arXiv MCP Server - A Model Context Protocol server for searching and retrieving academic papers from arXiv
Strong tool naming and descriptions across all 7 tools. Each tool has a clear verb-noun pattern (arxiv_search, arxiv_get_paper_details, etc.) and descriptions that explain WHAT the tool does, WHEN to use it, and dependencies on other tools. Input schemas are complete with type definitions and constraints (enums for citation formats). However, output schemas are NOT documented in the tool definitions, the source code does not show what fields are returned, their types, or structure. This is a critical gap per pattern:tool and pattern:response-shaper. Error handling is not visible in the tool definitions, and descriptions lack explicit guidance on recovery paths. Parameter descriptions are strong (average ~80 chars), but the lack of output documentation limits context for downstream reasoning and chaining.
Native integration capability. Downloads the full, original PDF of a paper directly to a specified directory on the user's local machine. Use this to save papers for the user's personal archives or if you need to pass a local file path to another tool for deep PDF extraction.
Citation generation tool. Takes a paper_id and an optional format (BibTeX, APA, MLA) and returns the perfectly formatted citation. Use this when finalizing a research report or bibliography for the user.
The deep-dive reading tool. Takes a specific paper_id and returns the full, unabridged metadata, the complete abstract, all categories, and the exact publication history. Use this *after* discovering a paper ID via search to understand its core methodology and findings.
Analytical summarization tool. Given a paper_id, this tool fetches the paper and distills the core thesis, methodology, results, and limitations into a dense, easy-to-understand bulleted summary. Ideal for rapidly digesting complex scientific literature.
The primary discovery tool for arXiv. Use this to search for papers by keywords, authors, or categories. It returns a concise, token-optimized list of results (IDs, titles, and primary authors) designed for scanning. Do NOT use this tool if you need the full abstract; use this to find relevant paper_ids first.
Output schemas not documented. Tool definitions do not specify what fields are returned, their types, or structure. LLMs cannot plan downstream calls or extract data reliably without knowing response shape.
Error handling and recovery guidance absent from tool definitions. No indication of what errors can occur, whether they are retryable, or what the LLM should do next (e.g., 'paper not found, try searching first').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
The knowledge graph tool. Given a paper_id, this automatically executes a semantic search to find similar literature. Use this to automate the "rabbit hole" workflow when a user wants more papers like a specific foundational work.
Proactive discovery tool. Retrieves the most recent or highly relevant papers across a specific domain (like Artificial Intelligence or Quantum Physics) without needing a specific keyword query. Perfect for "What's new today?" or broad literature reviews.
No pagination or result-limiting guidance. arxiv_search and arxiv_search_trending_papers may return large lists, but descriptions do not explicitly state maximum results or whether results are paginated.
arxiv_download_paper requires 'destination_dir' (absolute file path) as a parameter. This violates pattern:secret-injection and mxe:natural-identifiers, requiring agents to construct file paths is error-prone and ties the tool to local filesystem structure. Should accept a simpler identifier or use server-side path management.
Tool composition potential: arxiv_search returns paper IDs, which all other tools (get_details, get_summary, get_citations, download) expect. Descriptions mention this dependency ('use this to find relevant paper_ids first'), but output schemas are not documented, so LLMs cannot verify they're passing the correct field. This invites chaining errors.