A comprehensive Model Context Protocol (MCP) implementation for interacting with arXiv.org
The arXiv MCP server has 11 tools with explicit schemas and descriptions. Strengths: all tools have consistent verb_noun naming (search_arxiv, get_paper, summarize_paper, etc.), descriptions are present and reasonably informative (100-200 chars typical), and input schemas are complete with proper JSON Schema format including types, descriptions, and constraints (enums, min/max values). The server provides structured input validation but incomplete composition guidance for multi-tool flows.
Analyze publication trends in a specific field or category
Compare multiple papers side by side
Export papers in various formats (BibTeX, JSON, CSV, Markdown)
Find papers related to a given paper based on categories and keywords
Get detailed information about a specific arXiv paper
Get citation information and metrics for a paper (simulated)
Get the most recent papers from arXiv
OUTPUT SCHEMAS NOT DOCUMENTED: None of the 11 tools document their return type, structure, or fields. LLMs cannot reason about what data is available downstream or plan multi-step tool chains.
MISSING PAGINATION GUIDANCE: search_arxiv, search_by_author, search_by_category, get_recent_papers all accept max_results but lack documentation of how to handle large result sets, whether next_cursor or offset is returned, or how results are truncated.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 54 | - | v1 |
Search for academic papers on arXiv
Search for papers by a specific author
Search for papers in a specific arXiv category
Get a formatted summary of an arXiv paper
TOOL DISAMBIGUATION MISSING: Three search tools (search_arxiv, search_by_author, search_by_category) are similar enough that LLMs may confuse them. Descriptions lack explicit guidance on WHEN to use each.
NO ERROR HANDLING GUIDANCE: Tools lack error recovery instructions. E.g., if arxiv_id is invalid, what should happen? Should the LLM search first?
DESCRIPTIONS LACK ACTIONABLE CONTEXT: Many descriptions are functional but generic. E.g., 'Get a formatted summary of an arXiv paper' says WHAT but not WHY an LLM would choose it vs get_paper. State WHAT the tool does, WHEN to use it, and any prerequisites.' Descriptions should be 50-200 chars with clear intent signals.
PARAMETER CONSTRAINT DESCRIPTIONS INCOMPLETE: While enums are declared (e.g., sort_by in search_arxiv), some parameter descriptions lack actionable guidance. E.g., 'category' in search_by_category has examples but no explanation of what a valid category looks like or how to discover valid options.
NO TOOL ANNOTATIONS: None of the 11 tools use readOnlyHint, destructiveHint, or idempotentHint annotations. Per current MCP spec (2026-07-28), tool annotations enable better agent planning. While not critical, their absence is a missed signal for client optimization.
COMPOSITION RISKS: Tools like compare_papers and export_papers require arxiv_ids but lack clear guidance on what ID format is valid (new format '2301.07041' vs legacy 'cs.AI/0001001').