Search, download, and summarize academic papers from arXiv. Built for AI/ML researchers. Find papers by query, author, or category.
arXiv MCP demonstrates basic tool structure with clear naming and functional schemas, but falls short of production quality due to incomplete descriptions, missing parameter documentation, and lack of error guidance. All 6 tools follow verb_noun naming conventions (search, get_paper, download_pdf, save_paper, list_saved, update_status), which is good. However, parameter descriptions are sparse or absent in several tools, and output schemas are not formally documented. The server relies on MongoDB integration without clear fallback behavior or error messages. Tool descriptions average ~40 characters, which is below the baseline of 194 chars for A+ tools. Parameter descriptions are either missing entirely or generic. No tools provide actionable error messages or recovery guidance. The schema definitions visible in the source are basic JSON Schema with types and defaults, but lack the depth needed for robust LLM tool selection.
Download PDF for a paper
Get details for a specific paper by arXiv ID
List saved papers from MongoDB
Save paper to MongoDB reading list
Search arXiv for papers matching query
Update paper status (to-read, reading, read, cited)
Parameter descriptions are missing or minimal across all tools. For example, 'search' has a 'max_results' parameter with only 'Max results' as description, no guidance on valid range (1-100?), and no warning that high values may timeout. 'sort_by' lacks explanation of when to use each option.
Tool descriptions are too brief (average ~40 chars vs. baseline 194 chars). For example, 'get_paper' is described only as 'Get details for a specific paper by arXiv ID' with no guidance on when to use it vs. search, whether it returns the full abstract, or what happens if the ID is invalid.
No output schema documentation. Tools return unspecified dictionaries; LLMs cannot plan downstream calls without knowing what fields to expect. For example, 'search' returns a list of papers but the response structure is not documented in the tool definition.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Error handling provides no recovery guidance. When MongoDB is not configured or a paper is not found, the server returns bare error strings ('MongoDB not configured. Paper not saved.', 'Paper {arxiv_id} not found'). LLMs cannot reason about what to do next.
No input validation or actionable error messages. The 'arxiv_id' parameter accepts free-form strings and the code attempts to 'clean' them via string replacement, but provides no feedback if the ID format is invalid. LLMs cannot self-correct.
'download_pdf' and 'save_paper' are write operations but descriptions do not explicitly state they modify state. This violates the requirement that destructive/write tools declare their impact so agents know they cannot be safely retried without side effects.
'status' parameter in 'save_paper' and 'update_status' has an enum constraint (to-read, reading, read, cited) defined in the schema, but the enum values are not mentioned in the parameter description. LLMs must infer valid values from JSON Schema, which is error-prone.
'download_pdf' has an 'output_dir' parameter with a default of './papers', but no guidance on whether relative or absolute paths are supported, whether the directory will be created if it does not exist, or what happens if the agent lacks write permissions.
'search' limits abstracts to 500 characters in the response ('abstract: paper.summary[:500]'), but the tool description does not mention this truncation. 'get_paper' with full=True returns the complete abstract. LLMs cannot predict which tool provides the full text.
'list_saved' accepts an optional 'status' parameter but the description says 'Filter by status (optional)' without documenting what values are allowed or what happens if an invalid value is passed.