A flexible arXiv search and analysis service with MCP protocol support
The arXiv MCP server has clear tool names that follow verb_noun conventions (search_papers, download_paper, list_papers, read_paper), which is good. However, definition quality is held back by several critical gaps: (1) Parameter descriptions lack detail about format, constraints, and valid ranges, for example, 'date_from' and 'date_to' state 'ISO format' but provide no guidance on what happens if the format is wrong or if invalid dates are passed. (2) Tool descriptions are adequate but generic, they state what the tool does but lack context about when to use it, prerequisites, or downstream dependencies. (3) Output schemas are not documented anywhere in the code; we cannot see what fields search_papers returns, what structure download_paper produces, or how results are paginated. (4) Error handling is minimal, the catch-all in call_tool() returns plain text errors with no guidance for recovery, no categorization of retryable vs. fatal errors, and no suggestions for what the LLM should do next. (5) The 'download_paper' tool is WRITE-class (triggers a download/conversion) but has no dry-run, confirmation, or idempotent hint. (6) Parameter relationships are undocumented, e.g., what happens if both date_from and date_to are omitted? Does list_papers have pagination, and if so, how? (7) No tool has @readOnlyHint, @destructiveHint, or @idempotentHint annotations. The server meets a basic bar for usability but falls short of production quality.
Download a paper and create a resource for it
List all existing papers available as resources
Read the full content of a stored paper in markdown format
Search for papers on arXiv with advanced filtering
Output schemas are not documented. The code shows tool descriptions but no structured output schema for any tool. LLMs cannot plan downstream calls or extract specific fields (e.g., arXiv ID, title, authors) without knowing what search_papers returns.
Parameter descriptions lack actionable constraints. 'date_from' and 'date_to' say 'ISO format' but do not state the expected range, whether they are required or optional when used together, or what error the tool returns if the format is invalid or dates are out of range.
Error handling provides no recovery guidance. The catch-all exception handler in call_tool() returns a plain text error message with no classification (retryable vs. fatal), no suggestions for what to do next, and no indication of what input caused the error.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 54 | - | v1 |
download_paper is a WRITE tool with no dry-run or confirmation step. The tool downloads papers and creates resources (side effects), but agents have no way to preview the action or undo it if a mistake occurs. This violates the confirmation-request pattern for irreversible operations.
No tool annotations present. Tools lack @readOnlyHint, @destructiveHint, or @idempotentHint annotations. This makes it harder for agents and clients to reason about safety and side effects at call time.
Pagination is not documented. list_papers and search_papers (with max_results) do not document pagination behavior. Does search_papers return all results up to max_results, or is there a page/offset parameter? Can an LLM request page 2 if page 1 returns 20 results?
Tool descriptions are generic and lack context about downstream dependencies. 'Download a paper and create a resource for it' does not explain what a 'resource' is, whether download_paper must be called before read_paper, or what fields are returned for use in subsequent calls.
Parameter relationships are not documented. For search_papers, if both date_from and date_to are provided, what if date_from > date_to? If only one is provided, does the tool default the other to today or some other value? The categories parameter is an array but lacks guidance on valid category values.