An MCP server that fetches and processes arXiv papers using LaTeX source for accurate equation handling
The server provides four well-named tools with consistent schemas and descriptions. All tools follow verb_noun naming (get_*, list_). Descriptions are adequate (ranging 150-320 chars) and explain WHEN to use each tool vs alternatives, which is stronger than most community servers. Schemas are present and properly typed. However, there are notable gaps: (1) No output schema documentation, LLMs cannot see what fields to expect from each call; (2) Parameter descriptions lack format/constraint details (e.g., 'arxiv_id' accepts '2403.12345' but format and validation rules are undocumented); (3) No explicit error handling guidance, the code catches exceptions but returns generic text without actionable recovery steps; (4) No pagination support despite tools potentially returning large results (full paper, all sections); (5) Security: no rate-limiting, input validation, or timeout handling visible in the call_tool handler. The tool composition is sound (each tool does one thing), and naming clearly distinguishes tools (get_paper_prompt vs get_paper_abstract vs list_paper_sections vs get_paper_section). The descriptions demonstrate good intent for agent guidance, e.g., 'Use this for a quick preview when the user hasn't read the paper yet', but lack the depth required for production use.
Get just the abstract of an arXiv paper. Use this for a quick preview when the user hasn't read the paper yet, not when they provide an arXiv ID to discuss a paper.
Recommended default: fetch the full LaTeX source of an arXiv paper for precise interpretation of mathematical expressions.
Get a specific section of an arXiv paper by section path. Use when the full paper is too long for context or the user wants to focus on a particular section. Use list_paper_sections first to find available paths.
List section headings of an arXiv paper. Useful when the full paper is too long for context and you need to identify which sections to fetch individually.
No output schema documented. LLMs cannot predict what fields or structure will be returned from each tool call. This forces them to reason about response contents on the fly, increasing hallucination risk and wasting tokens on uncertainty.
Parameter 'arxiv_id' lacks format specification. Description says 'e.g., 2403.12345' but does not state if format is required, if leading zeros matter, if it can include version (v1, v2), or what happens if malformed. LLMs will guess and may pass invalid IDs.
Error handling in handle_call_tool() is generic. Code catches exceptions but the except clause (truncated in source) likely returns a bare error message without actionable recovery steps. LLMs need guidance: 'Try a different arxiv_id' or 'This paper may not have LaTeX source available' or 'The service is temporarily unavailable, retry in 60s'.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 48 | - | v1 |
No pagination or result limits declared. A full paper's LaTeX source could be tens of thousands of tokens. The tool description does not mention limits or how to handle papers too large for context. LLMs may request the full source and blow the context window.
No input validation or timeout handling visible in the call_tool handler. If arxiv_id is malformed or the arXiv API hangs, the tool will fail silently or with no guidance. Production tools should validate inputs early and set explicit timeouts for external calls.
Logging feature is present (logging=true) but relies on the deprecated 'logging/setLevel' pattern and 'set_logging_level' handler. The current MCP spec (2026-07-28) has removed stateful initialize handshake and server-initiated logging. Migration: log to stderr or adopt OpenTelemetry; remove 'set_logging_level' handler.