An MCP server for searching, reading, and analyzing arXiv papers.
The server has 6 tools with reasonable parameter schemas, but significant quality gaps prevent a higher score. Tool descriptions are present but often generic. Parameters lack depth in constraint documentation. Most critically, no output schemas are documented, the rubric requires '100% of A+ tools have documented return types.' Error handling exists but is minimal, and descriptions lack the LLM-optimized structure needed for effective tool selection. The codebase shows good intent (health_check returns structured output, search_arxiv has detailed parameter docs), but falls short of production quality. Naming is clear and verb-first throughout, which helps. However, descriptions are 50-200 chars in only 2 tools; most are under 100 chars and lack 'WHEN to use' and 'what it returns' guidance.
Download a PDF from arXiv and cache it locally.
Extract a specific section from a paper's full text.
Get the BibTeX citation for a specific arXiv paper.
Retrieve the full text of a paper from arXiv, with caching.
Return local server health and cache status without contacting arXiv.
Search for papers on arXiv and cache metadata for Resources. Args: topic: The search query. Supports advanced prefixes like 'ti:' (title), 'au:' (author), 'abs:' (abstract). max_results: Maximum number of results to return (default: 100, max: 300). offset: The index of the first result to return (for pagination). category: Optional category filter (e.g., 'cs.AI'). sort_by: Sort order. Options: 'relevance' (default), 'submitted', 'updated'. start_year: Filter by submission year (start). Set to 0 to ignore. end_year: Filter by submission year (end). Set to 0 to ignore.
No documented output schemas for any tool. The rubric requires '100% of A+ tools have documented return types.' Currently, tool responses are visible only in implementation code (e.g., search_arxiv returns a list of dicts with 'id', 'title', 'authors', etc.), but this structure is not declared in tool metadata. LLMs cannot plan downstream calls without knowing what fields to expect.
Tool descriptions lack WHEN and WHY context. E.g., 'Get the BibTeX citation for a specific arXiv paper' (52 chars) does not explain when to use it vs. alternatives, or what format the output is. Baseline expectation: 50-200 chars with clear context. Current average ~60 chars, too terse.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameters lack constraint documentation. E.g., 'sort_by' accepts 'relevance', 'submitted', 'updated', but the description states only 'Sort order. Options: ...' without explaining what each option implies (e.g., 'relevance = most similar to query', 'submitted = chronological from oldest arXiv submission'). Baseline: constraints should be actionable for LLM decision-making.
Error handling is minimal and not actionable. search_arxiv returns plain strings like 'Error searching arXiv: topic must not be empty.' but does not classify the error (retryable vs. user-fixable), nor does it suggest next steps. E.g., 'Invalid year range: start_year > end_year. Please swap them or omit end_year.'
No per-tool destructiveHint or readOnlyHint annotations visible in the tool registration. The rubric expects 'tool annotations (readOnlyHint/destructiveHint/idempotentHint)' in protocol readiness. 'download_pdf' modifies local state (WRITE risk), but no hint is set to warn agents.
No pagination guidance or result limits documented. search_arxiv accepts 'max_results' up to 300, but the description does not warn that returning 300 records wastes tokens and degrades LLM reasoning. Best practice: cap at 50, offer pagination via 'offset', and document this clearly.
Tool composition risk: extract_section requires 'paper_id' + 'section_name', but search_arxiv returns 'id' (not 'paper_id'). Field name mismatch (id vs. paper_id) forces LLM to reason about field mapping instead of directly chaining results. Baseline: 'if ReplyToMessage expects channel_id, SearchMessages must return channel_id'.