MCP server for searching and analyzing academic papers from arXiv
The server implements three read-only tools for arXiv research paper discovery and comparison. Tool naming follows good verb-noun conventions (search_papers, get_paper_details, compare_papers). All three tools have descriptions (78-96 chars), but several lack depth about when/why to use them. Input schemas are present and properly typed with JSON Schema, but parameter descriptions are generic and lack guidance on constraints or expected formats. Output is returned as JSON strings rather than structured objects, making it harder for LLMs to parse. Error handling returns JSON errors but lacks recovery guidance or actionable remediation steps. The in-memory caching and search history features add modest value but increase state complexity. No tool annotations (readOnlyHint, idempotentHint) despite all being read-only operations.
Compare 2–3 arXiv papers and return a structured comparison scaffold.
Fetch full metadata for one arXiv paper by ID.
Search arXiv for papers matching a query (optionally within a category).
Tool descriptions lack LLM optimization guidance. Descriptions are too brief (65-96 chars) and do not explain WHEN to use each tool or distinguish them from alternatives. 'Search arXiv for papers matching a query' fails to guide LLM on when search_papers is preferable over get_paper_details. Descriptions should be 100-200 chars and include use-case hints.
Parameter descriptions are minimal or absent. 'query' in search_papers just says 'keywords, e.g. "RLHF"', this is an example, not a constraint. 'category' lacks format hints (e.g., 'arXiv category codes: cs.AI, cs.LG, etc.'). 'max_results' says '1 - 25' in description but should state this as a formal constraint with type bounds. LLMs cannot infer what formats are valid.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Output schemas are not documented. Tools return JSON strings (e.g., {'query': ..., 'results': [...]}) but the response structure is not formally declared. LLMs cannot plan downstream operations without knowing what fields to expect. Baseline production tools document output schema as: 'Returns JSON with fields: query (string), category (string), results (array of {id, title, authors, published, categories, pdf_url})'.
Error responses lack recovery guidance. When search_papers returns {'error': 'max_results must be between 1 and 25'}, the LLM knows to retry but does not know if the value was too high or too low. Error responses should follow pattern:recovery-guide: 'max_results must be 1 - 25 (received: 50). Retry with a smaller value.'
No tool annotations despite read-only semantics. All three tools are read-only and idempotent, but neither readOnlyHint nor idempotentHint is set. This forces agents to assume tools may have side effects or require confirmation, leading to unnecessary caution.
In-memory state (SEARCH_HISTORY, PAPER_CACHE) violates stateless request handling. Caching improves performance but introduces hidden dependencies across tool calls. In a multi-instance agent environment, cache hits/misses become non-deterministic. Consider stateless design: either move caching to the agent/LLM context or make cache invalidation explicit.
compare_papers does not validate array bounds in input schema. The tool accepts 'arxiv_ids' as an array but the schema does not declare minItems: 2, maxItems: 3. Without formal bounds, LLMs can pass 0, 1, 4, or 100 IDs, forcing runtime validation errors instead of schema validation.