arXiv aggregator MCP server providing AI assistants with access to arXiv paper search, recommendation, and summarization capabilities through the Model Context Protocol
Arxiver has 9 tools with reasonable naming (all verb-first: search_, get_, import_, summarize_, recommend_). Descriptions are present but inconsistent in depth. Input schemas are defined with proper JSON Schema types and descriptions for all tools. Output schemas are documented via Pydantic models (Paper, SearchResponse, SummaryResponse, etc.), which is good, but not all tools explicitly declare their return type in the source. Error handling is minimal, no recovery guidance or actionable error messages visible. Security: no sensitive parameter exposure detected, but no permission gates or audit logging visible. Tool composition is generally sound (single responsibility), though some tools could be more focused (e.g., recommend_papers criteria param is vague). Parameters have reasonable defaults and descriptions. Major gaps: parameter validation is missing (no enums for 'category', no length limits for 'query'), and error messages are not visible in the code provided.
Get detailed information about a specific paper by its arXiv ID
Get recently published papers from the past N days
Import and store a paper from arXiv by its ID into the local database
Get recommended papers based on specified criteria or ML model predictions
Search for papers by author name
Search for papers by arXiv category
Search for papers in the arXiv database using text queries
recommend_papers has a vague 'criteria' parameter with no enumerated values or format constraints. The description says 'e.g., category, keywords, similar to paper_id' but does not formalize the syntax. LLMs will guess and pass invalid values.
search_papers and search_by_author accept 'query' and 'author_name' as free-form strings with no length limits, regex patterns, or validation rules documented. LLMs may pass extremely long queries or special characters that break the underlying API.
Error handling is not visible in the source code provided. No recovery guidance, actionable error messages, or error categorization (retryable vs. fatal) can be verified. Agents will not know how to recover from failures.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
Get or generate a concise AI-powered summary of a paper
Search for papers using semantic vector similarity search on paper summaries
search_by_category accepts 'category' as a free-form string. arXiv categories are well-defined (cs.AI, cs.LG, stat.ML, etc.), but no enum or validation is declared. LLMs will hallucinate invalid categories.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in the source. While tool risk is stated in metadata (import_paper and summarize_paper are WRITE), the MCP server does not register these hints in the tool definitions.
No audit logging or permission gates visible. Tools like import_paper and summarize_paper perform writes but do not log who called them, when, or from where. No permission checks are declared.
recommend_papers criteria parameter is ambiguous. Accepting 'similar to paper_id' as a criteria value mixes input types (string) with functional intent. Should be split into separate tools (recommend_similar_papers, recommend_by_category) or use a structured param object.
Numeric parameters (max_results, top_k, limit, days, days_back) lack explicit min/max bounds in descriptions. Unbounded parameters let LLMs pass absurd values (limit=999999, days=99999) that break performance or exhaust resources.