A comprehensive MCP server for searching and analyzing academic papers from arXiv with AI-powered relevance ranking, full text extraction, and caching
The arXiv Research MCP server has significant definition quality gaps. Tool naming is generally adequate (verb-forward: search_, get_, clear_) but descriptions vary widely in completeness and clarity. Most critically, the server exposes 5 tools with incomplete schema documentation and missing parameter descriptions. The inferred tool definitions (from integrations/langchain_tool.py) lack explicit visible registration in source control, making it difficult to verify actual implementation. While some tools like search_arxiv_papers have reasonable parameter documentation, others (clear_cache, get_cache_stats, arxiv_cache_management) have minimal or generic descriptions. No output schemas are documented in the tool definitions, LLMs cannot know what fields to expect in responses. Error handling guidance is absent entirely. This is typical of community research tools but falls well short of production quality.
Manage the cache for arXiv search results. Available actions: - 'stats': Get cache statistics - 'clear': Clear all cached results. Input should be either 'stats' or 'clear'.
Search for recent academic papers on arXiv with AI-powered relevance ranking. This tool searches arXiv for papers matching your query and returns: - Papers from the specified time period (default: last 4 years) - Relevance-ranked results using TF-IDF similarity - Full paper text when available - Abstracts, authors, publication dates, and arXiv links. Input should be a JSON object with: - query: research topic (required) - max_results: number of papers to return (default: 10) - years_back: years to search back (default: 4) - include_full_text: whether to extract full text (default: true)
Clear all cached search results
Get cache statistics and information
Search arXiv for academic papers with relevance ranking and full text extraction
No output schemas documented for any tool. LLMs cannot plan downstream operations or extract required fields. This forces empirical discovery and increases hallucination risk.
Destructive operations (clear_cache, arxiv_cache_management with 'clear' action) lack confirmation or dry-run patterns. Agents can accidentally wipe caches without warning.
No error handling guidance in any tool description. What happens if arXiv API is down? If query is malformed? If cache is locked? No recovery instructions for LLMs.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Tool definitions in integrations/langchain_tool.py appear to be wrapper integrations, not canonical MCP tool registrations. Actual tool implementation and schema details cannot be verified from source.
Duplicate/conflicting tools: search_arxiv_papers and arxiv_research appear to do nearly the same thing. No disambiguation in descriptions explains when to use which.
Parameter descriptions lack detail on formats, constraints, and valid ranges. E.g., 'years_back' (integer) has no minimum/maximum documented. max_results has no guidance on practical limits.