A flexible arXiv search and analysis service with MCP protocol support
The server defines 4 tools with visible schemas and descriptions, but significant quality gaps limit production readiness. Tool names follow verb_noun convention adequately (search_papers, download_paper, list_papers, read_paper), but descriptions are generic and lack LLM-optimization. Parameter descriptions exist but lack specificity around constraints, formats, and valid values. No output schemas are documented. Error handling is generic catch-all with minimal recovery guidance. The codebase shows tracing capability but lacks formal composition patterns.
Download a paper and create a resource for it
List all existing papers available as resources
Read the full content of a stored paper in markdown format
List available arXiv research tools.
Tool descriptions are vague and under 50 characters. 'List available arXiv research tools' for search_papers is generic and does not explain WHEN to use this tool vs. list_papers, WHAT constraints apply to 'query', or what response structure to expect. Under 20 chars rule triggers 0-20 cap; these are 25-60 chars but lack specificity.
No output schemas documented. LLMs cannot plan downstream calls or extract results. E.g., search_papers response structure is invisible, does it return paper_id, title, authors, date, abstract? Agents must guess or fail.
Parameter 'sort_by' (search_papers) lacks enum constraint. Description says 'relevance or date' but does not enforce it as an enum. LLM might pass 'newest', 'recent', or invalid values. No min/max on max_results either.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | <=2025-11-25 | v2 |
Destructive operation (download_paper) has no confirmation or dry-run mode. 'check_status' boolean is a weak proxy. No description explains the side effects (downloads to disk, consumes quota, creates resources). Agents may trigger unintended downloads.
Error handling is a generic catch-all in call_tool() that returns 'Error: {str(e)}'. No recovery guidance, no distinction between retryable vs. fatal errors, no actionable remediation. E.g., if paper_id is invalid, the LLM gets a raw exception instead of 'Paper not found. Try search_papers() first.'
No tool-to-tool composition guidance. The relationship between search_papers → download_paper → read_paper is implicit. No mention of required IDs, field naming consistency, or chaining. If search_papers returns 'arxiv_id' but download_paper expects 'paper_id', silent failures occur.
Parameter descriptions lack format/constraint details. 'paper_id' is described as 'The arXiv ID of the paper' but no format is specified (e.g., 'YYMM.NNNNN format, e.g., 2301.12345'). 'query' has no length limits, special character handling, or example.