An MCP server that provides access to arXiv papers through their API.
This server demonstrates good definition quality with mostly complete schemas and clear descriptions. All 5 tools have descriptions and input schemas visible. Tool naming follows verb_noun convention (search_papers, get_paper_data, get_full_paper_text, list_categories, update_categories). However, there are notable gaps in parameter descriptions, inconsistent schema completeness, and missing output schema documentation. The search_papers tool has an exceptionally detailed description (~1400 chars) that may exceed optimal length, while others are more concise. Descriptions average ~200-400 chars, within baseline range. All parameters have type definitions, but several lack detailed descriptions or constraints.
Retrieve the full text of a paper from arXiv (PDF converted to Markdown).
Get detailed information about a specific paper including abstract and available formats.
List all arXiv categories and subcategories.
Search for papers on arXiv.
Update the cached arXiv category taxonomy from the official arXiv website.
Output schemas not documented. Tools return strings but lack definition of response structure (paper fields, error format, pagination if applicable). LLMs cannot infer what fields to expect from results, preventing confident downstream tool chaining.
Parameter descriptions incomplete. 'sort_by' and 'sort_order' in search_papers lack explanation of valid values despite enum constraint. 'paper_id' in get_paper_data and get_full_paper_text has generic description without format guidance (arXiv IDs have specific patterns like '2103.08220' vs 'arxiv:2103.08220').
Missing error handling guidance. No description of what happens when paper_id is invalid, API rate limits are hit, or papers cannot be fetched. LLMs receive no recovery hints (e.g., 'Try refining your search query with ti: prefix').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
search_papers description is overly long (~1400 chars, exceeds recommended max of 1024). While content is valuable, excessive detail wastes tokens and risks burying critical constraints in token-limited contexts.
Pagination and result limits not addressed. search_papers accepts max_results (1-100) but does not document total result count, next_cursor, or how to iterate large result sets. LLMs cannot plan multi-call pagination strategies.
list_categories and update_categories lack descriptions explaining their purpose and relationship. Are they discovery tools? Do they return hierarchical data? When should an agent call update_categories vs rely on cached data?