An MCP server that provides your agents the ability to search, analyse, and explore academic papers from arXiv
Search Papers is a focused academic paper discovery tool with competent tool definitions and schemas. All 4 tools have clear verb-noun naming (search_*, advanced_search), descriptions in the 40-100 character range, and JSON Schema with type definitions and constraints. Parameters generally include descriptions and constraints (min/max for numeric fields, enums for sort options). However, there are notable gaps: (1) Parameter descriptions are terse, many lack context on expected format or use cases; (2) Output schemas are not documented, LLMs cannot infer what fields come back from search results, forcing inferred planning; (3) Error handling is present but generic, 'Error in advanced_search' provides no recovery guidance; (4) No tool annotations (readOnlyHint, idempotentHint) despite all tools being read-only operations; (5) Composition is weak, all tools return formatted text blocks rather than structured JSON, preventing downstream tool chaining. The server is competent for basic discovery but falls short of production-grade tooling patterns.
Advanced search with field-specific queries, date ranges, and sorting options.
Find all papers by a specific author with affiliation information.
Browse papers in specific ArXiv categories.
Find papers published within a specific time period.
Output schemas not documented. Handler code returns formatted text (formatPaperDetails, formatPaperSummary) rather than structured JSON. LLMs cannot know what fields to expect (e.g., paper ID, authors, date, URL) and cannot chain results to downstream tools. This violates pattern:tool and pattern:response-shaper.
Parameter descriptions are terse (most under 30 chars) and lack actionable context. E.g., 'author' says 'Author name to search for' but does not clarify: partial name allowed? Last name only? Does it accept 'Einstein, Albert' format? 'date_from' lacks timezone context. 'category' gives examples but no hint that users can call search_by_category first to discover options. This violates pattern:tool-description guidelines requiring 10-1024 chars with WHAT, WHEN, and HOW context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Error handling is generic and non-actionable. Handlers return 'Error in advanced_search: {error.message}' or 'Error searching by author: {error.message}' with no recovery guidance per pattern:recovery-guide. LLM receives no hint whether to retry, ask user, or abandon. Should return: 'No papers found, try broadening search terms or removing category filter. Available categories: cs.AI, cs.LG, ...' Example: search_by_author handler does have partial recovery logic ('No papers found for author: {author}') but others lack this.
All tools are read-only (queries to external arXiv API) but lack tool annotations. Per current MCP spec (2026-07-28), read-only tools should declare readOnlyHint to signal to agents that these are safe, non-destructive, and can be retried. None of the 4 tools declare this annotation, missing optimization signal to the LLM.
Composition is broken. Tools return formatted plain text (response += formatPaperDetails(paper, idx + 1)), making it impossible for agents to chain results. If an agent wants to 'search for papers on AI and then filter by author Einstein', it must parse the text response. Expected pattern: return structured JSON [{id, title, authors, date, url, ...}] so agents can programmatically process. This violates pattern:tool-chain.
Parameter 'max_results' is defined identically (1-50, default 10 or 20) across all tools but lacks guidance on when results are truncated and whether pagination is possible. Per pattern:paginated-result, tools returning lists should accept limit/offset and include total_count or next_cursor. Current implementation returns up to max_results but no indication if more results exist or how to fetch them.