MCP server for searching and analyzing arXiv research papers
This server demonstrates solid tool design with clear verb-based naming, mostly complete schemas, and coherent parameter definitions. All 8 tools are explicitly registered with names, descriptions, and input schemas. Descriptions are generally actionable (ranging 80-230 chars, well within the 10-1024 target). Parameters have type declarations and purpose-driven descriptions. However, there are notable gaps: (1) output schemas are not documented in tool definitions, LLMs cannot see what fields to expect from responses; (2) no pagination documentation despite search tools returning lists (search_papers accepts max_results and start, but no total_count or next_cursor is mentioned); (3) parameter descriptions lack format constraints (e.g., arxiv_id format, valid chunk_index ranges); (4) no error handling guidance in descriptions; (5) parameter validation rules not documented in descriptions. The naming is consistently strong (search_papers, get_paper, get_paper_pdf_text, search_by_author, get_citations, get_related_papers, summarize_paper, compare_papers, all action-verb driven). Schemas are valid JSON Schema with proper types. But absent output documentation and error guidance, the score plateaus at a solid B.
Fetch metadata for 2-3 papers and return them in a format suitable for comparison. The LLM can then analyze and create a comparison table of methods, datasets, and results.
Extract references and bibliography from a paper's PDF. Returns a list of citations found in the paper.
Fetch full metadata for a specific arXiv paper by ID. Returns complete details including abstract, authors, categories, PDF URL, comments, and journal references.
Download and extract full text from a paper's PDF. Returns structured text, chunked if necessary to fit context limits.
Find papers related to a given arXiv paper based on title keywords and category. Useful for literature review and finding similar work.
Find all papers by a specific author, sorted by publication date (most recent first).
Output schemas not documented. Tool definitions lack documented return types, forcing LLMs to guess the response structure. This violates the critical check: 'Document the output schema. LLMs need to know what fields to expect.' Without knowing field names, LLMs cannot confidently chain calls (e.g., extracting paper IDs from search_papers to pass to get_paper).
Parameter constraints and format specifications missing from descriptions. arxiv_id accepts both '2301.07041' and 'cs/0001234' formats but neither the description nor schema defines a regex pattern or format property. chunk_index defaults to 0 but no maximum is documented. author parameter accepts full names, last names, and abbreviations (per description example 'Bengio') but no disambiguation strategy is described.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Search arXiv for research papers by query, author, category, or date range. Returns paper metadata including title, authors, abstract, and arXiv ID.
Fetch a paper and return its content structured for summarization. The LLM can then analyze the text to extract objective, method, results, and conclusions.
Pagination not fully documented. search_papers and search_by_author accept max_results and start (pagination params), but descriptions do not clarify total_count or whether there is a next_cursor. For large result sets, LLMs cannot determine if they retrieved all results or need to paginate further.
No error handling guidance. Descriptions do not tell LLMs what to do if an arxiv_id is not found, PDF parsing fails, or an author name is ambiguous. Per critical check: 'Error responses must tell the LLM what to do next.' Absent this, LLMs cannot plan recovery.
Tool composition clarity: summarize_paper and compare_papers descriptions are ambiguous about whether the tool returns a summary/comparison or just structured input for the LLM to analyze. 'Fetch a paper and return its content structured for summarization' could mean either. This forces the LLM to infer the boundary between tool responsibility and LLM reasoning.