MCP server for searching and extracting academic papers from arXiv with support for multiple topics, paper information retrieval, and research prompts
Two tools with basic but incomplete schemas. Tool names are action-verb clear (search_, extract_), and descriptions exist but lack depth. Input schemas are present with proper types, but parameter descriptions are minimal. Output schema documentation is missing entirely, LLMs cannot reliably infer what fields to expect in returned JSON. No error handling guidance, no retry semantics, no indication of idempotency. Both tools perform external side effects (file writes, arXiv API calls) without explicitly documenting mutation impact.
Search for information about a specific paper across all topic directories.
Search for papers on arXiv based on a topic and store their information.
Missing output schema documentation. Both tools return complex structured data (lists, JSON strings) but no schema describes what fields, types, or structure the LLM should expect. LLMs cannot reliably parse or chain results without documented output schemas.
Parameter descriptions are trivial (1-2 word phrases). 'The topic to search for' and 'The ID of the paper to look for' provide minimal context. No format guidance, range constraints, or dependency hints. Does topic accept free text? Exact phrases? Boolean operators? The description leaves LLMs guessing.
No error handling guidance. search_papers may fail if the topic is empty, max_results is negative, or the arXiv API is unreachable. extract_info may fail if the PAPER_DIR doesn't exist or files are corrupted. Both return either structured data or a plain string error message, with no indication of how LLMs should interpret or recover from failures.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Mutation not explicitly declared. search_papers has side effects (creates directories, writes JSON files). The description says 'store their information' but never explicitly states 'This tool writes to disk and modifies state.' Agents cannot determine if the tool is safe to retry or if repeated calls risk duplicate files.
max_results parameter lacks constraints. The default is 5, and the description says 'Maximum number of results to retrieve (default: 5)' but does not specify min/max bounds. Can an LLM pass 0? 1000? 1000000? Unbounded parameters invite excessive API calls and timeouts.
extract_info returns a plain string message on failure ('There's no saved information related to paper {paper_id}.'), not a structured error object. LLMs cannot distinguish success from failure, a missing paper returns a message string, while found papers return JSON. This inconsistency invites parsing errors.
No pagination or result limiting. search_papers accepts max_results but extract_info performs no search, it scans the entire PAPER_DIR filesystem, which could be slow if many topics are stored. If the directory grows large, this becomes a performance liability.