An MCP server for searching and managing academic papers from arXiv with research topic organization and paper extraction capabilities.
The server has 2 tools with reasonable naming but significant gaps in descriptions, schema documentation, and error handling. Tool names follow verb_noun convention (search_papers, extract_info), which is good. However, descriptions lack depth about dependencies and prerequisites, parameter descriptions are minimal, and output schemas are not formally documented. The tools also lack error handling guidance for the LLM. The server includes prompts and resources, which adds some value, but core tool quality falls short of production baseline.
Search for information about a specific paper across all topic directories.
Search for papers on arXiv based on a topic and store their information.
extract_info has vague description ('Search for information about a specific paper across all topic directories'), does not explain when to use it vs. what it returns. Lacks dependency hint that search_papers must be called first to populate the data directory.
Output schemas not documented. Both tools return unstructured strings (List[str] for search_papers, plain string JSON for extract_info). LLM cannot predict the structure of the returned paper information without seeing the underlying JSON format. Response shape should be formally defined.
No error handling guidance. If paper_id is not found, extract_info returns a plain string message instead of a structured error that tells the LLM what to do next (e.g., 'Paper not found. Try searching with search_papers() first.'). No recovery hints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
search_papers parameter 'max_results' lacks validation hints. Description says 'default: 5' but does not state range (min 1? max 100?). LLM could pass unreasonable values like 1000, causing excessive API calls to arXiv.
search_papers description does not state that it has side effects (creates/writes files to disk). Per pattern:command-tool, descriptions of state-modifying tools must explicitly say so. This affects agent reasoning about retryability and idempotence.
Naming clarity: 'extract_info' is somewhat vague. Does it fetch, parse, or transform? A more specific name like 'get_paper_info' or 'retrieve_paper_details' would be clearer. Current name could be confused with a general-purpose extraction tool.