MCP server for searching and extracting information from arXiv papers
This server has two tools with basic but incomplete definitions. Both tools have short, generic descriptions (under 50 chars) that lack actionable context. Input schemas are present with type declarations, but parameter descriptions are minimal or missing entirely. No output schemas are documented. The naming is reasonably clear (search_papers, extract_info follow verb_noun convention), but the tool definitions lack the depth required for confident LLM selection and error recovery. Error handling is absent, there are no documented recovery paths, validation guidance, or actionable error messages in the source. This is a typical community MCP server in the 40-60 range.
Search for information about a specific paper across all topic directories.
Search for papers on arXiv based on a topic and store their information.
Descriptions are too short and lack actionable context. 'Search for papers on arXiv based on a topic and store their information' (61 chars) does not explain WHEN to use this vs extract_info, what 'store their information' means, or what fields are returned.
No output schemas documented. The source shows search_papers returns List[str] (presumably paper IDs), but the LLM has no visibility into what those strings contain, what format they use, or what downstream tools expect. This breaks tool chaining and forces guessing.
Parameter 'paper_id' in extract_info lacks a description. The LLM cannot infer whether to pass a full arXiv ID (e.g., '2024.01234'), a basename, or a UUID. This ambiguity will cause incorrect calls.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 39 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
No error handling or recovery guidance. The source does not document what happens when a topic yields zero results, when a paper_id is not found, or when file I/O fails. LLMs receive no actionable error messages to self-correct.
Tool definitions appear inferred rather than directly visible in the partial source. The research_server.py snippet is incomplete (cuts off mid-function). If tool registration and schemas are not explicitly shown, per-tool scores should be capped at 50.
Parameter 'topic' in search_papers has a basic description ('The topic to search for') but no guidance on format, constraints, or multi-word handling. Does 'machine learning' work? 'machine-learning'? 'ML'? The LLM must guess.
Default value for max_results (5) is not explained. Why 5? Is this a hard limit or a suggested default? Can users request more? The description should clarify bounds and trade-offs.