MCP server for searching and extracting information from arXiv papers using the FastMCP framework with SSE transport
The server exposes 3 tools via FastMCP with basic descriptions and schemas. However, significant gaps exist: parameter descriptions are minimal or missing, output schemas are not explicitly documented, error handling lacks recovery guidance, and tool naming could be clearer. The tools operate on arXiv research papers with local file storage, but the interface does not follow production-grade LLM agent patterns. Naming is mostly action-oriented (search_papers, extract_info), but descriptions are sparse and output structures are underdocumented.
Retrieve stored metadata for a given paper ID
Search arXiv for papers on a given topic, store metadata locally
Returns the server status including IP address, port, transport type, and initialization status
Output schemas not documented. None of the 3 tools formally declare their return types in the tool schema. Agents cannot plan downstream calls or extract fields reliably. search_papers returns List[str]; extract_info returns Dict with title/authors/summary/pdf_url/published; status returns Dict with ip/port/transport/status. These must be documented.
Parameter descriptions are incomplete or missing constraint information. search_papers max_results lacks bounds (code defaults to 5, but min/max are not specified). extract_info paper_id description does not state the required format '1234.56789' or 'v2' suffix rules, even though the code enforces a regex. Agents cannot validate input without explicit constraint documentation.
Generic tool name 'status' lacks action verb. Should be 'get_status' to clarify intent. Naming convention (verb_noun) helps LLMs disambiguate tools, 'status' alone is ambiguous.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Error handling lacks recovery guidance. extract_info returns {error: 'Paper not found'} for invalid IDs, but does not suggest alternatives (e.g., 'Did you mean to search for similar papers?'). Agents receive no direction on next steps.
Tool descriptions do not clarify relationships. search_papers stores metadata locally; extract_info retrieves it. But the descriptions do not explain that extract_info depends on prior search_papers calls, or how the tools interact. Agents may call extract_info without understanding the prerequisite.
No idempotency documentation. search_papers overwrites local JSON on each call; unclear if repeated calls with same topic are safe or create duplicates. extract_info is read-only and idempotent, but this is not declared via toolAnnotations.