A FastMCP-based research server providing arXiv paper search, topic-based paper management, and real-time SSE/REST API endpoints. Implements MCP tools to search, organize, and retrieve academic papers by topic, expose research data via HTTP/SSE, and support integration with agents and web clients.
This server has 2 tools with basic implementations but significant definition quality gaps. Both tools have descriptions and input schemas visible in the source code, but the descriptions lack depth and actionability. The schema definitions are present but minimal. Tool naming is reasonable (verb_noun pattern), but parameter descriptions are sparse. Error handling is minimal, extract_info returns generic error strings rather than actionable guidance. No output schema is documented, and there is no guidance on error recovery. The server implements prompts and resources in addition to tools, which adds some functional breadth, but core tool quality issues limit the overall score.
Search for information about a specific paper across all topic directories. Args: paper_id: The ID of the paper to look for Returns: JSON string with paper information if found, error message if not found
Search for papers on arXiv based on a topic and store their information. Args: topic: The topic to search for max_results: Maximum number of results to retrieve (default: 10) Returns: List of paper IDs found in the search
Minimal parameter descriptions in tool signatures. The 'topic' parameter in search_papers has a description ('The topic to search for'), but lacks format guidance, constraints, or dependency hints. 'max_results' defaults to 10 but provides no bounds (e.g., '1-1000') or rationale. 'paper_id' in extract_info is similarly sparse.
No documented output schema. search_papers returns List[str] (paper IDs), but the docstring does not explain what each ID format looks like or what downstream tools (like extract_info) expect. extract_info returns a JSON string, but no schema is documented for the JSON structure (title, authors, summary fields are visible in code but not in tool spec). LLMs cannot reliably chain these tools without explicit output documentation.
Poor error messages and no recovery guidance. extract_info returns a plain string 'There's no saved information related to paper {paper_id}.' when a paper is not found. This tells the agent nothing about what to do next: should it retry search_papers? Check available folders first? The error lacks actionable remediation steps.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Tool descriptions are below the 50 - 200 character baseline for LLM-optimized descriptions. search_papers description is ~190 chars (acceptable), but extract_info description is ~160 chars and lacks clarity on when to use it vs. search_papers. The descriptions do not explain WHEN or WHY to call each tool or what makes them distinct.
No input validation or constraint declarations. The max_results parameter has no min/max bounds in the schema (only a default of 10). An LLM could pass max_results=999999 without validation, leading to timeout or resource exhaustion. Numeric parameters should declare ranges in both the schema and description.
Composition risk: search_papers and extract_info are tightly coupled via paper_id strings, but the coupling is implicit. search_papers saves papers to the filesystem and returns IDs; extract_info reads from the same filesystem. This works but is fragile, the schema does not document this coupling or the persistence side effect.
Unclear side effects and idempotency. search_papers modifies filesystem state (creates directories, writes papers_info.json). The description does not warn the agent about this side effect. Is the tool idempotent? If called twice with the same topic, does it overwrite or merge? No guidance.