An MCP server that searches for academic papers on arXiv and extracts paper information. Provides tools for paper discovery and information retrieval, along with resources and prompt templates for research workflows.
This MCP server has 2 tools with basic schemas and descriptions, but significant quality gaps prevent a higher score. Both tools have descriptions and input schemas present, but descriptions are terse (both under 100 chars), parameter descriptions lack detail, and there is no documented output schema. Error handling is minimal, tools return plain strings without guidance for recovery. The tool names follow verb_noun convention ('search_papers', 'extract_info') but are somewhat generic. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present. The 'search_papers' tool modifies state (writes to disk) but this is not clearly stated in the description. Parameter constraints (e.g., max_results range) are absent from descriptions. Output is unstructured text rather than typed JSON objects, making it difficult for LLMs to extract specific fields for downstream tool chaining. This server lands in the 'Fair' range, functional but not production-ready.
Search for information about a specific paper across all topic directories.
Search for papers on arXiv based on a topic and store their information.
Missing output schema documentation. Neither tool documents what fields the LLM should expect in the response. search_papers returns a list of paper IDs as a comma-separated string; extract_info returns a JSON string. Without formal output schema definitions, LLMs cannot plan downstream calls or extract structured data reliably.
search_papers modifies state (creates directories, writes JSON files) but the description does not explicitly state this. Description should say 'Searches arXiv and stores paper metadata to disk in JSON files organized by topic.' Agents need to know which calls have side effects for safe retry logic.
Parameter descriptions lack actionable constraints. 'max_results: Maximum number of results to retrieve' does not specify bounds. Should state: 'Maximum number of results (1-100, default: 5)'. Unbounded or vague constraints invite LLMs to pass invalid values.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
No error handling guidance. Both tools can fail (network errors, file I/O errors, paper not found) but error responses are plain strings without recovery hints. Example: 'There's no saved information related to paper {paper_id}' does not guide the LLM on next steps. Should suggest 'If paper not found, try calling search_papers() first to index papers for this topic.'
No tool annotations. Neither tool has readOnlyHint (extract_info is read-only) or destructiveHint (search_papers writes to disk). Tool annotations help agents understand side effects and plan safer execution.
Unstructured output format. search_papers returns a comma-separated string of IDs; extract_info returns a raw JSON string. LLMs must parse these strings manually. Should return structured objects: search_papers returns {"papers": [{"id": "...", "title": "...", ...}]}, extract_info returns {"paper_id": "...", "title": "...", ...}.
Tool chaining risk. search_papers returns paper IDs but extract_info expects a paper_id parameter. The response field name matches, but without explicit documentation of the output schema, an LLM may not reliably extract the ID to pass to extract_info. This should be verified by documenting the exact output structure.