MCP server for searching and fetching academic papers from arXiv with PDF text extraction and graph query capabilities
This MCP server has critical definition gaps across all three tools. While tool names follow verb_noun convention (search_papers, fetch_paper, execute_cypher), the parameter schemas are incomplete, descriptions lack crucial context, and parameter documentation is minimal. The execute_cypher tool is particularly concerning, it accepts arbitrary Cypher queries with no input validation hints or safety guidance. No output schemas are documented. Error handling is claimed but not visible in the provided code. The server exposes arXiv integration and graph database access, both useful for research agents, but the definitions do not meet production standards for reliable agent reasoning.
Execute a Cypher query against a graph dataset
Fetch detailed information about a specific arXiv paper by ID, including PDF text extraction
Search for papers on arXiv by query with optional filtering by date range and category
execute_cypher accepts arbitrary Cypher query strings with no input validation, format constraints, or safety guidance. LLMs may generate syntactically invalid or unsafe queries. No mention of injection risks or query limits.
No output schemas documented for any tool. Agents cannot predict response structure, field names, or available chaining IDs. For search_papers, the response likely includes metadata and IDs needed by fetch_paper, but this is not formally declared.
Parameter descriptions are generic and lack detail. For example, search_papers 'query' parameter has no guidance on syntax (keyword, Boolean, phrase?). 'limit' has no stated range or default. 'category' example 'cs.AI, physics.cond-mat, math.NA' could be mistaken as sample input by LLMs rather than valid enum values, should be a formal enum constraint.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | 2024-11-05+ | v1 |
fetch_paper description states 'including PDF text extraction' but does not clarify what format the extracted text is returned in, or size/length limits. Agents cannot reason about downstream processing.
search_papers and fetch_paper lack pagination/limit documentation. search_papers has a 'limit' parameter but no stated default, maximum, or explanation of what happens if limit exceeds API capability. No 'total' count or 'next_cursor' mentioned in descriptions.
Date parameters ('after', 'before') in search_papers are described as 'YYYY-MM-DD format' but this is a free-form string with no regex pattern, type enforcement, or error guidance. What happens if an agent passes an invalid date? No recovery hint provided.
No error handling guidance visible. If arXiv API times out, returns 404 (paper not found), or rate-limits the client, what does the tool return? Are errors retryable? Should the agent wait or ask the user? No recovery paths documented.