An MCP server for searching and extracting information from academic papers on arXiv
This server has significant structural and quality issues. While the two tools are registered with basic schemas in arx.py, the tool definitions lack depth in descriptions, parameter documentation is minimal, and output schemas are not documented. The code shows incomplete implementation (truncated process_query function, typo in append). Tool naming follows verb_noun convention (search_paper, extract_paper_info) but descriptions are too brief (10 and 14 characters respectively, well below the 34-392 character baseline for A+ tools). The extract_paper_info function contains a critical bug (os.dir instead of os.listdir) that would cause runtime failures. Parameter descriptions exist but are too generic. No error handling guidance, no validation of inputs, and no documentation of output structure. The server is functional but not production-ready.
Extract information about a paper
Search for a paper
Tool descriptions are critically short (10-14 characters). The rubric baseline for A+ tools is 34-392 characters. 'Search for a paper' and 'Extract information about a paper' provide minimal context for LLM tool selection. Descriptions must explain WHAT the tool does, WHEN to use it, and any prerequisites.
Parameter descriptions are too generic and lack actionable constraints. 'The query to search for' does not explain valid formats, length limits, or examples. 'The ID of the paper' does not specify what format paper_id accepts (ArXiv ID format? Full URL? Filename?). Descriptions must state expected format and constraints explicitly.
Output schemas are not documented. The rubric requires 100% of A+ tools to have documented return types. search_paper returns a list of paper objects with fields like title, author, summary, url, etc., but the schema in arx.py does not declare this structure. extract_paper_info returns a JSON string or 'Paper not found', ambiguous. LLMs cannot plan downstream calls without knowing response structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 26 | - | v1 |
No pagination support. search_paper caps at max_results=20 but returns all results in one response without limit enforcement in the tool itself. If an agent calls search_paper for a very broad query, the response could contain dozens of papers, bloating the context window. Should implement page/limit parameters and return a next_cursor or total_count.
Error handling provides no recovery guidance. extract_paper_info returns 'Paper not found' (ambiguous string) and None (silent failure) rather than structured errors that tell the LLM what to do next. Should return 'Paper ID "xyz" not found. Try calling search_paper() first to find available papers.'
Critical bug in extract_paper_info: uses os.dir() which does not exist (should be os.listdir()). This function will fail at runtime with AttributeError. Also logic error: checking if paper_id in papers_info (list) will never succeed, should iterate and compare against a paper ID field.
Source code is incomplete and contains typos. arx.py truncates mid-function (process_query definition incomplete, missing function body). serverconnect.py has logical issue in tool call loop (breaks after first tool, never processes subsequent tools or completes conversation loop). Code quality prevents confident evaluation of actual runtime behavior.
No input validation. max_results parameter for search_paper has no min/max bounds declared in schema. Paper_id parameter has no format constraints. Unconstrained inputs allow LLMs to pass invalid values (negative max_results, empty paper_id, etc.). Should declare enum, minLength, maxLength, pattern, or min/max in schema.
Parameter response field naming mismatch. search_paper returns results with fields like 'author' (a list), 'published', 'url', 'categories'. If downstream tools or the agent need to reference a paper, they must use 'url' as the ID, but extract_paper_info expects 'paper_id' format undefined. Mismatched naming forces LLM reasoning about field mappings.
Tool composition unclear. After calling search_paper, what should the LLM pass to extract_paper_info? The url field? A generated ID? The description does not specify. Tool A's output must contain the IDs and references Tool B needs. Design should clarify whether paper_id is the ArXiv ID (e.g., '2301.00001') or the full URL.