MCP server for searching and extracting information from arXiv papers
Two tools are explicitly defined with basic descriptions and parameter schemas. However, critical gaps in output documentation, parameter descriptions, and error handling significantly impact quality. Descriptions are present but minimal (79 and 85 chars respectively, below the 194-char baseline). Tool names follow verb_noun convention but lack clarity around write-side effects for search_papers. Parameters have type definitions but are missing detailed descriptions and constraints. No documented output schemas, no error handling guidance, and no input validation visible.
Search for information about a specific paper across all topic directories.
Search for papers on arXiv based on a topic and store their information.
search_papers has side effects (WRITE) not reflected in description or naming. Description says 'store their information' but doesn't clarify that this writes to local disk/filesystem, creates directories, or has idempotency implications.
No output schema documented for either tool. search_papers returns List[str] (paper IDs) but the description doesn't specify format, structure, or what the LLM should do with them. extract_info returns a JSON string but type annotation suggests str, no schema clarity.
Parameter descriptions are absent or minimal. 'topic' for search_papers has a 28-char description; 'max_results' has 56 chars. Neither explains valid ranges, format constraints, or failure modes. 'paper_id' for extract_info (34 chars) doesn't explain format, source, or what happens if not found.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
No input validation visible. max_results has no bounds (could accept negative, zero, or 10000). topic is free-form, no guidance on length, character constraints, or invalid inputs. extract_info's paper_id accepts any string with no validation.
Error handling is missing. extract_info prints to stderr on JSON decode errors but returns a generic string message ('There's no saved information related to paper...'). No actionable recovery guidance. search_papers has no error cases documented, what if arXiv API fails? What if filesystem write fails?
Tool descriptions are 79 and 85 characters, well below the 194-char baseline. search_papers says 'Search for papers on arXiv based on a topic and store their information' but doesn't explain when to use it vs other tools, what information is stored, or what format is returned. extract_info's description is vague about where it searches or what 'specific paper across all topic directories' means.
No pagination support. search_papers accepts max_results but doesn't return metadata (total found, has_more, next_cursor). If an arXiv search finds 10,000 papers and max_results=5, the LLM has no way to discover or fetch more results.
Tool composition is weak. search_papers returns paper IDs; extract_info expects paper_id. But the flow assumes IDs from search_papers match keys in papers_info.json files. If paper_id format changes or lookup fails, there's no recovery path. No chaining IDs or cross-reference hints.