The server defines 4 tools with reasonable naming (verb_noun patterns: search_arxiv, get_paper, get_recent, search_papers) and detailed descriptions that explain when to use each tool. However, there are significant gaps in schema documentation, parameter descriptions, and error handling guidance. Tool descriptions are clear and actionable (ranging 150-250 chars), but output schemas are completely undocumented, LLMs cannot reason about downstream calls or field names. Parameter validation exists in code but is not reflected in JSON Schema constraints (enums, patterns, min/max bounds). Error messages in the source code are generic strings without recovery guidance. All tools are read-only but lack explicit idempotent/readOnly annotations.
Fetch a single paper by its arXiv ID (e.g. 2301.07041). Use when the user provides a specific arXiv ID and wants full metadata including title, authors, publication date, abstract, and PDF link. IDs can be with or without version suffix (e.g. 2301.07041v2).
Get the most recent papers in a given category (e.g. cs.AI, quant-ph). Use when the user wants to see the latest submissions in a specific arXiv category, sorted by submission date (newest first). Ideal for 'what is new' style queries.
Search arXiv papers by keyword, author, category, and/or date range. Use when the user wants to find papers matching specific terms, by a specific author, in a specific category, or within a date range. Supports Boolean-like searches via keyword. Returns a numbered list with title, arXiv ID, authors, publication date, and PDF link.
Search papers by query text, with optional category filter and date range. Simpler than search_arxiv - just a search query plus optional category and date_from. Use when the user provides a natural language query like 'papers about transformers' and optionally a category or start date. The query searches across all fields (title, abstract, authors).
Output schemas are not documented. Tools return formatted strings (_fmt() function) but LLMs cannot parse structured fields, plan downstream calls, or extract IDs reliably. No field names, types, or structure documented.
JSON Schema lacks constraints. Parameters like max_results, date_from, date_to, category lack enum, pattern, minLength, maxLength, minimum, or maximum in the schema. Validation happens in Python code but is invisible to LLMs at schema-inspection time.
Error messages lack recovery guidance. Code returns 'Error: <message>' strings without actionable next steps. E.g., 'No results found' does not suggest trying a broader query or different category.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 54 | 2026-07-28+ | v2 |
Missing idempotent/readOnly annotations. All tools are read-only GET operations, but tool definitions carry no readOnlyHint or idempotentHint. This prevents MCP clients and LLMs from understanding safe retry semantics.
Parameter descriptions in schema lack detail. The 'keyword' param in search_arxiv says 'Search keyword to find papers' but does not specify: wildcards allowed? Case-sensitive? Boolean operators supported (AND/OR/NOT)? Required if others omitted? This forces LLMs to guess.
Date format validation is strict (YYYYMMDDHHMM) but the parameter description uses an example '202401010000' instead of formally declaring the pattern as regex.
Pagination not documented in schema. search_arxiv and search_papers accept 'start' and 'max_results' but do not document total count, has_more, or next_cursor in output. LLMs cannot know if results are paginated or complete.
search_arxiv requires at least one of {keyword, author, category, date_from, date_to} but this constraint is not encoded in the schema as mutuallyExclusive/anyOf/oneOf. The runtime check returns an error string, but schema validation tools cannot prevent invalid calls.