Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The server defines 19 tools with basic naming (all verb-based), but suffers from significant quality gaps. Tool descriptions are present but mostly generic (20-50 chars, below the 194-char baseline). Input schemas are visible but lack depth: parameters have type declarations but descriptions are often minimal or absent. Output schemas are not documented. Error handling is basic. The tools cover a coherent academic research domain, but the implementation quality is below production standard. STDIO transport limits protocol readiness significantly.
Output schemas not documented. Tool definitions declare input schemas but nowhere in the visible source (server.py, tools module excerpt) do we see documented return types or response structure. LLMs cannot infer what fields a tool returns, forcing them to guess at downstream data extraction.
Generic and minimal parameter descriptions. Most parameters (e.g., 'query', 'paper_id', 'topic') have descriptions under 50 characters with no guidance on format, valid ranges, or constraints. 'Search query' does not tell an LLM whether to expect free-text keywords, boolean operators, or arxiv-specific syntax. Baseline for A+ tools is 72 chars per parameter.
Recommendations
Document the output schema for every tool. Specify the structure (object fields, array items, types) that each tool returns. Example: 'Returns {papers: [{id, title, authors, abstract, url, published}], total_count: integer, has_more: boolean}'.
Expand parameter descriptions to 60 - 100 characters with format and range guidance. Example: 'sort_by: How to order results. Valid values: relevance (default), submitted_date (newest first), submitted_date_asc (oldest first). Default: relevance.'
Add enums for multi-value parameters. For 'format' in export_citations, define enum=["bibtex", "apa", "mla", "ieee"] instead of free-form string.
Define pagination explicitly. For search_papers, update parameter description: 'max_results: Maximum number of papers to return (integer, 1 - 100, default: 20). Use with semantic_search to fetch next page.'
Enhance tool descriptions with WHEN and WHY context. Example: 'search_papers: Search arXiv for papers by keyword or arxiv ID. Use when you need to find papers on a topic. Returns top 20 matches ranked by relevance. Call list_papers to see already-downloaded papers without network latency.'
Add tool annotations to state destructiveness. Mark unwatch_topic, reindex, and watch_topic with destructiveHint=true in the Tool definition so the LLM knows these have side effects.
Include recovery hints for error cases. Add to tool descriptions: 'If paper not found, verify the arxiv ID (format: YYMM.NNNNN or YYMMNNNvN) or try search_papers(query=title).'
Spec posture evidence
Inferred effective spec: 2026-07-28+.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Tool descriptions are too brief (20 - 50 chars vs. 194-char baseline). Example: 'Download a paper from arXiv' does not explain when to call this tool vs. read_paper, what format is returned, or what happens on failure. LLMs need context to select tools correctly.
No enum constraints on multi-value parameters. 'sort_by' in search_papers accepts a string but no enum of valid values is defined. LLMs will hallucinate invalid sort orders (e.g., 'relevance-reverse', 'newest-first') instead of valid options. 'format' in export_citations also lacks enum definition.
No pagination or limit guidance in schema. search_papers declares 'max_results' as integer but no description of valid range (e.g., 1 - 100, default 20). semantic_search has the same gap. Unbounded integers let LLMs request 10,000 results, overloading the API or context window.
Error handling guidance absent from tool definitions. The call_tool() handler catches exceptions and calls _tool_error_message(), but the tool descriptions do not document what errors are possible, when they are retryable, or what the LLM should do next. A generic 'Error: resource not found' tells an LLM nothing.
No tool composition guidance. 19 tools exist but tool descriptions do not cross-reference related tools or explain preferred workflows. An LLM might call search_papers then separately call get_abstract for each result, wasting tokens, instead of understanding that search_papers could return abstracts if the tool were designed properly.
No destruction/destructive hints or confirmation patterns. Tools like 'unwatch_topic' and 'reindex' modify state but lack tool annotations (destructiveHint=true) or confirmation steps. An LLM might accidentally unwatch all topics without user approval.
unwatch_topicreindexdownload_paper
Document idempotency. Clarify in descriptions which tools are idempotent (safe to retry) and which have side effects. Example: 'read_paper is idempotent (no state change). download_paper is idempotent if the paper is already downloaded.'
Add batch variants for tools called in loops. A 'download_papers_batch' accepting paper_ids: [string] would be more efficient than calling download_paper 10 times.
Cross-reference related tools in descriptions. For search_papers, note: 'After searching, call read_paper(paper_id) to fetch full content, or get_abstract(paper_id) for a summary.'
Sanitize inputs to prevent prompt injection and arxiv ID spoofing. Validate paper_id format (YYMM.NNNNN) before calling arxiv API.
Return human-readable error messages with actionable next steps. Instead of 'Error: 404', return 'Paper 2024.99999 not found on arXiv. Double-check the arxiv ID or search by title using search_papers(query="your title here").'.