MCP server for searching and reading academic papers from arXiv and OpenAlex, with PDF to Markdown conversion and local caching
The Paper Reader MCP server defines 2 tools with generally complete schemas and reasonable descriptions. Tool names follow verb-noun conventions (search_papers, get_paper_content). Both tools have documented input parameters with types and descriptions. However, output schemas are not explicitly documented in code (only mentioned in docstrings), parameter descriptions contain example values (an anti-pattern), and there is no formal error handling guidance or structured output schema definition. The tools are read-only with clear, helpful descriptions that explain WHEN to use them and what they return, but lack recovery guidance for error cases. Overall quality is above average for community servers but falls short of production-grade documentation standards.
Get full paper text in Markdown format with pagination support. Downloads paper PDF and converts to Markdown. Papers are cached locally.
Search academic papers by keywords, returning title, abstract and other information. Supports sorting and category filtering.
Parameter descriptions contain example values (e.g., 'e.g. machine learning', 'e.g. 2301.12345'). LLMs tend to reuse example values literally rather than adapting to actual context. Use formal enum constraints and format declarations instead.
Output schemas are documented only in docstrings, not as formal JSON Schema. The tools return free-form string responses rather than structured JSON objects. LLMs cannot reliably parse unstructured text to extract data for downstream tool calls or reasoning.
Error responses are generated as unstructured text strings (e.g., '❌ 搜索失败: {str(e)}'). These lack recovery guidance. Error messages should tell the LLM what to do next (e.g., 'Try with different keywords', 'Check paper_id format').
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
No explicit documentation of return field structure. While docstrings mention fields like 'paper ID', 'title', 'abstract', 'authors', there is no JSON Schema that defines the exact structure, types, and optionality of returned fields.
The 'sort_by' parameter accepts 'smart', 'relevance', 'submitted', 'updated' but these are described as free-form strings. Should be formally declared as an enum to prevent hallucinated values and self-document valid options.
The 'source' parameter ('arxiv' vs 'openalex') lacks enum declaration. Both tools accept this parameter and treat it differently, but valid values are described inline rather than as formal constraints.
No pagination support documented or returned by search_papers. The tool accepts max_results (up to 50) but returns all matching results without a total_count, next_cursor, or pagination metadata. Large result sets are not trimmed or structured for chunked LLM processing.
The get_paper_content tool accepts 'page' parameter for pagination, but the output is a plain string without metadata indicating current_page, total_pages, or remaining_chars. An LLM cannot reliably determine when to call the tool again for the next page.