A MCP server layer for existing APIs from popular sources e.g. arXiv, DBLP, etc. to help researchers expedite literature review process using LLMs and MCP Clients like Claude, Cursor, etc.
lit-mcp provides two search tools (arxiv_search, dblp_search) with minimal but present definitions. Tool names follow verb_noun pattern (search_*). Descriptions exist but are generic (10-47 chars, below the 34-char p10 baseline, well below the 194-char average). Parameters have type declarations and brief descriptions, but lack detail on constraints, formats, or valid ranges. Output schemas are partially documented in docstrings but not formally declared in the tool definitions. No error handling guidance, recovery paths, or actionable error messages. Prompts are defined (3 total: latest_info, related_topics, author_spotlight) but lack input validation and error handling. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite all tools being read-only operations. No parameter constraints (enums, ranges, patterns) to guide LLM input validation. Missing: output structure documentation, pagination support (both tools default to max_results=10 but no offset/cursor mechanism), dependency hints, and examples of error recovery.
Search for papers on arXiv.
Search DBLP database for computer science papers.
Tool descriptions are below baseline and lack actionable context. 'Search for papers on arXiv' (28 chars) and 'Search DBLP database for computer science papers' (47 chars) do not answer: WHEN to use this tool instead of the other? What exact fields are returned? What formats are supported? The descriptions are shorter than the 34-char p10 baseline and far below the 194-char average for production tools.
Input parameters lack format constraints and valid value ranges. The 'query' parameter has no length limits, regex patterns, or guidance on what constitutes a valid search string (full author name, keywords, phrases, Boolean operators). The 'max_results' parameter lacks min/max bounds (is 0 valid? Is 10000 valid?). Unbounded integers invite absurd values that waste API quota or cause timeouts.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Output schemas are documented only in docstrings (not formally declared in tool definitions). The docstrings state 'Returns: A list of dictionaries with the following keys: [...]' but the MCP tool registration does not include a formal 'outputSchema' declaration. LLMs cannot reliably parse docstrings to infer output structure; formal schema declarations are required for robust tool chaining.
No error handling or recovery guidance. If a search returns zero results, times out, or fails due to invalid input, there is no documented error response, categorization (retryable vs user-fixable), or suggested next action. LLMs will have no context for recovery.
No tool annotations despite read-only semantics. The readOnlyHint annotation is missing, which signals to agents that these tools are safe to call multiple times without side effects. This is especially important for search tools that agents may invoke speculatively.
No pagination mechanism or result limits enforced in tool design. Both tools accept max_results up to 10 by default with no guidance on upper bounds or offset/cursor support. Large result sets will blow the context window. The baseline recommends capping results at 20-50 and enforcing pagination.
Prompt definitions lack input parameter validation and error handling. The three prompts (latest_info, related_topics, author_spotlight) accept a 'topic' string with no constraints, length limits, or validation. They return strings with no documented error cases or handling.