arXiv and DOI research for agents: search, full text, RAG, epistemic profiling
arxiv-mcp demonstrates moderate definition quality with significant strengths in naming and structure but notable gaps in parameter descriptions and output schema documentation. 33 tools show consistent verb-noun naming (search_papers, fetch_full_text, list_categories). Tool descriptions are present but vary in completeness and LLM-optimization, many are overly technical or include implementation details (e.g., 'HTML→Markdown with PDF fallback', 'Jina Reader full text') rather than user-intent framing. Parameter schemas are visible but lack comprehensive descriptions for many parameters (e.g., 'preferred_html' is described minimally, 'format' enums lack rationale). Output schemas are not documented, callers must infer return types from tool names alone. Error handling is implicit; no recovery guidance visible in descriptions. Security consideration: no apparent secret injection for API keys (Semantic Scholar, Jina Reader credentials likely needed). Composition is strong, tools chain logically (search→get_details→fetch_full_text→ingest). The epistemic analysis and depot RAG features are sophisticated but lack clear invocation guidance for users unfamiliar with the domain.
Classify evidence mode and what still needs bench/telescope/human review.
Research plan via MCP sampling (ctx.sample); names concrete tools.
Suggested queries/categories via MCP sampling.
Bundle abstracts for cross-paper synthesis.
Claim-level epistemic profile: rule tags + LLM claim table (MCP sampling or HTTP LLM).
LanceDB vector index status and indexed chunk count.
Job-based deep epistemic analysis for short-timeout clients: submit returns job_id immediately, analysis runs in background (HTTP LLM only), poll with status. Requires ARXIV_MCP_SAMPLING_BASE_URL.
Output schemas not documented. Tool definitions lack return type specifications, forcing callers to infer results from tool names and behavior. This violates pattern:tool and pattern:response-shaper, wasting agent reasoning cycles on type guessing and increasing hallucination risk.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
| 2026-04-07 | F | 23 | - | v1 |
Fetch an Anthropic blog post (backward-compat alias for fetch_lab_post).
FETCH_FULL_TEXT - arXiv HTML→Markdown with PDF fallback when HTML is missing. Tries experimental HTML first (bounded time/size). On 404, timeout, or oversize HTML, falls back to extracting plain text from the arXiv PDF when ``pdf_url`` is available. Rate limits on metadata are retried automatically; transient failures return structured recovery hints instead of hanging.
Fetch and parse a post from Anthropic, Google Research, DeepMind, or Google AI Blog.
Semantic Scholar citations and references.
Jina Reader full text; success + content/abs_url/jina_url or structured error.
Abs-page HTML metadata; success + paper dict or structured error.
Category recent list HTML; includes parse_stats; structured errors.
GET_PAPER_DETAILS - Full metadata: title, abstract, authors, links.
HTML-first ingest plus epistemic profile (knowing type + human/physical loop).
Persist Markdown + FTS + LanceDB chunks in local depot.
Static category catalog; success + categories array.
List Anthropic blog posts (backward-compat alias for list_lab_posts).
Recent papers in an arXiv category (rolling window).
Filter ingested papers by primary_mode and aggregate needs (bench, telescope, formal, deep claims).
List posts from any supported AI lab blog index.
Rebuild LanceDB embeddings for all ingested papers.
Scan recent category submissions, write firefront digest JSON, optional depot ingest.
arxiv.org HTML search; JSON with success, papers, parse_stats, or structured error.
arxiv.org advanced HTML search (field filters, dates).
Search ingested depot: fts, semantic (LanceDB), or hybrid RRF.
SEARCH_PAPERS - Query arXiv with optional category filters and sorting. PORTMANTEAU RATIONALE: Primary discovery surface for "firefront" scanning.
Prefab card: Semantic Scholar citations and references for a paper.
Prefab status card: LanceDB health, indexed chunks, embedding model.
Prefab stats card: depot papers, FTS chunks, favorites, RAG summary.
Prefab claims card: epistemic profile with evidence-mode flags per claim.
Render a rich Prefab card (title, authors, badges, abstract, links) in Claude Desktop.
Parameter descriptions inconsistently detailed. Tools like 'search' and 'searchAdvanced' expose raw HTML scraping parameters ('page', 'page_size') with minimal guidance on limits or defaults. 'sort_by' enum values lack rationale. Missing descriptions on some optional parameters invite hallucinated values.
No visible error recovery guidance. Descriptions do not indicate what to do if a tool fails (e.g., 'If HTML fetch times out, try prefer_html=false', 'If job submission fails, check ARXIV_MCP_SAMPLING_BASE_URL'). Agents have no path forward on errors.
Naming ambiguities reduce clarity. 'search' and 'search_papers' are similar; 'getPaper' vs 'get_paper_details' is inconsistent casing. 'getContent' is vague (content of what?). LLMs may conflate similar names when deciding which tool to call.
Descriptions embed implementation details and technical jargon unsuitable for LLM decision-making. E.g., 'fetch_full_text': 'arXiv HTML→Markdown with PDF fallback', 'jina_url', 'bounded time/size'. Should focus on user intent: 'Get the full paper text in readable format' with fallback strategies hidden inside.
Specialized tools lack context-setting descriptions. 'analyze_paper_epistemics' and 'deep_analyze_paper_epistemics' assume the LLM knows what 'epistemic profile', 'knowing type', and 'human/physical loop' mean. Users unfamiliar with the domain will not understand when to call these tools.
Backward-compatibility aliases ('fetch_anthropic_post', 'list_anthropic_posts') duplicate functionality already covered by 'fetch_lab_post' and 'list_lab_posts'. LLM sees multiple ways to do the same thing, wasting reasoning cycles. Violates pattern:tool (one tool per concern).
Numeric parameter constraints not stated. 'limit' params lack max values (capped at 100? 1000?). 'hours' in 'list_category_latest' has no min/max guidance. Unbounded numbers invite LLM to pass absurd values.