An autonomous research agent that plans, uses tools, stores memory via RAG, and produces structured, citation-backed research reports. It integrates web search, document fetching, and RAG-based retrieval to conduct research queries.
This MCP server has significant gaps in definition quality. All 6 tools are registered via FastMCP decorators with names and basic descriptions, but critical issues emerge: (1) Descriptions are present but minimal (10-60 chars, well below the 194-char baseline for A+ tools), providing insufficient LLM-optimizable context; (2) Parameter schemas lack depth, most parameters have only type and description, missing validation constraints like min/max for k (query_rag), format specifications, or enums; (3) No output schemas are documented anywhere; (4) No error handling guidance is provided, tools return bare success/failure strings without actionable recovery steps; (5) The clear_rag tool is destructive but has no confirmation mechanism or dry-run option; (6) Tool composition issues: fetch_page_content and fetch_pdf_content both do fetching AND indexing, violating single-responsibility; query_rag returns JSON strings that the LLM must parse rather than structured objects. Per-tool analysis shows web_search at 48, fetch_page_content at 35, fetch_pdf_content at 35, query_rag at 50, index_text at 45, clear_rag at 28 (due to destructive risk without safeguards).
Clear all documents and index from the RAG store.
Fetch text content from a URL (HTML, Blog, etc.) and index it.
Fetch text content from a PDF URL and index it.
Directly index a piece of text into the RAG store. 'metadata' should be a JSON string.
Query the internal RAG store for relevant documents. Returns results as a JSON string.
Search the web for relevant sources. Returns a list of URLs as a JSON string.
Destructive operation (clear_rag) has no confirmation mechanism, dry-run option, or permission gate. An agent could wipe the entire RAG store inadvertently.
No output schemas documented for any tool. LLMs cannot predict return structure; tools return bare strings or JSON-string-encoded results, forcing unstructured parsing.
fetch_page_content and fetch_pdf_content combine two responsibilities (fetch AND index) in a single tool, violating single-responsibility principle. This prevents agents from fetching without indexing or composing fetch/index steps independently.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 42 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 32 | - | v1 |
Tool descriptions are minimal (10 - 78 chars, avg 52). Baseline for A+ tools is 194 chars. Descriptions lack WHEN to use, prerequisites, and example outcomes. LLMs cannot optimize tool selection with such sparse context.
No input validation constraints (min/max for k in query_rag, length limits for query strings, format specs for metadata JSON). Parameters lack actionable constraint descriptions, inviting invalid LLM-generated values.
No error handling guidance. Tools return strings like 'Failed to fetch content from {url}' with no actionable recovery steps, categorization (retryable vs fatal), or suggestions for next steps.
query_rag returns JSON-encoded strings instead of structured objects. LLM must parse strings at runtime; no schema aids extraction. This wastes tokens and invites parsing errors.
index_text accepts 'metadata' as a JSON string parameter, which is error-prone and non-idiomatic. LLM must construct valid JSON; parse errors will occur. Consider accepting structured metadata (dict/object) or key-value pairs.