A RAG knowledge base server that indexes documents and answers questions using Corrective RAG pipeline with ChromaDB vector store and local LLM backend
mcp-rag provides 5 tools with complete parameter schemas and reasonable descriptions. Naming follows verb_noun conventions (index_folder, ask_question, find_relevant_docs, summarize_document, index_status). All tools have parameter type definitions and descriptions. However, descriptions lack depth about error cases, prerequisites, and recovery paths. Output schemas are documented in text but not formalized. No tool annotations (readOnlyHint/destructiveHint) despite clear semantic distinctions (index_folder is destructive/WRITE, others are READ_ONLY). Error handling guidance is minimal, no mention of what to do when LLM connection fails or when retrieval returns empty results. The mcp.instructions are helpful for agent guidance but do not substitute for per-tool documentation.
Ask a question against the indexed document collection. Runs the full Corrective RAG pipeline: rewrite → retrieve → grade → generate → hallucination_check Args: question: Natural language question in any language. Returns: Dict with keys: - answer (str): The generated answer with source citations. - sources (list[str]): Source file paths referenced in the answer. - is_grounded (bool): Whether the answer passed hallucination check. - retrieve_retries (int): Number of retrieval retry loops performed. - generate_retries (int): Number of generation retries performed.
Retrieve the most relevant document chunks for a query without generating an answer. Useful for inspecting what the index contains or for debugging retrieval quality. Args: query: Natural language search query. top_k: Maximum number of chunks to return (default: 5). Returns: Dict with key 'results': list of chunks, each with text, source, chunk_index, and distance (lower = more similar).
Index all supported documents in a folder into the vector store. Supported formats: .md, .txt, .rst, .py, .js, .ts, .json, .yaml Args: folder_path: Absolute or relative path to the folder to index. glob_pattern: Glob pattern to filter files (default: all files recursively). Returns: Summary dict with files_indexed, chunks_added, skipped_files, errors.
Return current vector index statistics.
index_status description is too brief (26 chars) and lacks context about what fields are returned and when to call it.
Output schemas are documented in natural language (e.g., 'Dict with files_indexed, chunks_added, skipped_files, errors') but not formalized in JSON Schema format visible to MCP clients. This forces LLMs to infer structure.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite semantic clarity: index_folder is destructive (WRITE), others are read-only. Missing annotations prevent clients from offering context-specific affordances (confirm before index, safe to retry reads).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 49 | - | v1 |
Summarise a single document file using the local LLM. The file is read and fed directly to the summarisation prompt — it does not need to be indexed first. Args: file_path: Absolute or relative path to the file to summarise. Returns: Dict with keys: summary (str), filename (str).
Error handling lacks recovery guidance. If ask_question fails due to LLM connection error, the description does not say 'call llm_status to diagnose', only the mcp.instructions mention this. Per-tool descriptions should be self-contained.
find_relevant_docs description mentions 'distance (lower = more similar)' but does not state the distance metric (cosine, euclidean, etc.) or typical range. LLMs cannot reason about distance values without context.
No guidance on what happens if index_folder is called twice on the same directory. Is it idempotent? Does it deduplicate chunks or add duplicates? This matters for agent planning and retry safety.
summarize_document description says 'does not need to be indexed first' but does not explain why you would use this vs ask_question for single files, or what the summary is optimized for. No context on summary length or format.