RAG Knowledge Base MCP Server - A Model Context Protocol server that provides semantic search capabilities over a RAG (Retrieval-Augmented Generation) knowledge base. Enables LLMs to query, retrieve, and manage documents using vector similarity search with configurable chunking and embedding strategies.
RAG Knowledge MCP provides well-structured tool definitions with clear naming conventions and comprehensive parameter schemas. All 4 tools follow verb_noun naming (rag_search_knowledge, rag_list_documents, rag_get_document, rag_get_stats). Each tool includes a detailed description explaining what it does, when to use it, and what it returns. Input schemas are properly defined with type constraints, min/max bounds, and enums. However, output schemas are not explicitly documented in the source, and error handling guidance is absent. The codebase shows a production-minded architecture with abstract backends and configuration management, but MCP tool responses lack structured error categories and recovery hints.
Retrieve a specific document from the knowledge base. Fetch the full content and metadata of a document by its ID. Returns the complete reconstructed document from its indexed chunks.
Get statistics about the knowledge base. Returns information about the indexed documents, chunks, embedding model, vector dimensions, and storage configuration.
List documents in the knowledge base with pagination. Retrieve a paginated list of all indexed documents with their metadata. Use limit and offset parameters to navigate through results.
Search the knowledge base using semantic similarity. Performs vector similarity search over indexed documents to find the most relevant content matching the query. Results are ranked by similarity score and can be filtered by a minimum threshold.
Output schemas not documented in tool definitions. While tools describe what they return conceptually (e.g., 'ranked by similarity score'), the actual response structure with typed fields, field names, and data types is not formally defined in the MCP registration. This forces LLMs to infer output structure and makes downstream tool chaining fragile.
No error handling guidance in tool descriptions. Tools do not explain what errors can occur (e.g., 'document not found'), how to recover (e.g., 'call rag_list_documents to find available IDs'), or whether failures are retryable. This leaves LLMs without a recovery path when calls fail.
No tool annotations present. Tools are marked as READ_ONLY in the spec but do not include readOnlyHint/destructiveHint/idempotentHint annotations in the MCP registration. This omission means LLMs cannot infer safety properties directly from the tool definition and must reason about safety separately.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Response format parameter (markdown/json) is exposed but output documentation does not clarify field structure for each format. Returning 'markdown' vs 'json' is a presentation choice, but the underlying data structure (what fields, types, nesting) should be stable and documented.
No batch or bulk operation variants. Agents iterating over search results to fetch individual documents must make N sequential rag_get_document calls instead of one batch_get_documents call. This wastes tokens, latency, and invites mid-chain failures.