RAG MCP Server with LangChain, Qdrant, and Gemini - provides retrieval-augmented generation capabilities with document ingestion, vector storage, and AI-powered querying
This RAG MCP server has moderate tool definitions but significant quality gaps. Tool names follow the verb_noun pattern well (initialize-rag, add-document, query-rag). However, descriptions are inconsistent in detail and depth. The input schemas provided are basic but present for most tools. Several tools lack parameter descriptions entirely, and output schemas are not documented. The server relies on environment variables for secrets (good practice), but error handling is minimal and recovery guidance is absent. Of the 7 tools evaluated, most have descriptions in the 50-150 character range (reasonable), but schema detail is sparse for several tools. Security practices are sound (no API keys in parameters), but the overall polish and LLM-optimization of definitions is below production baseline.
Add a document from text content to the RAG system, splitting it into chunks and storing in the vector database
Get the current status of the RAG system, including initialization state, document count, LLM model, embedding model, and vector store information
Initialize the RAG system with environment variables from .env file, including Qdrant vector store, Google Generative AI embeddings, LLM, and retrieval chain
List all documents currently stored in the RAG system with their metadata (name, type, chunk count, upload timestamp)
Query the RAG system with a question and receive AI-generated answers based on retrieved documents
Remove a document from the RAG system by its document ID, deleting associated vector embeddings
Output schemas not documented for any tool. LLMs cannot predict what fields will be returned, forcing guesswork about response structure and downstream chaining.
list-documents and get-rag-status have empty input schemas ({}), but lack parameter descriptions in the input specification. Even empty inputs should explain what the tool returns.
initialize-rag has empty input schema but is described as 'Initialize the RAG system with environment variables from .env file'. The description implies it reads .env, but there are no parameters to control this behavior. Clarify whether environment variables are mandatory or configurable.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Upload a document file from the filesystem to the RAG system, supporting PDF, DOCX, and TXT formats
Error handling is minimal. The code catches exceptions but returns no structured error responses to guide LLM recovery. For example, if initialize-rag fails, there is no guidance on what to do next (e.g., 'Check GOOGLE_API_KEY in .env').
No pagination support for list-documents. The description states it returns 'all documents', but with no limit or offset parameters. If the vector store grows, returning thousands of documents could exhaust context windows.
query-rag returns include_sources as a boolean parameter, but the response structure is not documented. The LLM cannot know whether sources appear as an array, nested object, or markdown text.
remove-document lacks idempotent/destructive annotations. The Risk is marked DESTRUCTIVE in metadata, but the tool registration does not include destructiveHint or idempotentHint annotations visible in the code.
add-document's 'type' parameter uses an enum (txt, pdf, docx, doc), but the description does not explain what happens if the content does not match the declared type (e.g., passing PDF binary as 'txt').