A Retrieval-Augmented Generation (RAG) MCP server that provides tools for PDF embedding, document management, and semantic search across ChromaDB vector collections with reranking capabilities.
RAG-MCP-Server has 11 tools with definitions visible in Modules/ToolDefinition.py. While all tools have descriptions and basic schemas, quality is inconsistent. Tool names are generally action-oriented (embed, retrieve, list, delete, search) but lack consistency in verb-noun conventions. Descriptions are present but often generic (e.g., 'Returns a dictionary...' for multiple tools). Input schemas show parameter types and required flags, but lack constraint documentation (enums, ranges, formats). A critical typo exists: 'seacrhAcrossCollections' (tool #8) is misspelled. Parameter descriptions are minimal and do not follow the 'actionable context' pattern, they describe WHAT but not WHEN or WHY. Error handling and output schema details are absent. The server exposes a local document filesystem interface and ChromaDB integration but lacks safeguards for destructive operations (deleteCollection accepts only a collection name with no confirmation step).
Adds citation references to a given LLM-generated answer based on retrieved documents.
Retrieves statistics about a specific ChromaDB collection.
Deletes the specified ChromaDB collection if it exists.
Lists and describes all tools available on the MCP server.
Embeds the given PDF and stores it in Chroma vectorstore.
Retrieves metadata for a specified local document.
Lists all the collections in the ChromaDB vectorstore.
Tool name typo: 'seacrhAcrossCollections' (tool #8) is misspelled; should be 'searchAcrossCollections'. This breaks discovery and LLM-based tool selection.
deleteCollection tool lacks confirmation mechanism. Destructive operations should support dry-run or explicit confirmation to prevent accidental data loss. Currently no safeguard.
embedPDF has conflicting required parameters (filepath vs file_url vs filename). All three are optional, but the tool needs at least one. This is undocumented and will cause agent confusion. Parameter relationships not stated in descriptions.
No output schema documentation. Tools return dicts with unspecified fields (e.g., 'Returns a dictionary...' without listing keys, types, or availability). LLMs cannot plan downstream calls or extract data reliably.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 34 | - | v1 |
Lists all locally stored documents in the DOCUMENT_DIR.
Retrieves relevant documents from ChromaDB based on a query.
Searches for relevant documents across multiple ChromaDB collections based on a query.
Updates the filename of a local document in DOCUMENT_DIR.
Parameter descriptions lack actionable context. E.g., 'User query for retrieval' does not explain query format (plain text? supported operators?), length limits, or language support. Descriptions should follow 'what+when+constraints' pattern.
No input validation or error handling guidance. Tools lack error response definitions and recovery paths. If a collection name is invalid or file does not exist, no guidance on next steps for the agent.
Generic descriptions reused across tools. Multiple tools return 'Returns a dictionary...' without specifying contents. This violates the baseline that 100% of A+ tools have documented return types.
Boolean and integer parameters lack range/constraint documentation. 'use_reranker' (boolean) and 'top_k' (integer) have no guidance on valid ranges (e.g., top_k 1 - 100) or default behavior if omitted.
citationProvider expects 'results' as a dict from retrieveDocs, but no formal contract is documented. If retrieveDocs output schema changes, citationProvider breaks silently. This violates tool-chain pattern.
updateDocument tool accepts arbitrary new filenames with no path validation. Vulnerable to path traversal if attacker-controlled. Accepts 'filename' but no constraints on special characters or directory traversal patterns documented.