A Retrieval-Augmented Generation (RAG) server that indexes documents from multiple sources (PDFs, webpages, Git repositories, folders) and provides semantic search capabilities via vector embeddings
RAG MCP Server has 4 tools with foundational schemas and descriptions, but multiple critical gaps prevent higher scoring. All tools have Pydantic BaseModel definitions with typed parameters and descriptions, which is good. However, output schemas are not documented in the tool registration layer, descriptions lack depth about when/why to use each tool, error handling is minimal (no recovery guidance, error classification, or actionable messages), and security concerns around file system operations are not addressed. The server operates as a knowledge base manager with write/read/destructive operations but lacks comprehensive error handling and output documentation that production systems require.
Add a new source to the RAG knowledge base. Supports PDF files, webpages, Git repositories, and folders. The source will be processed, chunked, embedded, and stored in the vector database for semantic search.
List all indexed sources in the vector database with metadata including source type, path/URL, number of chunks, and last indexed timestamp.
Query the vector database for relevant document chunks using semantic search. Returns the most relevant chunks based on the query embedding similarity.
Remove a source and all its associated chunks from the vector database.
Output schemas not documented to LLM. All tools define Pydantic response models (RetrievedChunk, QueryContextResponse, SourceInfo) in Python but these are never exposed as tool output documentation. LLMs cannot plan downstream calls without knowing what fields are returned.
Destructive tool (remove_source) lacks confirmation/dry-run pattern. No mention of irreversibility, no recovery guidance, no pre-execution confirmation. Agents cannot safely invoke deletion without explicit safeguards.
No error handling documented. Code raises ValueError, RuntimeError without actionable recovery messages. Errors do not classify as retryable/user-fixable/fatal. Examples: 'Could not process PDF: {e}', the LLM cannot determine if retrying helps, or if the PDF is corrupted.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
Missing parameter enums. source_type in add_source is a free-form string despite having fixed valid values ('pdf', 'webpage', 'git_repo', 'folder'). LLMs will hallucinate invalid types like 'arxiv' or 'gdrive'.
File system operations lack input validation and path traversal protection. VECTOR_DB_PATH is environment-configurable but paths in add_source (folder, file_path, git_repo) are unsanitized. An agent could be tricked into indexing arbitrary system files or executing arbitrary git commands.
Pagination missing on list_sources. No mention of how many sources are returned, no limit/offset parameters, no total_count or next_cursor. If an agent indexes 10,000 sources, list_sources returns all 10,000 entries, exploding context window.
Tool descriptions lack 'when to use' context. An LLM has four tools and limited guidance on composition. E.g., should it always call list_sources before query_context? If add_source is in progress, can query_context be called? No composition guidance.