RAG MCP playground - a Model Context Protocol server that wraps RAG (Retrieval-Augmented Generation) operations for document ingestion and semantic search using Milvus vector database
Server has 2 tools with documented descriptions and basic input schemas. Naming is clear and action-oriented (verb_noun pattern). However, parameter descriptions lack specificity about constraints, ranges, and error handling. Output schema is not formally documented in the tool definitions. Error handling is present in code but not surfaced in tool descriptions to guide LLM recovery. Tool descriptions are well-written (194-240 chars) but lack explicit guidance on when to use each tool vs. the other, and missing documentation of the returned response structure.
Add a document to the knowledge base from raw text content with custom metadata. This tool allows direct ingestion of text content without requiring a physical file. Perfect for dynamic content, API responses, user input, or programmatically generated text that needs to be made searchable in the knowledge base. Features: - Direct text ingestion without file system dependency - Custom metadata support for enhanced categorization and filtering - Same intelligent chunking as file-based ingestion (800 chars with 200 char overlap) - Automatic content length validation and processing Metadata benefits: - Add source attribution and document provenance - Include categorization tags (topic, type, priority) - Store creation timestamps and author information - Enable filtered searches and document organization
Add a document to the knowledge base from a file with intelligent processing. This tool ingests documents from files, automatically processing and chunking them for optimal semantic search. The system supports 15+ file formats including text, markdown, code files, JSON, YAML, and web formats. Features: - Automatic file type detection and processing - Intelligent chunking (800 chars with 200 char overlap) - Metadata extraction (filename, file type, size) - Error handling for missing or corrupted files Supported formats: .txt, .md, .py, .js, .ts, .json, .yaml, .css, .html, .xml, .toml, .ini, .log and more.
Output schema not documented. Code returns a dict with keys like 'success', 'file_path', 'filename', 'file_size', 'processing_time_seconds', but tool docstrings do not explicitly describe the response structure. LLMs cannot plan downstream operations without knowing what fields to extract.
Parameter 'metadata' in add_document_from_content is typed as 'object' with no JSON Schema definition. LLMs cannot infer the structure of nested keys. The description lists example keys ('source', 'topic', 'type', 'author', 'created_at') but does not enforce them as enum, required properties, or type constraints.
No minimum/maximum length constraints for 'file_path' or 'content' parameters. Parameter 'content' has a vague 'minimum recommended: 50+ characters' in the description but no formal minLength schema constraint. 'file_path' has no length limit stated.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 13 | - | v1 |
Error handling in code (file not found, connection failures) is robust, but error guidance is not surfaced in tool descriptions. LLM cannot plan recovery without knowing what errors are possible, what they mean, or what to do next. E.g., 'If file not found, check that the path is absolute or relative to the working directory.'
Tools lack idempotency guidance. If an LLM retries add_document_from_file with the same file_path, is the document duplicated or updated? Code does not show deduplication logic. Without idempotency guarantees, agents may create duplicate chunks.
No distinction or guidance on when to use add_document_from_file vs add_document_from_content. Both tools chunk documents identically (800 chars, 200 char overlap). The description explains the difference (file vs. text) but does not clarify use cases or performance trade-offs, forcing LLMs to reason about which to choose.