A retrieval-augmented generation (RAG) backend with FastAPI and MCP server integration for document processing, chunking, embedding, and semantic search using Qdrant vector database
Optim-RAG demonstrates a functional RAG tool suite with complete parameter schemas and reasonably detailed descriptions. However, there are significant gaps in schema completeness for nested objects, missing type definitions on array items, inconsistent parameter naming conventions, and underdocumented error conditions. Tool descriptions are adequate (10 - 24 sentences) but could be more precise about dependencies and error recovery. The server uses FastMCP to auto-convert a FastAPI app to MCP, which works but does not guarantee high-quality tool metadata. The tools follow a clear semantic structure (session management, document editing, chat), but composition could be tighter: create_session handles file extraction/chunking/embedding in one call, which couples concerns and limits reuse.
Create and process a new session from an uploaded document archive. The uploaded file may be a ZIP containing multiple documents or a single standalone document. Files are extracted, categorized, chunked, and stored in the Qdrant vector database as embeddings. Args: archive: The uploaded document archive (ZIP or single file). session_name: human-readable name for the session. Returns: A `SessionMeta` object describing the created session, including generated session ID, creation timestamp, archive name, and size.
Delete an existing session and all its stored chunks. Removes all associated embeddings and payloads from the vector store (Qdrant) corresponding to the specified session ID. Args: session_id: Unique identifier of the session to delete. Raises: HTTPException(404): If the session does not exist. Returns: A `DeleteSessionResponse` confirming the deletion status.
Retrieve all document chunks for a specific session. This endpoint returns the full list of stored text chunks (and their metadata) belonging to the provided `session_id`. Useful for debugging or inspecting what content was indexed for retrieval.
Retrieve metadata for a specific session. Returns minimal session metadata such as ID and creation time. This endpoint can be extended to include additional metadata (like document count or vector stats) as needed. Args: session_id: Unique identifier of the target session. Raises: HTTPException(404): If the session does not exist. Returns: A `SessionMeta` object describing the session.
Array item schemas lack type definitions. 'documents' in update_chunks and 'files' in upload_files declare type:array but do not properly type items, items are objects with properties, but array cardinality, item constraints, and required fields are missing.
Nested object properties lack descriptions. In update_chunks, the 'documents' array contains properties like 'chunk_id', 'filename', 'filetype', 'page_number', 'page_content', 'status', 'lastEdited', 'originalHash', none of which have descriptions. LLMs cannot infer what each field controls or whether it is required.
send_chat parameter 'messages' array item properties ('role', 'content') lack field-level descriptions. The enum values for 'role' are ['developer', 'user', 'assistant'], unusual (most tools use 'system', 'user', 'assistant'), but undocumented. LLMs will guess at semantics.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 62 | - | v1 |
Retrieve metadata for all stored sessions. Each session corresponds to a distinct uploaded dataset or document archive that has been processed and indexed in the vector store (Qdrant). Returns: A list of `SessionMeta` objects containing ID, name, creation timestamp, and basic archive information for each session found.
Execute a retrieval-augmented chat query for a given session. This endpoint performs a semantic search across the vectorstore (Qdrant) associated with the specified `session_id`, retrieves the most relevant document chunks, and generates an answer using the active LLM pipeline. Args: req: A `SendChatRequest` containing: - `session_id`: The target session to query. - `query`: The user's natural-language question or message. - (Optional) `top_k`: Number of chunks to retrieve. - (Optional) `history`: Prior chat context (if multi-turn chat). Returns: A `SendChatResponse` containing: - `answer`: The LLM-generated response based on retrieved context. - `sources`: Metadata for the retrieved chunks (document, score, etc.). - `session_id`: The session ID used for the query. - (Optional) `latency_ms`: Time taken to generate the response. Typical use case: This route enables a conversational interface on top of the session's indexed knowledge base. It is used by frontend chat UIs and by MCP clients to perform retrieval-augmented generation queries.
Update or insert text chunks for a given session. This endpoint receives a list of chunked documents along with their metadata inside a `ChunkUpdateRequest`. It reprocesses and stores them in the Qdrant vector database under the specified `session_id`. Use this endpoint when you need to refresh or manually modify chunks for an existing session.
Upload and process new files for a session. The files are automatically chunked and embedded into the vector store (Qdrant). This replaces or extends existing session data depending on configuration.
No output schemas documented. Tool descriptions mention return types (SessionMeta, DeleteSessionResponse, SendChatResponse) but the actual field structure is not visible in the provided source. Without documented output schemas, LLMs cannot plan downstream operations or extract the right fields.
No error recovery guidance. Tool descriptions do not specify what errors are possible, what causes them, or how to recover. E.g., delete_session mentions HTTPException(404) in the docstring but no guidance on what the LLM should do if a session is not found.
Parameter naming inconsistency. 'session_name' is used in create_session, but 'session_id' is the unique identifier. In update_chunks and upload_files, both 'session_id' and 'session_name' are parameters, but it is unclear whether 'session_name' is a display name or a lookup key. Ambiguous naming forces LLMs to guess.
No parameter constraints on numeric or string fields. 'top_k' in send_chat is implied from context but not declared as a parameter with min/max constraints. 'session_id' and 'chunk_id' have no length, format, or regex constraints documented.
Monolithic operation: create_session couples file upload, extraction, chunking, embedding, and vector storage into one call. An LLM cannot decompose this if one step fails or needs customization. Consider splitting into extract_archive → chunk_documents → embed_and_store.
No idempotency guidance. Neither create_session nor update_chunks specify whether they are idempotent. If an LLM retries due to a transient error, are duplicate sessions or chunks created?
send_chat 'top_k' parameter is not declared in the input schema shown. The docstring mentions it as optional, but the schema does not include it. LLMs cannot pass parameters not in the schema.