A FastAPI-based RAG (Retrieval-Augmented Generation) server that ingests PDFs and CSVs into a persistent Chroma vector database and provides semantic search with optional LLM summarization (Ollama/OpenAI). Exposes /tools endpoints for document ingestion and semantic search.
The server provides three tools with reasonable structure and descriptions, but exhibits significant gaps in parameter documentation, output schema clarity, and error handling guidance. Tool naming follows verb_noun convention (ingest_*, search_*), which is good. Descriptions are present and substantive (avg ~140 chars), meeting baseline minimums. However, parameter descriptions are sparse or missing (e.g., 'k' parameter lacks explanation of what the score means), output schemas are not formally documented (SearchHit is a Pydantic model but not described in the tool metadata), and error handling lacks recovery guidance. The server uses HTTP (good for protocol readiness) but does not expose schema/output metadata to MCP clients. No tool annotations, no pagination hints for search_docs, and no guidance on how to interpret relevance scores.
Ingest a local CSV into the vector store. If 'text_col' is provided, each row uses that column as its text. If not, the loader concatenates row values into a single text string. Metadata will contain 'row' indices and 'source' path if available.
Ingest a local PDF into the vector store. Uses PyPDFLoader to extract page text (no OCR). Normalizes metadata and tags chunks with 'doc_type'. Splits pages into chunks before adding to Chroma.
Semantic search with relevance scores. Returns a list of hits (content + score + normalized metadata). Client UIs can show badges (doc_type / page / row) and preview pane (e.g., PDF page text, CSV row).
Output schemas not formally documented in tool definitions. SearchHit and SearchResp are Pydantic models internal to the server, but the MCP tool metadata does not declare what fields are returned. Clients and LLMs cannot infer the structure of results without reverse-engineering the response or guessing.
Parameter 'k' in search_docs lacks description of what the score metric represents (Chroma's distance metric, normalized 0-1, etc.). LLMs cannot reason about result quality or threshold interpretation without this context.
No pagination support or result limit enforcement documented. search_docs accepts k=1..50 but does not explain what happens when k=50 returns a large result set. No guidance on token cost or context window impact.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Error responses lack recovery guidance. HTTPException with 400 status for missing files does not tell the LLM what to do next (e.g., 'Try listing available files' or 'Check the path and retry'). Stack traces are not actionable for agents.
No tool annotations present. ingest_pdf and ingest_csv should declare destructiveHint=true (they modify state). search_docs should declare readOnlyHint=true. This blocks agents from reasoning about side effects and retry safety.
No dry-run or confirmation pattern for ingest operations. Agents cannot preview what will be added before committing chunks to the vector store. High risk of accidental ingestion of wrong files or duplicates.
Metadata normalization is not documented in tool descriptions. LLMs do not know that ingest operations will add 'source', 'page', 'row', 'doc_type' fields to returned results. Client UIs depend on this contract but agents have no visibility.
search_docs description mentions 'relevance scores' but does not explain the metric or interpretation. Is a lower score better? Can scores exceed 1.0? What does a score of 0.3 vs 0.8 mean in practice? LLMs cannot set thresholds or reason about confidence without this.