MCP server providing document ingestion, semantic search, and vector database operations using ChromaDB with Ollama or OpenAI embeddings. Supports multiple document formats (PDF, DOCX, TXT, HTML, PPT, XLSX, ODT, MD, CSV, JSON, XML) and integrates with LangChain and Streamlit.
The server has 6 tools with moderate description quality but significant schema and parameter documentation gaps. Tool naming follows verb_noun conventions (ingest_document, retrieve, db_info, check_ollama, clear_db), which is good. However, most tools lack explicit input parameter type definitions visible in the source, schemas are not formally documented, and output schemas are entirely absent. The ingest_pdf tool is marked as deprecated, creating redundancy. Critical security and safety issues: clear_db is a destructive operation with no confirmation mechanism or permission gating. Error handling is present in code but not described in tool documentation. The descriptions are reasonable in length (most 50-150 chars) but lack actionable recovery guidance. Schema registration via fastmcp is likely present but not visibly detailed in the provided code excerpt.
Check Ollama service health and embedding model availability. Use this tool to diagnose connection issues with Ollama.
Clear all data from database and reinitialize an empty vectorstore.
Get ChromaDB and collection information. Samples up to 100 documents for source list.
Ingest document(s) from URL, folder, or file path. Supports multiple formats: PDF, DOCX, TXT, HTML, PPT/PPTX, XLSX, ODT, MD, CSV, JSON, XML, and more. Uses thread pool to process documents concurrently, batches DB writes for performance.
Legacy tool. Use ingest_document instead. Ingest PDF(s) from URL, folder or file path.
Retrieve top N chunks for given query. Clamps n to reasonable bounds.
Destructive tool (clear_db) has no confirmation mechanism, permission gating, or dry-run support. No explicit warning that the operation is irreversible.
Input parameter schemas lack complete type definitions and constraints visible in source. While fastmcp may auto-infer from Python type hints, explicit JSON Schema definitions are not documented in the provided code.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract return fields without explicit response structure documentation.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 36 | - | v1 |
ingest_pdf is marked deprecated in favor of ingest_document but remains exposed as a tool. This creates LLM confusion about which tool to use.
Parameter descriptions lack detail on constraints and valid formats. E.g., 'n' in retrieve says 'clamped to 1-100' in docstring but parameter description does not specify this range or default.
Error handling messages in code (e.g., 'Download failed: {exc}') are generic. No actionable recovery guidance (e.g., 'retry with a different URL' or 'check network connectivity').
Tool descriptions do not explicitly state which operations are side-effect-free vs. which modify state. ingest_document and clear_db clearly modify state, but this is not consistently documented.