Local agentic RAG system using Model Context Protocol with LLM reasoning, vector search, and document retrieval powered by ChromaDB and Ollama
This RAG server has moderate definition quality with significant gaps. Tool naming follows verb-noun conventions (query_agent, search_documents, add_documents), which is good. However, descriptions are inconsistent, some tools have adequate descriptions (query_agent, search_documents, add_documents) while others lack sufficient detail (health_check, get_stats are extremely brief). Parameter descriptions are present but often generic. Input schemas are visible and properly typed (string, boolean, integer, array), but output schemas are not documented. Error handling is largely absent, no recovery guidance, no actionable error messages, no dry-run patterns for destructive operations like clear_documents. The most critical issue: clear_documents is destructive with ZERO protection, confirmation, or validation. Security concerns include no audit logging or permission gates visible. Tool composition is reasonable (separate tools for search vs query), but clear_documents should require confirmation. Overall, this is a C-range server with functional definitions but significant production readiness gaps.
Add documents to the vector store
Clear all documents from the vector store
Get statistics about the vector store and system configuration
Health check endpoint to verify API and vector store status
Query the RAG agent with optional context retrieval
Search for documents in the vector store using semantic similarity
clear_documents is destructive (WRITE_DESTRUCTIVE) with no confirmation, dry-run, or permission gate. Agents can wipe the entire vector store without any safeguards.
Descriptions for get_stats and health_check are under 20 characters and provide minimal LLM guidance. 'Get statistics about the vector store and system configuration' (67 chars for get_stats) and 'Health check endpoint to verify API and vector store status' (59 chars for health_check) are borderline but lack WHEN/WHY context.
No output schemas documented. LLMs cannot infer what fields are returned from query_agent (is it {response: string, sources: []}?), search_documents (does it return relevance scores?), or add_documents (success count?). This forces agents to guess downstream field names.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
No error handling guidance. Tools lack recovery suggestions (e.g., 'If query_agent returns empty context, try search_documents with simpler keywords'). No classification of errors (retryable vs. user-fixable vs. fatal).
Parameter descriptions are present but generic. 'Number of results to return' (n_results in search_documents) lacks constraints (min/max). 'Number of documents to retrieve for context' (n_results in query_agent) similarly unconstrained. Unbounded values can cause performance issues.
No idempotency guarantees or retry safety documented. If add_documents fails mid-operation, are documents partially added? Can an agent safely retry? This is critical for multi-step RAG pipelines.
No pagination or result limits documented. search_documents defaults to n_results=3, which is good, but no explicit cap stated in description. If an agent sets n_results=10000, is there a server-side limit? Unbounded results can exhaust context windows.
No permission gates or audit logging visible. No log of which agent/user called clear_documents or add_documents. No permission checks (e.g., is this agent allowed to clear the store?). Critical for compliance and incident response.