MCP server for ChromaDB vector database management with per-collection permissions, read-only mode, and structured tool access control
Strong foundation with well-structured tool definitions, clear descriptions, and proper input schemas. All 8 tools have descriptions (194-char baseline met), input schemas with types, and permission declarations. Tool names follow verb_noun pattern (list_, describe_, search_, get_, add_, delete_, collect_). However, output schemas are undocumented for most tools, parameter descriptions lack constraint details (ranges, formats, examples), and error handling guidance is minimal. The codebase shows sophisticated permission modeling (MCPToolPermission enum) and structured response types (MCPToolOutcome, MCPCollectionSummary), but these patterns are not fully leveraged in tool definitions visible to LLMs.
Adds new documents to a collection. Documents are embedded using the collection's configured model. Supports providing explicit IDs or letting the system generate them. Returns the IDs of added documents. Respects per-client daily document limits and per-document byte limits.
Finds all documents in a collection that contain a specific text substring, returning them in order of appearance. Useful for locating all references to a term without semantic search. Returns paginated results with total count.
Deletes documents from a collection by ID. Returns which documents were deleted, which were not found, and whether any were moved to trash instead of permanently deleted. Requires explicit delete permission separate from write permission.
Returns detailed schema information for a collection: field names, types, whether fields are required, example values from actual documents, and whether the collection allows extra fields beyond the schema. Provides the metadata structure needed for intelligent filtering and document composition.
Retrieves documents by ID or by pagination through a collection. Returns documents without distance scores (those only come from search). Supports offset-based paging with a hasMore flag to indicate whether more results exist.
Output schemas undocumented. MCPToolDefinition supports outputSchema field, but it is null for all 8 tools. LLMs cannot predict response structure, forcing them to parse unstructured text and plan downstream calls blindly.
Parameter constraints missing. 'nResults' (search), 'limit' (get_documents, get_file, collect_mentions), and 'offset' parameters lack min/max bounds. LLMs can pass absurd values (limit=999999) causing timeouts or memory exhaustion.
Error handling lacks recovery guidance. MCPToolOutcome.failure() returns isError=true with text only. No categorization (retryable vs user-fixable vs fatal), no actionable next steps, no invalid-value echoing for self-correction.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2025-06-18+ | v2 |
Retrieves all chunks of a specific file in order, paginated by chunk_index metadata. Returns chunks in the order they were split, with total count and hasMore flag. Supports returning either raw text concatenation or structured chunk list. Essential for reading complete files that were chunked during ingestion.
Lists all collections accessible to the client, showing document count, embedding model, distance metric, and vector dimension for each. Returns a summary view suitable for agent decision-making about which collection to query.
Searches one or more collections by embedding the query text and finding semantically similar documents. Returns documents ranked by distance metric (cosine, L2, IP), with the metric and embedding model identified so the agent understands the relevance scores. Supports filtering by metadata and respects per-client result limits.
Destructive operations lack confirmation. delete_documents has no dry-run or confirmation step. Agents can permanently delete documents without a safety gate, risking data loss.
Parameter descriptions lack format/constraint details. 'filter' (search) is described as 'Metadata filter conditions' with no syntax guidance. 'as' (get_file) enum is documented but other params lack similar clarity on valid ranges and formats.