MCP server for querying and retrieving documents from a ChromaDB vector database collection
This server has three read-only ChromaDB tools with basic schemas and descriptions. All tools have input schemas with types and descriptions, and all have tool descriptions. However, the descriptions are verbose (200-250 chars) and contain arglist-style documentation (Args: ...) that duplicates parameter names already visible in the schema. More critically, output schemas are completely undocumented, tools return free-form strings with no structure, requiring LLMs to parse unstructured text. The tools lack error handling guidance, permission declarations, and detailed parameter constraints. Tool names are verb-noun style (acceptable), but the parameter naming could be more explicit (e.g., 'max_results' instead of 'n_results' for consistency with get_all_company_docs' 'limit'). No tool annotations (readOnlyHint, etc.) are present. The error handling returns raw stack traces, which violates the pattern:recovery-guide requirement. For a vector search tool, pagination is not implemented (get_all_company_docs has a hard limit of 100, no offset/cursor). Overall, these are minimal read-only tools that work but lack production-grade polish.
Get all documents from the company-docs collection Args: limit: Maximum number of documents to return (default: 100)
Get the total number of documents in the company-docs collection
Query the company-docs collection in ChromaDB Args: query: The query text to search for n_results: Number of results to return (default: 5)
Output schemas completely undocumented. All three tools return free-form strings with no structured schema documentation. LLMs cannot plan downstream operations or extract specific fields. Violates pattern:tool requirement that tools must have documented output schemas.
Error handling returns raw stack traces (traceback.format_exc()). Stack traces are opaque to LLMs and provide no recovery guidance. Violates pattern:recovery-guide, errors must tell the agent what to do next.
Descriptions contain Args: docstring sections that duplicate parameter information already in the schema. Description text should explain WHAT the tool does and WHEN to use it, not restate parameter names. Descriptions are 200-250 chars, at the upper bound of the 10-1024 range, and diluted by redundant content.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
No pagination implementation. get_all_company_docs accepts a hard limit but no offset or cursor. For large collections, agents cannot iterate through results without loading 100+ documents into context. Violates pattern:paginated-result.
No tool annotations (readOnlyHint). While all tools are read-only (low risk), annotating them with readOnlyHint=true would signal to LLM clients that these are safe to retry and require no confirmation. Missing a simple quality signal.
Parameter naming inconsistency. query_company_docs uses 'n_results' while get_all_company_docs uses 'limit'. Both are numeric limits, should use consistent naming (e.g., 'limit' or 'max_results' throughout). Inconsistency increases LLM confusion when switching between similar tools.