An MCP server that provides semantic search and RAG (Retrieval-Augmented Generation) capabilities for documentation using ChromaDB and embeddings
The server defines 5 tools with mostly complete schemas and descriptions. All tools have clear action-verb names (search_, list_, get_, ingest_) and explicit JSON Schema input parameters with type definitions. Descriptions are present for all tools and most parameters. However, there are critical gaps in output schema documentation, error handling guidance, and parameter descriptions lack specificity around formats and constraints. The server would benefit from more detailed parameter documentation and explicit output structure definitions for LLM planning.
Detailed stats for a specific documentation collection.
Retrieve the full stored chunk for a given chunk ID.
Index new documentation from the local filesystem into a collection.
List all available documentation collections and their document counts.
Search indexed documentation semantically.
Output schemas not documented. All tools return dict/list but LLM cannot plan downstream calls without knowing exact field structure. For example, search_docs returns a list of objects with 'content', 'source', 'section', 'url', 'relevance_score', but this structure is only visible in implementation code, not in tool definition. LLMs need explicit return schema to extract the right fields.
Parameter descriptions lack format constraints and ranges. 'query' in search_docs has no guidance on length, special characters, or expected language. 'n_results' defaults to 5 but no min/max bounds stated. 'path' in ingest_docs is described as 'Absolute path' but no validation guidance for traversal attacks or file existence. Baseline from rubric: 100% of A+ tools specify format, range, and allowed values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 8 | - | v1 |
Error handling provides no recovery guidance. get_doc_context returns {'error': 'Chunk not found'} with no suggestion for the LLM to retry search_docs or check the collection name. Baseline: error responses must tell the LLM what to do next.
search_docs 'collection' parameter lacks enum constraint. Defaults to 'samsung_tv' but no list of valid collection names provided. LLMs may hallucinate collection names. Should enumerate available collections or require a prior call to list_collections.
ingest_docs is destructive (writes to store) but description does not explicitly state this. Baseline: 'If the tool modifies state (creates, updates, deletes, sends), the description must say so.' No mention of idempotency or confirmation pattern.
Tool descriptions do not include dependency hints or explain when to use which tool. E.g., 'search_docs searches existing docs, if you want to index new docs first, call ingest_docs.' No guidance on the workflow or when to invoke each tool in sequence.