MCP server for semantic search using local Qdrant and Ollama (default) with support for OpenAI, Cohere, and Voyage AI
The server provides 11 tools with reasonable naming conventions (most start with action verbs: index_, search_, list_, get_, delete_, add_). However, definition quality is inconsistent. Tool descriptions are present but often generic (e.g., 'Search indexed code with semantic similarity' lacks detail on output structure, when to use vs other tools, or prerequisites). Parameter schemas are visible in the tool list but lack full JSON Schema validation details (no min/max bounds, enum constraints, or format specifications for date fields like 'sinceDate'). Output schemas are not documented anywhere in the provided source. Error handling patterns are absent, no guidance on recovery or categorization of failures. The codebase shows sophisticated implementation (TreeSitterChunker, BM25SparseVectorGenerator, path validation), but the MCP interface itself is under-specified for agent usage.
Add documents to vector store
Delete a vector collection
Delete documents from vector store
Search across code and git history simultaneously
Get detailed information about a collection
Index a codebase from scratch or force re-index
Index git history for a repository
List all vector collections in Qdrant
Search indexed code with semantic similarity
Output schemas are not documented for any tool. LLMs cannot plan downstream operations or extract required fields (e.g., does search_code return a 'relevance_score' field needed by subsequent tools?). This forces agents to guess at response structure.
Parameter descriptions lack constraint details. 'sinceDate' accepts ISO 8601 format but this is stated nowhere, LLMs may pass invalid formats. 'limit' parameters have no min/max bounds, inviting calls with limit=0 or limit=999999. 'extensions' and 'ignorePatterns' are arrays but lack examples of expected format (file extension syntax, glob patterns, regex?).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 55 | - | v1 |
Search indexed git history
Perform semantic search across indexed collections
Destructive tools (delete_documents, delete_collection) lack error handling guidance and confirmation patterns. No indication that these operations are irreversible or that the LLM should confirm before executing. No recovery guidance (e.g., 'deleted items cannot be recovered; consider exporting first').
Tool descriptions are generic and lack 'when to use' guidance. 'Search indexed code with semantic similarity' does not explain when to prefer this over federated_search, what 'useHybridSearch' does, or what output format to expect. Descriptions average ~40-50 characters; rubric baseline is 194 chars for A+ tools.
Path traversal validation is implemented in CodeIndexer.validatePath() but only partially documented. Tool descriptions for index_codebase do not warn about path constraints or symlink handling, leaving agents unaware of security boundaries.
No pagination or result limits documented. Tools like search_code accept 'limit' but do not specify min/max (e.g., is limit=1 valid? limit=1000000?). No indication of default behavior or whether results are paginated. For large codebases, unbounded searches could return massive result sets, overwhelming context windows.
Error handling is not visible in the MCP schema definitions. The code shows sophisticated error detection (secret detection in metadata.ts, file scanner error handling) but no documentation of how errors are reported to the agent or how to recover from common failures (e.g., 'Collection not found', 'Invalid syntax in code chunk').