An MCP server that provides semantic code search and indexing capabilities for multiple programming languages (Python, Java, C++, JavaScript, TypeScript, Go) using vector embeddings and incremental change tracking.
This server has three tools with basic input schemas and descriptions, but significant quality gaps prevent it from reaching production-grade status. All three tools accept project_path as a string parameter (present with description), but lack comprehensive parameter validation, output schema documentation, and error handling guidance. Tool descriptions are present but generic (10-100 chars range). The 'query' tool has a numeric parameter (top_k) with a default, which is good, but lacks min/max bounds. No tool descriptions explain what happens on error or what the return structure contains. Critical patterns missing: structured output documentation, error recovery guidance, idempotency clarification, and permission gates. The codebase shows background threading for indexing, but no indication of how failures are handled or reported to the LLM. Overall design treats tools as simple wrappers around Python methods without LLM-specific guidance.
Ensure full indexing of a project, starting initial full index in background if not already present, and starting periodic incremental indexing (every 5 minutes)
Search for code blocks semantically similar to the provided text query using vector embeddings, returning top_k results with metadata
Get the index status of a project, including last index time, total files indexed, file hashes, and index metadata path
No output schemas documented. Tools return untyped dicts (status responses, query results, error responses). LLMs cannot predict response structure, forcing them to reason about fields generically. This violates pattern:tool and breaks tool chaining.
Error handling is minimal. All three tools wrap calls in bare try-except returning {"error": str(e)}. No error categorization (retryable vs. fatal), no recovery guidance, no invalid-value feedback. An LLM seeing 'error: Collection not found' has no idea whether to retry, call full_index first, or ask the user.
Parameter validation missing. project_path accepts any string; no check for existence, read permissions, or valid project structure. text in query accepts any string; no length limits (could trigger huge embeddings). top_k has no min/max constraints (could request 10,000 results, causing context explosion or timeout).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Tool descriptions lack WHEN/WHY guidance. 'full_index' description does not clarify idempotency or resource cost (spawns background threads, scans entire project). 'status' does not explain whether it's safe to call during indexing or if it blocks. 'query' does not warn that it requires a built index (calls to query before full_index will fail silently).
No idempotency declared. full_index spawns threads, calling it twice in quick succession may cause race conditions or duplicate indexing. No deduplication or atomicity guarantees mentioned. Agents retry on ambiguous failures; non-idempotent tools risk duplicate side effects.
No permission checks or audit trail. Tools do not verify the caller has permission to index or query a project. No logging of who indexed what, when, or what was returned. For a code indexing tool, this is a security and compliance gap.
API key exposure risk. Code shows OpenAI client initialized via environment variables (OPENAI_API_KEY, OPENAI_BASE_URL, OPENAI_MODEL_NAME), which is correct server-side. However, no validation that these are set; if missing, tools silently fail with cryptic API errors rather than clear guidance.
No pagination for query results. If top_k=5 is insufficient, LLM cannot request the next 5 results without re-querying with identical text. Baseline pattern requires limit/offset and total_count or next_cursor for tools returning lists.