A local code indexing server that uses ChromaDB and LlamaIndex to index and search code files with file system monitoring capabilities
This MCP server has significant definition quality gaps across multiple dimensions. While 8 tools are registered with basic descriptions, most lack proper input schema documentation, parameter descriptions are minimal or absent, and output schemas are entirely undocumented. Tool naming is reasonable (verb-noun pattern: index_directory, search_code, list_collections, delete_collection, get_file_content, start_watching, stop_watching, analyze_code), but descriptions are generic and lack the depth needed to guide LLM tool selection. Parameter definitions exist in the source but are under-described, for example, 'top_k' parameter in search_code lacks explanation of what the integer represents or valid ranges. No error handling guidance is visible in the code. Security concerns exist (file path parameters could enable path traversal attacks without visible sanitization). Output schemas are completely undocumented, forcing LLMs to guess what fields are returned. The server shows foundational structure but fails to meet production-grade documentation standards.
Analyze code snippets for syntax issues, complexity metrics, and other code quality metrics
Delete a ChromaDB collection and all its indexed data
Get the content of a specific indexed file
Index a directory of code files into a ChromaDB collection
List all available ChromaDB collections with their metadata
Search for code snippets in indexed collections using semantic search
Start file system monitoring for a directory to automatically update the index when files change
Stop file system monitoring for a directory
Output schemas completely undocumented. No tool documents what fields are returned, types of those fields, or structure of the response. LLMs cannot plan downstream operations or extract needed data.
Parameter descriptions are minimal or missing entirely. 'collection_name' appears in multiple tools but is never explained. 'top_k' parameter in search_code lacks explanation of valid ranges or meaning. No guidance on which parameters are required vs optional.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | <=2025-11-25 | v2 |
Tool descriptions are generic and lack context for LLM selection. 'Search for code snippets in indexed collections using semantic search' does not explain when to use search_code vs get_file_content, what preprocessing occurs, or what format results use. Description length is 81 chars (below the 10-1024 baseline but lacking depth).
No error handling guidance visible. Code snippet shows file system operations and ChromaDB interaction but provides no error classification, recovery hints, or actionable error messages. Agents cannot know if failures are retryable, user-fixable, or fatal.
Path traversal vulnerability. get_file_content and start_watching accept 'directory' and 'file_path' parameters with no visible sanitization. Malicious paths like '../../../etc/passwd' could expose sensitive files. No validation rules documented.
Destructive operations (delete_collection) lack confirmation or dry-run mechanism. No evidence of permission gating or audit logging. An agent could irreversibly delete indexed data.
Tool composition issues. start_watching and stop_watching accept 'directory' but do not document the relationship to 'collection_name' or explain when a watch auto-updates which collection. Multi-step workflows (watch → collection creation → indexing) are unclear.
No input validation rules documented. The 'language' parameter in analyze_code has no enum or pattern constraint. 'top_k' in search_code lacks min/max bounds. Unconstrained parameters invite hallucinated invalid values.