Advanced MCP server for code intelligence using Tree-sitter, DuckDB, and ChromaDB
This server has significant structural issues. While tool names follow verb-noun conventions (find_usages, analyze_impact, etc.), parameter and output schema documentation is incomplete or missing. Several tools lack proper input parameter documentation in the visible code. The server registers 14 tools across two files (server.py and server_minimal.py), but only partial schema information is visible. Most critically: output schemas are not documented anywhere in the provided source, descriptions are present but many are generic (e.g., 'Find all usages of a function, class, or variable' lacks context for when to use vs. similar tools), and error handling guidance is absent. The codebase shows professional structure (pydantic, async/await, configuration), but tool definitions lack the precision required for reliable LLM integration.
Add a test symbol to the database.
Analyze the impact of changing a symbol.
Find circular dependencies in the codebase.
Find all dependencies of a symbol.
Find code similar to the given snippet.
Find a symbol by name.
Find all usages of a function, class, or variable.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract fields without knowing return types and structure. E.g., find_usages returns 'List of locations where the symbol is used', but what fields comprise each location? Is it {file, line, column}? {file, range, context}?
Parameter descriptions lack actionable constraints. E.g., 'depth' in find_dependencies is described as 'Maximum depth to traverse' with default=3, but no validation guidance. What is the minimum? The maximum? If depth=0, what happens? Same issue with 'limit' in find_similar_code and semantic_search.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Get the call graph for a function.
Get current index statistics.
Index a single file into the code graph.
Search code using tree-sitter patterns.
Search for code using semantic similarity.
Test that the server is working.
Update the code index.
The visible code shows `Input: {}` but provides no documentation of what these tools do beyond the name.
No error handling or recovery guidance in any tool description. E.g., 'find_usages' does not explain: What happens if the symbol is not found? Should the LLM try a different search strategy? Call find_symbol first? Or is this a fatal error?
Descriptions are too generic and lack context for tool selection. E.g., 'find_usages' vs 'semantic_search', when should the LLM choose one over the other? 'Find all usages of a function, class, or variable' does not differentiate it from search tools. Similarly, 'analyze_impact' does not explain its relationship to 'find_dependencies'.
No indication of which tools modify state vs. read-only. The 'Risk' field in the metadata shows 'update_index', 'add_test_symbol', and 'index_file' are WRITE operations, but the descriptions don't warn agents of side effects. Per pattern, modifying tools must state consequences and reversibility.
Resource definitions are incomplete. The visible server.py excerpt shows @self.server.list_resources() is declared but the resource details (URI, MIME type) are truncated mid-line ('descripti'). Cannot assess resource schema completeness.
Tool 'add_test_symbol' with no input parameters (Input: {}) and name 'add_test_symbol' suggests test-only functionality. This should not be a public tool; it pollutes the tool namespace and misleads LLMs into thinking they should invoke it in production scenarios.
Parameter naming inconsistency: 'change_type' in analyze_impact uses enum-like values (modify, delete, rename) but is not declared as an enum in the visible schema. Free-form string invites hallucinated values like 'update', 'refactor', or 'rewrite'.
No pagination guidance. Tools like find_usages, find_similar_code, and semantic_search return lists without limit/offset or cursor parameters in their signatures, risking context window exhaustion. Default limit=10 is mentioned for find_similar_code and semantic_search, but find_usages has no stated limit.