MCP server for local code indexing — semantic search across any codebase
The server provides 6 tools with reasonable naming and mostly adequate descriptions, but suffers from incomplete output schema documentation, missing parameter descriptions on some fields, and inconsistent error handling guidance. Most tools are READ_ONLY which limits complexity, but the lack of explicit output schema definitions and parameter constraints prevents this from reaching a higher tier. Tool names follow verb_noun conventions (search_code, search_symbol, get_file_overview, index_status, get_project_summary, reindex) which is good, but descriptions lack the specificity and LLM-optimization guidance found in A-grade tools. The semantic search and symbol lookup tools are well-conceived for code exploration, but parameter descriptions are sparse and output structures are documented only in code, not in schema declarations.
List all symbols (functions, classes, methods, routes) in a file. Pass the relative path, e.g. 'mcp_server.py' or 'code_index/parser.py'
Return the auto-generated project summary (languages, structure, purpose, frameworks). Generated/updated automatically on each reindex.
Check if the code index exists and show statistics.
Rebuild the code index. By default does incremental update (only changed files). Set full=True to rebuild from scratch.
Semantic search across the codebase. Find code by meaning, not just text. Examples: "Chrome WebDriver cleanup", "database connection handling", "Flask route for login"
Search for a symbol by name. Optionally filter by type: function, method, class, interface, trait, enum, namespace, route, constant, file_summary. Examples: search_symbol("start_analysis"), search_symbol("cleanup", "method"), search_symbol("parsePage", "function")
Output schemas are not explicitly documented. Tools return markdown-formatted strings or mention returning structured data (e.g., search_code returns 'score', 'symbol_type', 'symbol_name', 'file_path', 'line_start', 'line_end', 'source_code') but these are not declared in a schema. LLMs cannot predict response structure and must parse unstructured output.
Parameter descriptions are minimal or missing for optional fields. The 'symbol_type' parameter in search_symbol lists valid enum values ('function, method, class, interface, trait, enum, namespace, route, constant, file_summary') but does not explain what filtering by type means or when to use it. 'limit' parameter in search_code lacks guidance on result quality vs. token cost tradeoff.
No explicit enum constraints in schema for symbol_type. The description lists valid values as comma-separated text rather than as a JSON Schema enum, forcing LLMs to parse freetext and inviting hallucinated values.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 51 | - | v1 |
Error handling lacks recovery guidance in several cases. search_code catches EmbeddingError and returns a message about 'model still loading', which is helpful, but generic Exception errors return 'Try running reindex', vague guidance that may not address the actual issue. index_status and get_project_summary do not document what errors might occur or how to recover.
Pagination is not addressed. search_code and search_symbol both accept a 'limit' parameter with a sensible default (10), but there is no 'offset' or 'cursor' mechanism to retrieve additional results if the user needs more than the default. Tools should support pagination per pattern:paginated-result.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). The server correctly marks reindex as WRITE risk and search/get tools as READ_ONLY in comments, but these are not declared as tool annotations in the MCP protocol, so clients cannot reliably determine tool safety without parsing descriptions.