MCP server for efficient code indexing and symbol retrieval — cuts token costs by up to 99%
astllm-mcp demonstrates solid definition quality with well-structured tool schemas and clear descriptions across all 11 tools. All tools have explicit registrations with name, description, and inputSchema visible in src/index.ts. Naming follows verb_noun convention consistently (index_*, get_*, search_*, list_*, invalidate_*). Descriptions are substantive (100-180 chars average, well within the 10-1024 baseline). Most parameters are typed and described. However, there are gaps in output schema documentation, error handling guidance, and some parameter descriptions could be more prescriptive about format/constraints. Security considerations for the index operations and API key injection are underdocumented.
Get all symbols in a specific file as a hierarchical outline (classes containing methods, etc.). Much cheaper than reading the file.
Get the file/directory structure of an indexed repository. Much cheaper than reading files — returns the tree with per-file language and symbol counts.
Get a high-level overview of an indexed repository: directory tree, file counts, language breakdown, symbol kind distribution.
Get the full source code of a specific symbol (function, class, method, etc.) by its ID. Uses byte-offset seeking for O(1) retrieval — much cheaper than reading the entire file.
Get full source code for multiple symbols in one call. More efficient than multiple get_symbol calls.
Index a local source code folder. Recursively discovers source files, parses ASTs, and stores symbols for fast retrieval.
Output schemas are not documented. While input schemas are complete and typed, there is no indication of what fields/structure each tool returns. LLMs cannot plan downstream calls or extract required data without knowing the response shape (e.g., what fields does index_repo return? Is there a task_id for polling? Does search_symbols return a count?). This forces LLMs to guess or hallucinate response structure.
Error handling guidance is absent. Tools do not document what errors can occur, what they mean, or what the LLM should do in response. For example: what happens if a GitHub repo URL is invalid? If tree-sitter parsing fails? If the index store is corrupted? LLMs receive no recovery guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Index a GitHub repository's source code. Fetches files, parses ASTs with tree-sitter, extracts symbols (functions, classes, methods, types), and saves to local storage for fast retrieval.
Delete the index for a repository, forcing a full re-index on the next index_repo or index_folder call.
List all indexed repositories with metadata (file count, symbol count, last indexed time).
Search for symbols by name, kind, language, or file pattern across the indexed repository. Returns matching symbols with signatures and summaries — no source code loaded unless you call get_symbol.
Full-text search across indexed file contents. Useful for finding string literals, comments, configuration values, or patterns not captured as symbols.
Parameter descriptions lack prescriptive format constraints. For example, 'repo' is described as 'Repository identifier: "owner/repo" or just "repo" if unique', but the description does not state whether it accepts full GitHub URLs, whether matching is case-sensitive, or how uniqueness is determined. 'file_path' says 'relative to repo root, e.g. "src/auth/login.ts"' but does not clarify whether forward slashes are required or if backslashes are allowed. Format and validation rules should be explicit in descriptions since LLMs do not read JSON Schema patterns reliably.
Security implications of API key injection for generate_summaries are underdocumented. The index_repo and index_folder tools mention 'requires API key' for summaries but do not state how the key is injected, what service it targets, or what permissions it requires. This is a security gap, agents need to understand credential scope.
List/search tools lack pagination and limit enforcement guidance. search_symbols defaults to limit=50 (good), but search_text defaults to limit=100 and neither tool documents what 'limit' means or what the maximum allowed value is. If an LLM passes limit=99999, does the tool cap it or error? Output schema is missing, unclear whether results include a total_count or next_cursor for pagination.
Destructive operation (invalidate_cache) lacks confirmation/dry-run pattern. Invalidating an index forces full re-indexing, which is expensive and potentially lossy if storage is corrupted. The tool has no confirmation step, no preview of what will be deleted, and no undo mechanism. Agents can accidentally nuke indices.
Tool chain references are incomplete. Tools like get_symbol accept a 'symbol_id' parameter but no tool description documents the exact format of symbol_id or guarantees that search_symbols returns symbol_ids in the correct format. This creates ambiguity, the LLM cannot reliably chain search_symbols → get_symbol without trial and error.
Idempotency and retry semantics are not declared. index_repo and index_folder support 'incremental' indexing, but it is unclear whether these tools are idempotent if called twice with the same input. Do they return the same result? Can they be safely retried? This matters for agents deciding whether to retry on transient failures.