Semantic code search MCP server for searching indexed code repositories by natural language meaning using embeddings and AST-aware chunking
The semantic-search-workspace-files tool exhibits good schema completeness with 6 well-defined parameters, each with type declarations and descriptions. The tool description is comprehensive (478 chars, above the 194-char average for production tools) and explains the WHAT, WHEN, and HOW effectively. Parameter descriptions are detailed and technical, averaging ~110 chars each. However, the response schema is not documented in the visible source code, only input parameters are specified. The description includes substantial technical depth about AST alignment and tree-sitter chunking, which aids LLM understanding but borders on implementation detail. No output schema, pagination behavior, or error responses are documented. The tool follows verb_noun naming ('semantic-search-workspace-files') and uses constrained inputs (searchMode enum, languages array), meeting modern standards. Risk annotation (READ_ONLY) is present and correct.
Call this tool when the user wants to search internal / indexed code by meaning—not only the repo open in the editor. Trigger intents include: search across or through the Lichens codebase; search in the company codebase; search the organization's private code. Similarity search over already-embedded chunks (not whole files). The index must exist on this server (built outside this MCP surface). `queryText` is a normal question or phrase in the user's language; optional `repository` filters to the same key used at index time (basename for single-repo index, or relative path for multi-repo); omit to search everything. See resource `code-crawler://workspace-repositories` for repo keys under the configured root. `nbResults` defaults to 10 (max 50). Each hit is a chunk with file path, line range, and preview—open files in the editor for full context.
Output schema not documented. The tool description does not specify the structure of returned search results (fields, types, example response). LLMs cannot predict what fields to extract or how to chain results into downstream tools without this.
Pagination and result limits not explicitly stated in tool description. Description mentions 'nbResults defaults to 10 (max 50)' but does not explain what happens when query matches >50 chunks across multiple files, or whether pagination cursors are available. Agents cannot reason about multi-round retrieval strategies.
Error handling not documented. No description of what errors the tool may return (e.g. 'index not found', 'repository key invalid', 'embedding service unavailable'), or how LLMs should recover. Missing actionable recovery guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 78 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Chaining IDs incomplete. If an LLM wants to open a result in the editor, it needs file path (provided), line range (provided), and repository context. The description mentions 'path prefix for retrieval context' but it is unclear whether repository key is returned or must be inferred from the original query parameter.
Parameter 'languages' is an unconstrained array of strings. Description lists valid values as 'javascript, typescript, python, c-sharp, cpp' but the schema does not enforce an enum. LLMs may pass invalid language IDs, forcing validation errors.