Precise local semantic code search plugin — indexes with Go AST/tree-sitter and embeds with Ollama or LM Studio
Lumen exposes 2 tools with basic descriptions and partial schemas. Both tools have descriptions (index: 29 chars, search: 43 chars) but fall below the 50-char production baseline. The 'index' tool description is critically short ('Index a project for semantic search') and lacks context on WHEN to use it, WHAT happens, or prerequisites. The 'search' tool is similarly terse ('Perform semantic search on indexed code'). Input schemas are partially visible with type information (strings, integers, booleans) and brief parameter descriptions, but lack important constraints: the 'backend' parameter accepts arbitrary strings instead of an enum constraining to ['ollama', 'lmstudio']; the 'limit' parameter has no bounds specified (should be 1-100 or similar); no defaults are documented; no output schemas are visible. Error handling is absent, there is no guidance on what happens if the project_path doesn't exist, if embedding fails, or if search returns zero results. The tools follow basic naming (verb_noun: index, search) but lack the context density that LLMs need for reliable tool selection and composition.
Index a project for semantic search
Perform semantic search on indexed code
Tool descriptions are critically short (29 - 43 chars vs. 50 - 200 char production baseline). 'Index a project for semantic search' does not explain WHEN to use this tool, what embedding model is needed, or what happens on success/failure.
The 'backend' parameter accepts free-form strings instead of a constrained enum. Description says '("ollama" or "lmstudio")' but this is not enforced as a JSON Schema enum. LLMs may hallucinate other backends like 'vllm' or 'huggingface'.
The 'limit' parameter in 'search' has no bounds specified. Unbounded integers let LLMs pass absurd values (limit=1000000) that may timeout or exhaust resources. Should document min=1, max=100 (or appropriate limit).
No output schemas are documented for either tool. LLMs cannot know what fields to expect (e.g., does 'search' return file_path, line_number, match_score, code_snippet?). This prevents downstream reasoning and tool chaining.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 27 | - | v1 |
No error handling guidance. What happens if the project_path does not exist? If embeddings fail? If index is corrupted? If search returns no results? Agents have no recovery path.
Parameter 'model' in 'index' has a default ('$LUMEN_EMBED_MODEL or all-MiniLM-L6-v2') documented in description but not in JSON Schema. Schema should explicitly state the default value so LLMs know they can omit it.
The 'search' tool returns results but no pagination is documented. If a query matches 1000 code snippets, how many can be returned? Is there a next_cursor or offset? This risks context window overload.