Codebase Graph + Vector RAG Indexer for Java, TypeScript, JavaScript, Kotlin, Rust, Python, Groovy, C/C++, C#, Build Systems, and HTML/CSS codebases
knot provides 6 well-scoped, READ_ONLY tools for codebase exploration. All tools have descriptions and documented input schemas with typed parameters. However, descriptions lack depth and strategic guidance for LLM selection. Parameters lack granular constraints (enums, min/max bounds). Output schemas are not documented in visible source. Error handling and recovery guidance are absent. The tools follow verb_noun naming correctly but descriptions don't explain WHEN to use each tool vs. alternatives or dependencies between them. Most descriptions are 30 - 60 chars, below the 50 - 200 char optimal range for LLM reasoning.
Inspect file structure and entity declarations
Reverse dependency lookup (impact analysis)
List files in the indexed codebase
List dependencies for a repository
List all indexed repositories with optional name filtering
Find entities by semantic meaning with dependencies
Output schemas not documented. Tool descriptions do not specify what fields are returned or their types. LLMs cannot plan downstream calls or extract needed data without knowing the response structure.
Parameter descriptions lack strategic context. 'Natural language search query' for search_hybrid_context does not explain: What kind of semantic meaning? How does this differ from keyword search? When should I call this vs explore_file? This forces LLMs to guess.
No numeric constraints (min/max) on max_results parameter. Unbounded integers let LLMs pass absurd values (max_results=999999) that may timeout or exhaust memory.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
repo_name and pattern parameters are optional but lack clarity on default behavior. What happens if repo_name is omitted? Search all repos? Return an error? This ambiguity forces LLMs to reason about side effects.
No error handling guidance in tool descriptions. What if a file_path does not exist? What if a repo_name is invalid? How should the LLM recover? Agents need classification (retryable, user-fixable, fatal) to plan alternatives.
Pagination and result limits not addressed. list_files and list_repositories could return large datasets. No documentation of max results, offset/limit parameters, or total_count fields. Large responses risk exhausting LLM context.
Tool composition guidance missing. No documentation of typical workflows (e.g., 'search_hybrid_context → explore_file → find_callers'). LLMs must infer dependencies from names alone, increasing reasoning overhead.