Semantic search across Kiro CLI and IDE conversation history
The server defines 4 well-structured search tools with consistent parameter sets and clear semantic distinctions. All tools have descriptions (50-100 chars), proper JSON input schemas with typed parameters, and descriptive parameter labels. However, descriptions lack depth on use-case differentiation, output schemas are not documented, and error handling is not specified. Tool naming follows verb_noun pattern correctly (search_*), but descriptions do not explain WHEN to use one search tool vs. another or what recovery looks like on failure.
Search Kiro CLI conversation history only. Use this to find conversations from Kiro CLI sessions specifically.
Search conversation history across ALL WORKSPACES. Use this to find cross-project knowledge: user preferences, coding patterns, common solutions, and insights from all previous work.
Search Kiro IDE conversation history only. Use this to find conversations from Kiro IDE sessions specifically.
Search conversation history for the CURRENT WORKSPACE only. Use this to find workspace-specific context: past decisions, implementation details, bugs discussed, architecture choices in this codebase.
Output schemas not documented. Tool descriptions state 'Returns: Search results...' but do not specify field names, types, or structure. LLMs cannot infer what fields to extract or pass downstream.
Tool descriptions do not distinguish use cases clearly. 'Search conversation history across ALL WORKSPACES' (global) vs 'CURRENT WORKSPACE only' (project) is clear, but no guidance on WHEN to prefer CLI over IDE search, or whether similarity threshold 0.2 is appropriate for the domain. LLMs must guess.
Error handling and recovery not documented. Descriptions do not explain what happens if index loading fails, if a query is empty, if date filters are invalid, or how to retry. Agents have no recovery guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter 'threshold' lacks guidance on valid range and domain semantics. Description states 'Minimum similarity 0-1 (default: 0.2)' but does not explain: is 0.2 a cosine similarity score? Should users lower it for broader recall or raise it for precision? What does the model's embedding space mean to a domain expert?
Parameter 'context_size' description says 'Messages to include before AND after each match' but does not specify: is this per result or total? Is it symmetric? What if there are fewer messages available than requested? Ambiguity invites wrong usage.