Fully-offline semantic search over local files — powered by Ollama. Provides MCP tools for semantic search and agent context retrieval.
skylakegrep exposes 2 semantic search tools with reasonable parameter schemas but uneven description quality and missing output documentation. Tool naming is clear (search, agent_context) but descriptions vary in actionability. Parameter schemas are well-structured (types, descriptions, ranges) but output schemas are not formally documented. Error handling and recovery guidance are absent. This is a C-grade server, solid parameter definitions but lacks the output documentation, error guidance, and description clarity expected of production tools.
Agent-optimized context retrieval. Returns JSON snippets with agent summaries. Delegates to run_agent_context_search for structured agent-mode retrieval with top_k=8 default, snippets enabled, no rerank.
Semantic search over indexed project files. Returns ranked results with snippets. Delegates to storage.search core retrieval pipeline.
Output schemas are not documented. Tools return ranked results with snippets (search) and JSON snippets with agent summaries (agent_context), but the actual response structure (fields, types, array vs object, pagination) is not declared anywhere in the schema or description.
Descriptions lack actionable detail. 'Semantic search over indexed project files. Returns ranked results with snippets.' does not explain WHEN to use search vs agent_context, what quality of results to expect, or what prerequisites must be met (e.g., must index be built first?). Descriptions should be 50-200 chars and answer WHAT, WHEN, and any prerequisites.
No error handling or recovery guidance. If indexing fails, if ripgrep times out, or if no results are found, the tools do not document what error the LLM will receive or what to do next. Error responses must tell the LLM what to do: 'Index not found. Try running: skygrep index /path/to/repo'.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2026-07-28+ | v2 |
Parameter descriptions do not explain constraints or expected formats. Example: 'rg_timeout' has description 'Ripgrep timeout in seconds (default: 0.75)' but does not state if it accepts fractional seconds, what the min/max bounds are, or what happens if a timeout is exceeded. Descriptions should include format, range, and side effects.
Distinction between search and agent_context is unclear. Both tools accept nearly identical parameters (query, path, top_k, languages, include, exclude) but agent_context adds 'strict', 'semantic_only', 'support_per_path', 'rg_timeout'. The descriptions do not explain when an LLM should call search vs agent_context. 'Agent-optimized context retrieval' is vague, does it mean different ranking, different output format, or both?
No pagination declared. If search returns 10+ results by default, the tools should document: do results return an array, how many items max, is there a cursor or next_token for pagination, what happens if top_k is omitted? Large result sets can exhaust context windows, tools should cap results and offer pagination.