Semantic search over a local document corpus (on-device embeddings; nothing leaves the host). Use search_corpus for 'do we know anything about X' recall — results carry source paths and cosine scores; treat low scores as weak matches. get_file returns a full indexed document.
Two well-named, read-only tools with clear descriptions and proper input schemas. search_corpus and get_file follow verb_noun naming and include parameter types/descriptions. However, output schemas are not formally documented in the code, only described in docstrings. Error handling returns structured dicts but lacks recovery guidance (e.g., 'try search_corpus to discover paths'). Parameter constraints (k: 1-20, query: min 3 chars) are enforced at runtime but not declared in schema. Tool annotations (readOnlyHint) are present, which is good. Missing: explicit output schema documentation, per-parameter validation hints in descriptions, and actionable error messages that guide the LLM to next steps.
Full text of an indexed document. Only paths present in the index are readable — this tool is not a general filesystem reader.
Semantic search of the corpus. Returns up to k chunks with source path, heading path (for chunked docs), cosine score, and text.
Output schemas not formally documented. Docstrings describe results (chunks, paths, scores, text) but no JSON Schema definition is visible in code. LLMs cannot reliably parse or chain results without explicit output type declarations.
Parameter constraints (k: max 20, query: min 3 chars) are enforced in do_search() but not declared in the input schema. Schema shows no minLength, maxLength, or minimum/maximum. LLMs cannot see these bounds and may pass invalid values.
Error messages lack recovery guidance. 'query must be at least 3 characters' is clear, but 'path not in the indexed corpus' should suggest 'use search_corpus to discover valid paths', which the code does in a hint field, but not consistently across all errors.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 72 | 2026-07-28+ | v2 |
get_file returns different response structures depending on success/failure path (live file vs indexed chunks vs error). No unified output schema means LLMs must handle multiple response shapes, increasing parsing errors.
search_corpus truncates results to 2000 chars and get_file to 50000 chars, but these limits are not documented in tool descriptions. LLMs may assume full text is returned and fail to request pagination or multiple calls.