Claude Code plugin for local semantic search over document folders
The server has 4 well-named tools with generally strong descriptions and good schema definitions. Naming is action-verb oriented (index_folder, semantic_search, get_index_status, reindex_file). Descriptions are comprehensive and explain WHAT, WHEN, and dependencies. Tool annotations are present for safety hints. However, there are gaps: the output schema for semantic_search is not fully documented in the source (results structure is implied but not explicit); error handling guidance is minimal; parameter relationships between folder_path and db_path across tools lack coordination; and no validation constraints on mode enum values in semantic_search.
Get status information about the current search index. Returns total chunks, list of indexed files with chunk counts, and file type distribution.
Index or re-index all documents in a folder for semantic search. Scans the folder for supported document types (.txt, .md, .pdf, .docx, .pptx, .csv), extracts text, splits into chunks, computes embeddings, and stores them in a local vector database. Only processes files that have changed since the last indexing run. Safe to call multiple times — unchanged files are skipped automatically.
Force re-index a single file, ignoring the content hash cache. Deletes existing chunks for this file, re-parses, re-chunks, re-embeds, and stores new chunks. Useful when you know a file has changed or when parsing was updated.
Search indexed documents using natural language. Finds the most relevant document chunks matching the query using semantic similarity. Returns ranked results with source file paths and relevance scores. Use mode='hybrid' to combine vector search with full-text search via Reciprocal Rank Fusion for better keyword + semantic matching. The folder must be indexed first with index_folder.
semantic_search output schema not documented. The tool returns results but the exact structure (fields, types, nested objects) is not visible in the source. LLMs cannot infer the shape of search results without an explicit schema.
mode parameter in semantic_search accepts string with description 'vector' (default) or 'hybrid' but no enum constraint is enforced. LLMs may pass invalid values like 'bm25' or 'semantic' without validation. Use Pydantic Literal or JSONSchema enum.
db_path parameter is optional with default=null across multiple tools (semantic_search, get_index_status, reindex_file) but is NOT exposed in the tool signatures shown in server/main.py. The parameter exists in indexer.py but not in the FastMCP-decorated functions. This creates a mismatch: the underlying logic expects db_path but the tool interface doesn't expose it. Inconsistent parameter availability across tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 61 | 2026-07-28+ | v2 |
get_index_status has a generic description ('Get status information about the current search index') that lacks actionable WHEN guidance. When should an agent call this vs semantic_search? Clarify the distinction and add a use-case hint.
No error recovery guidance in any tool description. If indexing fails on a file, what should the agent do? If semantic_search returns zero results, should it retry with a different query or check index status first? Add recovery hints.
Tool annotations use destructiveHint=False for index_folder, but this tool DELETES chunks and modifies the database. This is a destructive operation. Should be destructiveHint=True.
file_type parameter in semantic_search accepts a string (e.g. '.pdf') with no validation. No enum, regex pattern, or length constraint visible. LLMs may pass invalid extensions like '.xyz' without feedback.