VecGrep presents a mixed quality server with well-defined core semantic search functionality but significant gaps in error handling, output schemas, and security considerations. All 7 tools have basic descriptions and input schemas with type information, placing it above the median (~45-55). However, descriptions lack LLM-optimized guidance for when/why to call tools, output schemas are undocumented, and error recovery paths are absent. The semantic search domain is niche but well-served by clear naming (index, search, inspect, graph_query). STDIO transport caps protocol readiness at 50. Key strength: input schemas are present and typed. Key weakness: no documented output structure, no error handling guidance, no parameter constraints (enums), and missing descriptions for several parameters.
Clear the index for a project. Removes all indexed chunks and resets the embedding metadata.
Query the semantic knowledge graph for structural relationships between code entities (functions, classes, imports).
Index a directory or project for semantic search. Creates vector embeddings of all code files.
Inspect the current index for a project. Returns metadata about indexed files, chunks, and embedding provider.
Search indexed code using semantic vector similarity. Returns chunks ranked by relevance.
Stop watching a project directory for changes.
Start watching a project directory for file changes and automatically re-index when files are modified.
No output schemas documented for any tool. LLMs cannot know what fields to expect (e.g., does search return 'relevance_score' or 'score'? is 'chunk_text' or 'text'?). This forces LLMs to parse unstructured responses and blocks downstream tool chaining.
search() and inspect() parameters lack type information visibility in provided schema (limit is integer but no min/max bounds; query is string but no length constraint documented). Results could be unbounded, causing context window exhaustion.
index() has a 'watch' boolean parameter but separate 'watch' and 'unwatch' tools also exist. This creates ambiguity, LLM must decide between index(watch=true) vs calling watch() separately. Tool composition is muddled.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 34 | - | v1 |
No error handling guidance. What does index() return if the path is invalid? If embedding fails? If storage is full? No recovery instructions, no error categorization (retryable vs fatal), no actionable messages.
clear_index() is a destructive operation but lacks confirmation/dry-run support. Agents can permanently delete indexed data without review. No idempotence hint or confirmation request pattern.
provider parameter in index() accepts enum ['local', 'openai', 'voyage', 'gemini'] but no guidance on API key injection, cost, or when each is appropriate. API keys must not be parameters, server-side injection is required.
No audit logging in tool definitions. Code has logging module imported but no per-call tracing visible (who indexed what, when, what results). Compliance and debugging require audit trails.
watch() and unwatch() parameters lack description text stating what 'watching' does, frequency of checks, disk/CPU impact, or how to configure filters. 'Path to watch' is self-evident but the WHEN/WHY/HOW for LLM is missing.
No documentation of what data search() returns when limit is hit, is the full ranked list truncated, or only top-N? Is there a cursor for pagination? Default limit=10 is not stated anywhere.
graph_query() parameter 'query' is described as 'Natural language query about code structure' but no guidance on expected format, example queries, or what 'structural relationships' means. LLM must guess what this tool accepts.