FastAPI-based semantic search server with cluster-aware caching, vector embeddings, and similarity-based document retrieval
The server exposes 3 tools with basic FastAPI endpoints. Tool names follow verb_noun convention (query_endpoint, get_stats, clear_cache), which is positive. However, definitions lack rigor. Descriptions exist but are minimal (90-120 chars), falling short of the 194-char production baseline. Input schemas are present and typed, but lack detail in parameter descriptions. Output schemas are entirely undocumented, the LLM cannot predict response structure. Error handling is absent, no recovery guidance for semantic cache misses or vector search failures. The semantic_cache.py and vector_store.py reveal implementation details (threshold=0.85, 'all-MiniLM-L6-v2' model) that should be documented in tool descriptions to inform LLM decision-making. Per-tool analysis shows moderate naming clarity but thin descriptions and missing output specs.
Clears all cached semantic search results and resets cache statistics
Returns cache statistics including hit count, miss count, and total entries
Performs cluster-aware semantic search on indexed documents with cache lookup and fallback to vector search
Output schemas completely undocumented. The LLM cannot know that query_endpoint returns {cache_hit, dominant_cluster, similarity_score, results} or that get_stats returns {hit_count, miss_count, total_entries}. This forces the LLM to guess field names and types, increasing downstream errors.
Parameter descriptions are minimal or absent. 'query' is documented as 'The search query text' (good), but 'top_k' lacks explanation of its effect (e.g., 'Controls the number of search results returned; higher values increase latency but improve recall. Range: 1-50.'). The LLM cannot reason about appropriate values.
No error recovery guidance. If semantic cache lookup fails (cluster not found, similarity below threshold), the tool silently falls back to vector search. The LLM cannot distinguish between 'cache hit' and 'cache miss' in the response structure and thus cannot learn when to expect which behavior. The description should explicitly state: 'Returns cache_hit=true if a similar cached query is found (similarity >= 0.85); otherwise, performs full vector search and returns cache_hit=false.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
get_stats has no input parameters documented (empty schema {}), yet the description does not explain what the returned stats represent. Baseline tools include sample stat structures in descriptions or return type hints. An LLM cannot infer that hit_count and miss_count track cache performance.
clear_cache is a destructive operation (WRITE risk) but has no confirmation or dry-run option. The description does not warn that this is irreversible. An agent mistakenly calling this would erase all cached results with no recovery path. Should include: 'Warning: This clears all cached search results and resets statistics. This action is irreversible. Consider calling get_stats() first to log current cache state.'
Tool descriptions do not document prerequisites or discovery flow. An LLM does not know whether to call get_stats before query_endpoint, or whether documents must be loaded first. The startup event initializes global state, but this is invisible to the MCP client. The description should hint: 'Call this after the system has loaded documents (typically automatic on server startup).'