A conversational agent powered by RAG to answer client-specific questions by performing semantic search on a vector database using a MCP server with hybrid search capabilities
Single tool 'search_documents' with good naming and clear description, but critical gaps in schema validation and error handling. Description is well-written (195 chars, within baseline 194 avg) and explains the hybrid search mechanism clearly. However, input schema is present but lacks explicit validation constraints (top_k has no min/max bounds). Parameter descriptions are adequate but sparse. No documented output schema structure beyond 'list[dict]'. Error handling is absent, no guidance for LLM on recovery paths. Tool is read-only (good), but lacks any error classification or validation feedback mechanisms. Overall follows basic patterns but misses production-grade composition and error resilience.
Perform hybrid search on indexed documents using vector and text search. This tool searches through the document collection using hybrid search that combines semantic similarity (vector embeddings) and keyword matching (text search) using Reciprocal Rank Fusion (RRF). This provides the best of both worlds: conceptual understanding from embeddings and precise keyword matching from text search. The results can be used to ground the user's answer with the most relevant documents.
Input schema lacks validation constraints: 'top_k' parameter has default=3 but no minimum (e.g., 1) or maximum bounds (e.g., 100). Unbounded integers allow LLMs to pass absurd values that could overwhelm the MongoDB query or exhaust memory.
Output schema not documented in code. Function signature shows 'list[dict]' but never specifies field names, types, or structure. LLMs cannot reliably extract data (e.g., document_id, content, score) without seeing the actual schema. Baseline requires 100% of A+ tools to document return types.
No error handling or recovery guidance. If MongoDB is unreachable, embedding service fails, or query times out, the function returns an unhandled exception. No error categorization (retryable vs. fatal) or actionable guidance for the LLM. Baseline pattern requires: 'User not found. Try search_users() with a partial name.'
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Parameter 'query' lacks format/length constraints in description. Description does not specify: minimum/maximum length, language assumptions, special character handling, or example valid queries. Leaves LLM guessing about edge cases (empty string, SQL injection patterns, very long strings).
No idempotency guarantee documented. If the tool is called twice with identical parameters, will it return the same results, or could caching/indexing changes cause variance? Agents retry on ambiguous failures, unclear idempotency semantics risk duplicate side effects if combined with multi-attempt strategies.
Result limit not enforced or documented. Tool accepts top_k but does not state hard cap (e.g., 'max 50 results'). Baseline pattern: 'Even if the API allows returning thousands, cap results at 20-50 and offer pagination.' Returning unbounded results risks context window exhaustion.