A server for interacting with the Retrieval-Augmented Generation project
Single tool with minimal schema, generic description, and critical gaps in parameter validation and error handling. The tool exists and is registered with fastmcp, but lacks the rigor expected of production-grade RAG adapters. Input schema is present but bare-bones. No output schema documentation. No error guidance. No parameter constraints or validation hints.
Function to search documents by query in vector database and return result with Retrieval-Augmented Generation
Output schema not documented. Tool returns `str` (result.json()['text']), but LLM has no specification of what fields to expect, what the text represents, or how to chain downstream operations.
Parameter description is generic and lacks actionable constraints. 'query to search documents' does not explain: format expectations, max length, special characters, required pre-processing, or what constitutes a valid query.
No error handling or recovery guidance. Tool calls `requests.get()` with a 20-second timeout and parses `result.json()['text']` without validation. If RAG_URL returns error status, malformed JSON, or missing 'text' field, the tool crashes with a cryptic exception. No guidance for LLM on what to do next (retry, ask user, try alternative).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Hardcoded dependency on environment variable RAG_URL. No documentation of what this service expects, its latency, failure modes, or SLA. If the service is down or slow, the tool provides no fallback or graceful degradation. LLM has no visibility into these constraints.
No pagination support. Tool does not accept limit or offset parameters, nor does it document result size limits. If the underlying RAG service returns 500+ documents matching a broad query, the full response floods the LLM context, wasting tokens and degrading reasoning.
Input parameter lacks type information in description and no enum/constraint enforcement. 'query' is a string, but the parameter description does not specify: Is it a free-form natural language query? A structured query language? Case-sensitive? Max 1000 chars? Presence of these constraints is missing, inviting LLM to pass invalid or oversized queries.
Tool description does not answer key questions: When should LLM call this vs. other search tools? What does 'Retrieval-Augmented Generation' mean in this context, does the tool apply RAG logic, or just return raw search results? Does it filter results by relevance threshold? How recent is the vector database? These gaps force the LLM to infer behavior.