An agentic RAG (Retrieval-Augmented Generation) server that provides COVID-19 FAQ retrieval from a vector database and web search capabilities via Firecrawl API
Two tools are registered with FastMCP and have descriptions visible in source code. However, both tools exhibit significant quality gaps: descriptions lack critical context about when to use them vs. alternatives; input parameters have minimal documentation; output formats are poorly specified; no error handling guidance is provided; and no defensive validation is shown. The naming follows verb_noun convention (covid_faq_retrieval_tool, firecrawl_web_search_tool), which is acceptable, but the descriptions are vague about resource selection logic and prerequisites. Parameter schemas are minimal (only 'query' string with basic description). No output schema is documented, the first tool returns 'str' and the second returns 'List[str]', but the actual structure of returned content is opaque to an LLM. Error handling is present in the code (type checks, logging) but does not guide the LLM on recovery ('No embeddings matched' returns a fallback message, but there is no structured error response). This is a below-average server typical of early-stage RAG tooling.
Retrieve the most relevant documents from the Covid FAQ collection. Use this tool when the user asks about covid related questions. OR Says `I want the covid FAQ referred`
Search for information on a given topic using Firecrawl. Use this tool when the user asks a specific question not related to the Covid.
Output schemas are not documented. covid_faq_retrieval_tool returns 'str' and firecrawl_web_search_tool returns 'List[str]', but the actual structure, field names, and content format are completely opaque. An LLM cannot plan downstream tool calls or extract structured data without knowing the response shape.
Tool descriptions lack WHEN-to-use context and selection criteria. The decision logic ('Use this tool when the user asks about covid related questions' vs. 'when the user asks a specific question not related to the Covid') is heuristic and ambiguous. What if a user asks 'Is COVID still a problem in 2025?', which tool should the LLM select? Descriptions do not clarify overlapping use cases or disambiguation.
Parameter descriptions are minimal. 'query' in both tools is described as 'The user query to retrieve/search for information', this is nearly tautological and offers no guidance on format constraints, length limits, language, or expected structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
No input validation or error handling guidance for LLMs. Both tools check isinstance(query, str) and raise TypeError, but this is not documented in the tool description. An LLM has no way to know what inputs are invalid before calling. Error messages like 'I couldn't find a relevant answer in my knowledge base' for the COVID tool are user-friendly but not structured, making it hard for the LLM to distinguish between 'no results' and 'service error' for retry logic.
No result limits or pagination. firecrawl_web_search_tool returns 'List[str]' with no upper bound. If the web search returns 1000 results, the LLM gets all of them, which will blow the context window and waste tokens.
Prerequisites and dependencies are undocumented. covid_faq_retrieval_tool depends on a Qdrant vector database, HuggingFace embeddings, and a collection named 'covid-faq' being pre-populated. None of this is mentioned in the tool description. If Qdrant is unavailable, the LLM will see a cryptic error. Descriptions should state: 'This tool queries a pre-indexed COVID FAQ collection via Qdrant vector search. Requires network access to Qdrant at [URL].'