An MCP server that provides Retrieval-Augmented Generation (RAG) capabilities using a vector store (Chroma) and Google Gemini, with web search fallback via Serper API.
The server defines 4 tools with basic descriptions and minimal input schemas. All tools accept a single 'query' string parameter. Descriptions are present but generic and lack specificity about when/why to call each tool relative to others. No output schemas are documented. Input validation exists (isinstance checks) but error messages are minimal. Tool names follow verb conventions but lack clarity about interdependencies.
Answer a question using the RAG system with Gemini. Use this tool when you want to generate an answer based on vector store content. Input: query: str -> The user question Output: answer: str -> Generated answer from Gemini
Run the complete RAG pipeline: 1. Try to answer from the vector store 2. Fall back to web search if needed 3. Generate a response using Gemini Input: query: str -> The user query Output: answer: str -> Generated answer
Search the vector store for relevant documents. Use this tool when the user asks about machine learning topics. Input: query: str -> The user query to search for information Output: context: str -> Relevant documents from the vector store
Search the web using Serper API. Use this tool when the user asks about topics not covered in the vector store. Input: query: str -> The user query to search for information Output: context: str -> Web search results
No output schemas documented. LLMs cannot predict what fields to expect from tool responses (e.g., does search_vector_store return 'context' as a string or object? Is there relevance scoring? Page info?). This forces LLMs to guess structure and wastes tokens parsing responses.
Ambiguous tool selection guidance. Docstrings say 'Use this tool when the user asks about machine learning topics' for search_vector_store, but no clear distinction from rag_pipeline which also searches the vector store first. LLMs will struggle to choose between them.
Generic parameter descriptions. 'The user query to search for information' is repeated across all tools and doesn't clarify constraints: Is the query limited to keywords? Full sentences? Max length? Does RAG support boolean operators? These details matter for LLM planning.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
No error handling guidance. Input validation raises ValueError for non-string queries, but the response to an LLM is just a Python exception. There's no recovery hint or suggested next step. Pattern: 'Query must be a string' doesn't tell the LLM what to do.
Tool composition confusion. answer_question calls rag_system.answer_question() which may internally call the vector store. rag_pipeline also 'tries to answer from the vector store [then] falls back to web search'. The boundary between these tools is unclear, when should an agent call answer_question vs rag_pipeline?
No pagination or result limits documented. search_web and search_vector_store return context as a single string with no mention of limits or pagination. If the vector store has 1000 relevant documents, does this return all of them? That could blow the context window.