A RAG (Retrieval-Augmented Generation) assistant that combines local vector database search with real-time web search via MCP, using LangChain and FastAPI
CRITICAL: Tool names lack verb clarity and contain duplicates. 'query' appears at least twice with different signatures; 'health_check' and 'get_health_status' are near-duplicates.
Tool definitions are not explicitly visible in source code. Most tools are inferred from method signatures in example files (cli-demo.py, advanced_usage.py) and placeholder endpoints (FastAPI server.py).
Input schemas lack constraints and format guidance. Parameters like 'max_results' have type and default but no min/max bounds; 'documents' in add_documents has no item schema; 'force_web_search' in query has no enum or explanation. LLMs cannot read JSON Schema pattern fields, they rely on description text.'
Recommendations
IMMEDIATE: Consolidate near-duplicate tools. Merge 'health_check' and 'get_health_status' into one canonical 'get_health_status'. Distinguish 'query', 'search', and 'rag_search' with explicit descriptions of when to use each (e.g., 'query' orchestrates the decision, 'search' does web-only, 'rag_search' does vector-db-only).
CRITICAL: Make tool definitions explicit in code. Add explicit MCP tool registration (e.g., using an MCP SDK decorator or registration function) with full JSON Schema for inputs and outputs. The current code mixes FastAPI endpoints, example usage, and method signatures, none of which constitute a formal MCP tool definition.
CRITICAL: Document output schemas for all tools. For 'search' and 'rag_search', specify: { type: 'array', items: { type: 'object', properties: { title, url, snippet, source, relevance_score }, required: ['title', 'url'] } }. For 'add_documents', return { type: 'object', properties: { added_count, failed_count, errors }, required: ['added_count'] }.
Add constraints to all numeric parameters: 'max_results' should be min=1, max=100 with description 'Maximum number of results to return (default 5, max 100 to avoid context explosion)'. 'confidence_threshold' should be min=0.0, max=1.0.
Example for 'add_documents': 'Add new documents to the RAG vector store. Documents are parsed (PDF, TXT, DOCX), embedded using sentence-transformers, and stored in FAISS. Returns count of successfully added documents and any parse errors. Supports batch operations.'
Output schemas are not documented for any tool. Rubric requires: 'Document the output schema. LLMs need to know what fields to expect so they can plan downstream tool calls and extract the right data.' Without documented returns, LLMs cannot chain tools or validate responses.
Descriptions are too brief and lack actionable guidance. State WHAT the tool does, WHEN to use it, and any prerequisites.' Examples: 'Refresh the vector database' (30 chars, no explanation of what happens), 'FastAPI health check endpoint' (36 chars, generic), 'Format search results for the client' (44 chars, no format details). Baseline: average A+ tool description is 194 chars; these average ~50 chars.
WRITE tools (add_documents, update_vector_store) lack clear state-change descriptions and error guidance. Agents need to know which calls are safe to retry and which have irreversible consequences.' add_documents doesn't explain: Does it replace or merge? Deduplicate on content hash? 'Refresh' is ambiguous, does it clear and rebuild or incrementally update?
No error handling guidance in tool definitions. Rubric requires: 'Error responses must tell the LLM what to do next' and 'Categorize errors as retryable, user-fixable, or fatal.' No visible error recovery paths, retry logic, or actionable error messages in tool specs. If add_documents fails to parse a file, what should the LLM do? Retry? Ask user? Skip that file?
Composition issues: Multiple tools with overlapping semantics ('search', 'rag_search', 'query') suggest unclear responsibility boundaries. Rubric: 'Each tool should do exactly one thing.' When should LLM call 'search' vs 'rag_search' vs 'query'? No clear distinction in descriptions. Additionally, 'process_query' and 'format_results' appear to be internal utilities that should not be exposed as MCP tools.
Parameter naming inconsistencies: 'query' vs 'query_text' vs 'query_string' used interchangeably across tools. Rubric: 'When a parameter could be an ID, name, email, or position, suffix it with the type.' This inconsistency forces LLM to reason about field mappings, increasing errors.
Dangerous boolean defaults in query tool: 'force_web_search' defaults to false, 'include_sources' defaults to true. Rubric: 'Defaults must not cause data loss or unintended side effects.' If an LLM omits 'force_web_search', does it silently skip RAG? If it omits 'include_sources', does it suppress attribution? These need explicit guidance.
query
Add explicit error handling specs: 'If document parsing fails, return { added_count: N, failed_count: M, errors: [{ file: 'x.pdf', reason: 'unsupported format' }] }. Recoverable errors (e.g., network timeout on embedding service) should include retry_after guidance.'
Rename 'process_query' and 'format_results' to internal-only functions (prefix with '_' or move to utils module). These are implementation details, not user-facing tools. If they must remain exposed, document their purpose and when agents should call them.
Add parameter dependency documentation: In 'query' tool, clarify: 'If force_web_search=true, the RAG database is bypassed regardless of confidence_threshold. If force_web_search=false, the system queries RAG first; if confidence < confidence_threshold (default 0.7), it falls back to web search.'
Implement pagination for 'search' and 'rag_search': Add 'offset' (default 0) and 'limit' (default 5, max 50) parameters. Return { results: [...], total_count: N, has_more: bool, next_offset: M }.
Add missing parameter descriptions. For example, 'documents' in add_documents should be: 'Array of file paths (absolute or relative to working directory). Supported formats: .pdf, .txt, .docx, .md. Each file processed independently; parsing errors do not block other files.'
Establish idempotency semantics: 'add_documents' should be idempotent (calling twice with same file paths does not duplicate content). Document how deduplication works (content hash? file path?). 'update_vector_store' should document if it's a full rebuild (lossy) or incremental (safe to retry).
Add security guidance in tool descriptions: 'add_documents does not accept file paths outside the configured document directory (prevents path traversal). All file I/O is sandboxed.' Similar for any tools accepting user input.