Model Context Protocol server exposing RAG (Retrieval-Augmented Generation) capabilities, document indexing, knowledge graph querying, calendar/task management, filesystem operations, code analysis, and web search through MCP tools
Shodh-RAG presents a moderately functional tool set with decent naming conventions and basic schema structure, but suffers from inconsistent description quality, incomplete parameter documentation, and missing output schema specifications. The server has 16 tools covering RAG operations, knowledge graphs, file operations, and task management. Naming generally follows verb_noun patterns (rag_search, knowledge_graph_query, read_file, create_task), which is positive. However, critical gaps emerge: descriptions vary in depth and helpfulness; many parameter descriptions are minimal or missing (e.g., 'filters' in rag_search is described only as 'Optional metadata filters' without clarifying what specific filter keys are accepted); output schemas are not documented in the source code, forcing LLMs to guess the response structure; several tools lack clear guidance on when to use them vs. similar tools (e.g., rag_search vs. search_documents vs. hybrid_search). Error handling and recovery guidance are absent from all tool definitions. The codebase shows proper use of JSON Schema for input validation (type, required, properties, enum fields), but descriptions do not leverage this to guide LLM reasoning. Duplicate tool names (read_file appears twice: #10 and #16 with nearly identical descriptions) suggest poor tool composition and increase LLM confusion.
Ask questions about code in the indexed codebase with full context
Create a task or reminder for the user. Use when the user asks to remember something, schedule a follow-up, or when you find an actionable deadline in a document. The task appears in the user's Tasks tab.
Retrieve all chunks from a specific document by its doc_id. Use this after search_documents when you need the full content of a document, not just the matching chunk. Provide the doc_id from a search result.
Get RAG system statistics including document count, index size, etc.
Perform hybrid search combining local RAG and web search
Index a new document into the RAG system
Query the knowledge graph for entities, relationships, and context
Duplicate tool 'read_file' defined twice (tools #10 and #16) with nearly identical descriptions. LLMs cannot distinguish between them and will randomly pick one, causing unpredictable behavior. Violates single-responsibility and composition patterns.
Output schemas not documented for any tools. The source code shows input schemas (type, properties, required) but no response schema specifications. LLMs cannot infer the structure of results and must guess what fields are returned (e.g., does rag_search return 'relevance_score' or 'score'? Is it 'document_id' or 'doc_id'?). This violates the 'WHAT does it return' requirement of pattern:tool-description.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
List all indexed documents in a space
List all indexed document sources in the knowledge base. Returns source names, document counts, and basic statistics. Use this to understand what documents are available before searching.
Search through indexed documents using semantic search, hybrid search, or keyword search
Read the contents of a file from the filesystem. Requires user permission.
Read the full content of an indexed file
Retrieve relevant memories from the memory system based on query or time range
Search the user's indexed documents using hybrid semantic + keyword search. Returns relevant document chunks with scores, sources, and citations. Use this whenever the user asks about information in their documents.
Generate an interactive call graph visualization showing how functions interact in the codebase. Use this when users ask to visualize, trace, or understand execution flow, architecture, or function relationships.
Search the web using DuckDuckGo and return results
Parameter descriptions are sparse and lack actionable detail. Examples: 'filters' in rag_search is described only as 'Optional metadata filters' without listing valid keys (source_type, author, date_start, date_end are properties but not explained in description). 'max_results' has no guidance on valid range or consequences of large values. 'format' in visualize_call_graph mentions 'interactive' and 'ascii' but not what each produces or when to use each.
Tool composition confusion: three highly similar search tools exist without clear guidance on when to use each. rag_search (indexes semantic/hybrid/keyword), search_documents (hybrid semantic+keyword), and hybrid_search (combines local RAG + web) sound like they do different things but descriptions do not clarify the distinction. LLMs will waste reasoning cycles and may pick the wrong one. violates pattern:tool-chain principle of clarity.
No error handling or recovery guidance in any tool description. Pattern:recovery-guide and pattern:error-classification mandate that tool descriptions tell LLMs what to do on failure (retry? ask user? fatal?). For example, if web_search times out, is that retryable? If index_document fails due to file not found, should the LLM try a different path or ask the user? No guidance provided.
Missing permission or scope declarations. Tools like index_document (WRITE) and create_task (WRITE) perform state changes but include no description of what permissions are required. Pattern:scope-declaration and pattern:permission-gate require this for audit and least-privilege configuration.
Parameter 'space_id' appears in multiple tools but descriptions do not explain what a space is, how to find a space_id, or what happens if an invalid space_id is passed. LLMs will guess or pass placeholder values. Should be consistently defined or split into discovery tools (list_spaces) and then space_id references.
get_statistics has an empty input schema (no properties), but the description does not clarify whether any parameters are accepted or if calling it always returns system-wide stats. Also, description is under 20 chars: 'Get RAG system statistics including document count, index size, etc.' This is borderline (exactly 73 chars after 'Get RAG system...') but does not explain WHEN to call it or what to do with the results.
Tools lack pagination or result limit guidance. rag_search, knowledge_graph_query, list_documents, and others accept max_results but descriptions do not warn about context window explosion if the default is too high or the user passes a large value. Pattern:paginated-result requires this.