LLM Powered Advanced RAG Application with MCP support for semantic search, document embedding, and LLM-based question answering
The pyLLMSearch server has 4 tools with basic functionality but significant definition quality gaps. All tools have names starting with verbs (rag_*) and descriptions are present. However, descriptions are generic and lack actionable guidance for LLM selection. Input schemas are minimal, most parameters are undescribed (missing parameter descriptions). Output schemas are not documented. Three tools are nearly identical in naming and purpose (rag_retrieve_chunks, rag_generate_answer, rag_generate_answer_simple), creating ambiguity for LLM tool selection. The tool names use prefixes that don't clearly distinguish their roles from each other. No error handling guidance, no parameter constraints, and no indication of prerequisites or when to use one tool vs another.
Retrieves answer to the question from the embedded documents, using semantic search.
Retrieves answer to the question from the embedded documents, using semantic search.
Retrieves chunks of information relevant to the question from the embedded documents, using semantic search.
Updates the index with the latest documents.
Three nearly identical tools (rag_retrieve_chunks, rag_generate_answer, rag_generate_answer_simple) with unclear differentiation. Names do not disambiguate purpose. LLMs will struggle to choose between 'generate_answer' and 'generate_answer_simple'.
Parameter descriptions missing for all tools. 'question' parameter in rag_retrieve_chunks has description, but 'label' parameter in rag_generate_answer lacks context on what filtering is applied. rag_update_index has zero input parameters documented despite being a destructive operation.
No output schemas documented for any tool. LLMs cannot plan downstream operations or extract relevant fields. What structure does rag_retrieve_chunks return? What format are the 'chunks'? What fields are in the answer response?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Tool descriptions are generic and lack guidance on when to use each. rag_retrieve_chunks says it 'retrieves chunks...using semantic search' but doesn't explain: What is the difference from rag_generate_answer? When would an LLM choose chunks over a generated answer? What does 'chunks' mean in context?
rag_update_index is marked WRITE (destructive) but has no description of what 'updating' entails. Does it re-embed documents? Reload from disk? Add new documents? No error handling guidance or confirmation/dry-run support for irreversible operation.
Parameter 'label' in rag_generate_answer and rag_generate_answer_simple has description 'Optional label to filter documents' but no guidance on what valid labels are, format (enum, regex), or what happens if label doesn't match any documents.
No error handling guidance in any tool. What if no chunks are found? What if the index is corrupted? What if semantic search times out? Descriptions tell LLMs nothing about recovery paths.
Tool naming uses common prefix 'rag_' which reduces clarity. Names like 'semantic_search_chunks' vs 'semantic_search_answer' vs 'semantic_search_simple' would be marginally clearer, but the core problem is functional overlap, not naming convention.