A Model Context Protocol (MCP) server implementation for txtai with causal boost and multilingual support
The server provides 3 tools (search, qa, retrieve) with basic structure but significant quality gaps. All three tools have descriptions and input schemas visible in the code, but descriptions are generic and lack specificity for LLM decision-making. Parameters have type information but lack detailed constraints and format guidance. Output schemas are not documented. Error handling is minimal, no recovery guidance or categorization. The tools are read-only but lack explicit readonly annotations. Parameter descriptions are present but brief (averaging ~40 chars, below the 72-char baseline for production tools). Tool names follow verb_noun convention (search_*, retrieve_*) but qa is ambiguous, does it answer questions, retrieve Q&A pairs, or validate answers? No pagination support visible despite semantic search returning multiple results. No output structure documentation means LLMs cannot plan downstream chaining. STDIO-only transport is a hard cap at 50 for protocol readiness, but definition quality itself is mediocre across all three tools.
Question answering tool that retrieves relevant documents and generates answers
Retrieve documents from the knowledge base by ID or semantic match
Search the knowledge base using semantic similarity
Tool 'qa' has ambiguous naming. Does it answer questions, retrieve Q&A pairs, validate answers, or generate summaries? The name alone does not clearly convey the action. LLMs will struggle to distinguish it from 'search' and 'retrieve' for similar queries.
No output schemas documented for any tool. LLMs cannot infer what fields (id, score, metadata, etc.) the responses contain, forcing them to guess or make follow-up calls. This breaks tool composition and chaining.
Tool descriptions are generic and lack decision guidance. 'Search the knowledge base using semantic similarity' does not explain when to use search vs qa vs retrieve, what similarity metric is used, or what 'limit' defaults to. Descriptions should be 50-200 chars and answer WHAT, WHEN, and WHY.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 30 | - | v1 |
Parameter 'limit' lacks constraints. No minimum/maximum specified. Can an agent pass limit=1000000? Should it? Without bounds, LLMs may request excessive results, causing timeouts or memory issues. Add description like 'Maximum number of results (1-100, default 20)'.
No pagination support visible. search and retrieve accept 'limit' but no 'offset', 'cursor', or 'page' parameter. Results larger than limit cannot be iterated, agents get truncated results or must retry with smaller limits, wasting tokens and context.
No error handling guidance. Tools are documented as READ_ONLY but provide no error messages telling LLMs what to do if a query returns no results, an ID is invalid, or the knowledge base is unavailable. Agents need recovery paths.
Tool descriptions do not clarify expected input format or constraints. E.g., 'query' parameter, does it accept free text, keywords, boolean operators, entity names? 'question', is it a full sentence? Can it be multi-line? No format guidance forces LLMs to guess.
Tools lack readonly or idempotent annotations. While marked READ_ONLY in Risk, the tool schema itself (as shown) does not carry MCP toolAnnotations (readOnlyHint, idempotentHint). This is a protocol readiness gap but also a definition clarity issue, LLMs cannot see the intent in the schema.