An agentic RAG server with MCP that provides tools for entity extraction, query refinement, and relevance checking for Sri Lankan news articles using OpenAI
This MCP server exhibits significant definition quality issues across all four tools. While tool names follow verb_noun convention (get_, extract_, refine_, check_), descriptions are generic and lack actionable context for LLM selection. Input schemas are present but parameter descriptions are minimal (1-2 words). Output schemas are entirely undocumented, LLMs cannot infer return types or structure. Error handling is absent; exceptions return JSON error strings without recovery guidance. The server conflates the MCP server role (server.py using FastMCP) with client-side tool invocation logic (mcp-client.py), creating confusion about tool ownership and where validation occurs. Most critically, three tools (extract_entities_tool, refine_query_tool, check_relevance) are defined in mcp-client.py, which is a CLIENT file, not the server, suggesting incomplete tooling exposure or code organization issues.
Check the relevance of a text chunk to a given question using an LLM. Returns a relevance score between 0 and 1.
Extract entities from a given text query using OpenAI.
Get the current date and time.
Refine a given text query using OpenAI.
Output schemas completely undocumented. Tools return strings, floats, or JSON error objects with no schema definition visible. LLMs cannot determine field names, types, or structure for downstream tool chaining.
Parameter descriptions are minimal (1-2 words max). 'query' described as 'The text query to extract entities from' lacks context on format, length, language, or when to call this vs refine_query_tool.
Tool descriptions lack 'when to use' context and do not disambiguate between similar tools. extract_entities_tool and refine_query_tool both operate on text queries but serve different purposes, LLMs cannot reliably choose between them based on 50-char descriptions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 51 | - | v1 |
Error handling is silent or returns JSON error strings without recovery guidance. check_relevance returns {'error': str(e)} on exception, LLM cannot determine if error is retryable, user-fixable, or fatal. No actionable next steps provided.
Naming inconsistency: get_time_with_prefix is unclear, 'prefix' is never explained in description or used in implementation (just returns datetime string). Ambiguous name invites misuse.
Tool definitions split across server.py (get_time_with_prefix) and mcp-client.py (other three tools). Per MCP spec, tools should be registered on the server, not in client code. Code organization suggests tools may not be fully exposed via the MCP server interface.
No parameter validation or constraint documentation. check_relevance expects 'question' and 'text_chunk' strings but does not document max length, encoding, or language. Truncation to 1000 chars is hardcoded with no explanation to LLM.
LLM API keys (OPENAI_API_KEY) are referenced via os.getenv() in tool implementations. While not exposed as parameters (correct), the code does not validate key presence at startup. Missing keys will fail silently at runtime.