An agentic RAG (Retrieval-Augmented Generation) server for processing queries about Sri Lankan news using entity extraction, query refinement, and relevance checking with OpenAI integration.
This server exhibits major definition quality gaps. Of 4 tools, only 1 has a properly visible input schema in server.py (get_time_with_prefix). Three tools (extract_entities_tool, refine_query_tool, check_relevance) are defined in mcp-client.py, which is NOT a server definition file, these are client-side invocations, not MCP tool registrations. The source shows tools being called as Python functions with OpenAI, but no explicit MCP tool registration with @mcp.tool() decorators for the three LLM-based tools is visible in server.py. This violates the fundamental requirement that tool definitions must be visible in the server codebase. Descriptions are present but generic (e.g., 'Extract entities from a given text query using OpenAI' does not explain WHEN to use this vs refine_query, what the output format is, or what categories of entities are expected). Input schemas are missing for 3 of 4 tools in the server definition. Error handling is minimal, tools return JSON error objects inconsistently, and there is no guidance to the LLM on recovery steps.
Check the relevance of a text chunk to a given question using an LLM. Returns a relevance score between 0 and 1.
Extract entities from a given text query using OpenAI.
Get the current date and time.
Refine a given text query using OpenAI.
Three tools (extract_entities_tool, refine_query_tool, check_relevance) defined in mcp-client.py, not server.py. This is a client file, not a server tool registration. No @mcp.tool() decorators visible for these tools in the server codebase.
Input schemas completely missing from server.py for 3 of 4 tools. Only get_time_with_prefix has an empty input schema {}. extract_entities_tool, refine_query_tool, and check_relevance have no visible schema definitions in the server code, they appear only as client-side function calls.
Output schemas not documented. Tools return raw strings or floats (e.g., check_relevance returns float, extract_entities_tool returns string containing JSON). LLM cannot infer expected output structure, making it impossible to chain tools or parse results reliably.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 35 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Descriptions are generic and lack context for LLM tool selection. 'Extract entities from a given text query using OpenAI' does not explain: (1) What entity types are supported? (2) What is the output format? (3) When should this be called vs refine_query_tool? (4) What are the prerequisites? Description is 62 characters, below the 10 - 1024 character guideline for effective LLM prompt engineering.
Parameter descriptions missing entirely. For extract_entities_tool, the 'query' parameter has description 'The text query to extract entities from', this is minimal and does not explain expected input format, length constraints, or language requirements. Similar gaps for refine_query_tool and check_relevance parameters.
Error handling is inconsistent and provides no recovery guidance. extract_entities_tool and refine_query_tool return JSON {"error": str(e)}, but check_relevance returns {"error": str(e)} as well, mixing structured JSON errors with float returns. No error messages guide the LLM on next steps (e.g., 'OpenAI API key missing. Check environment variables.' or 'Rate limit hit. Retry in 60 seconds.').
Tool names are ambiguous and overlap in responsibility. extract_entities_tool and refine_query_tool both pre-process a query using OpenAI. The distinction between 'extract entities' and 'refine query' is unclear, an LLM may struggle to choose between them. Consider renaming to clarify: query_extract_entities and query_refine_for_search.
No parameter type validation or constraints documented. check_relevance accepts 'question' and 'text_chunk' as strings with no size limits, yet truncates text_chunk to 1000 characters internally. This limit should be documented in the parameter description to set LLM expectations.
Secrets (OPENAI_API_KEY, OPENAI_MODEL_NAME) used via environment variables, correct pattern. However, no audit logging visible. Tools call OpenAI but do not log who invoked them, when, or with what parameters. Compliance and incident response require traceability.