Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
This server exhibits severe quality gaps across naming, descriptions, schemas, and error handling. While 11 tools are defined, most lack proper MCP integration patterns. Tool names are action-based but descriptions are minimal (10-60 chars, well below the 194-char baseline). Input schemas are present but extremely basic, each tool accepts only a single string parameter with no validation, constraints, or structured output. No output schemas are documented. Error handling is generic ('Error al procesar la solicitud'). The server appears to be a FastAPI/Express backend exposing LLM inference endpoints, not an MCP server following the protocol spec. Tool definitions exist in source code but are not explicitly registered via MCP's tool management mechanism.
Tools (11)
chatread onlyauthsource verified38/100
Chat endpoint using OpenAI GPT-3.5-turbo model
postChatread onlyauth43/100
Chat completion endpoint using OpenRouter API with configurable model selection
query_gemmaread onlyauthsource verified40/100
Query endpoint for google/gemma-2b-it model
query_hermesread onlyauthsource verified40/100
Query endpoint for OpenHermes-2.5-Mistral-7B model
query_llamaread onlyauthsource verified40/100
Query endpoint for Meta Llama 3 (8B Instruct) model
No output schemas documented. Tools return raw API responses without defining expected fields, types, or structure. LLMs cannot infer what to extract or how to chain results.
Tool descriptions are too short (10 - 60 characters, baseline is 194 chars). Descriptions lack context on WHEN to use each tool, WHAT it returns, or HOW it differs from the 10 other query tools. LLMs cannot disambiguate between query_qwen, query_gemma, query_hermes, etc.
Add comprehensive output schemas to every tool. Document the expected response structure (fields, types, sample values). Example: query_qwen should document that it returns {"response": string, "tokens_used": number, "model": string}.
Expand tool descriptions to 150 - 250 characters. Explain: (1) What this tool does vs. the 10 other query_* tools, (2) When an LLM should prefer this tool (e.g., 'Use for lightweight, fast inference on 1.8B-parameter model'), (3) What models or APIs it uses.
Consolidate or document the purpose of 10 nearly-identical query_* tools. If they all query different HuggingFace models, create ONE 'query_huggingface_model' tool that accepts a 'model_name' enum parameter (qwen, gemma, hermes, llama, mpt, phi, phi3, tinyllama). If they serve different use cases, document those use cases in each description.
Add constraints to input parameters. For 'query' string, specify max length (e.g., 2000 chars), allowed characters, and format (e.g., 'Natural language text, no JSON'). Use JSON Schema pattern, minLength, maxLength, enum fields.
Implement proper error categorization and recovery guidance. Instead of 'Error al procesar la solicitud: <exception>', return: {"error": "timeout", "retryable": true, "message": "HuggingFace API did not respond in 30s. Try again in 10s."} or {"error": "invalid_input", "retryable": false, "message": "Query too long (3500 chars). Max 2000. Shorten query and retry."}
Add pagination/limit parameters to tools returning potentially large results. E.g., postChat should accept 'max_tokens': 256 (documented, with min=1, max=4096). Return {"response": string, "tokens_used": number, "truncated": boolean}.
Score history
Overall score trend
↑ 0 points across a rubric change (v1 → v2)
30/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
30
2026-07-28+
v2
2026-03-09
F
30
-
v1
query_phi3read onlyauthsource verified40/100
Query endpoint for microsoft/Phi-3-mini-4k-instruct model
query_qwenread onlyauthsource verified40/100
Query endpoint for Qwen1.5-1.8B-Chat model
query_tinyllamaread onlyauthsource verified40/100
Query endpoint for TinyLlama-1.1B-Chat-v1.0 model
query_togetherread onlyauthsource verified40/100
Query endpoint for Together AI API (references Meta Llama 3 8B Instruct)
Input schemas are severely underconstrained. Single-string parameters with no length limits, patterns, enums, or validation rules. A 'query' parameter accepting any string invites hallucinated inputs that may cause API failures or timeout cascades.
No error recovery guidance. The Python code returns generic HTTPException with 'Error al procesar la solicitud: <str(e)>'. LLMs receive a raw error string with no hints: is this retryable? Can I fix it? Should I try a different tool? No categorization or actionable steps.
Tool proliferation without clear differentiation. 10 query_* tools exist (qwen, gemma, hermes, llama, mpt, phi, phi3, tinyllama, together, chat). No clear guidance on which to prefer or when to use each. LLMs will struggle with disambiguation.
No pagination or result limits documented. Tools may return unbounded responses (HuggingFace API, OpenRouter API). No mention of max tokens, result count, or pagination parameters in tool descriptions or schemas. Large responses risk context window exhaustion.
Tool definitions not explicitly registered via MCP protocol. Code shows FastAPI routes (@app.post('/query')) and Express endpoints, but no MCP-compliant tool registration (no CallToolRequest handling, no tool metadata response, no protocol handshake visible in provided code).
Parameter descriptions are trivial. 'User query string' and 'User chat message' provide no guidance on query length, format, content type, or expected behavior. The 'messages' parameter for postChat is described as 'Array of message objects' but no schema for message structure is provided.
Verify MCP protocol compliance. Ensure tools are registered via MCP's ListTools response with proper input schemas, descriptions, and output documentation. The current code shows HTTP routes but no MCP-compatible tool serialization. See https://spec.modelcontextprotocol.io/latest/basic/tools/
Add tool annotations (readOnlyHint, idempotentHint) to document that query_* tools are read-only and safe to retry. This helps LLMs reason about safety and retry strategies.
Document HuggingFace API key injection. Confirm that HUGGINGFACE_API_KEY is loaded from environment (.env) and never exposed in logs, responses, or error messages. Add a note: 'API credentials are managed server-side via environment variables; no API keys are exposed to the client.'
Add examples of valid and invalid inputs to parameter descriptions. E.g., 'query' parameter: 'Valid: "What is Python?", "Explain photosynthesis". Invalid: JSON payloads, code snippets, more than 2000 characters.'