MCP server for LLAMATOR with ARQ/Redis task broker
The server defines 2 tools with reasonable naming (verb_noun pattern) and basic descriptions, but has significant gaps in parameter documentation, schema completeness, and error handling guidance. Both tools have non-trivial input schemas, but parameter descriptions are minimal or missing, making it difficult for LLMs to understand what values are valid or required. The descriptions are functional but generic (143 and 61 chars respectively), lacking the LLM-optimized context that production tools provide. No output schema documentation is visible. Error handling provides no recovery guidance. The server relies on async job polling with timeouts but doesn't expose this pattern clearly to the LLM.
Create a LLAMATOR job and return the aggregated result after completion.
Return aggregated LLAMATOR results for a finished job.
Parameter descriptions are missing or severely incomplete. The 'req' parameter for create_llamator_run contains nested object properties (tested_model, run_config, plan) with only bare type declarations and no descriptions of what each nested field controls, valid value ranges, or required constraints. This forces LLMs to guess at valid inputs.
No output schema is documented for either tool. The code shows tools return structured responses (LlamatorRunToolResponse, aggregated results with per-model metrics), but the tool definitions do not declare what fields agents should expect, forcing LLMs to infer structure and risking incorrect downstream field access.
Error handling provides no recovery guidance. The code includes comprehensive error handling (timeouts, job failures, artifact retrieval errors), but none of this is surfaced to the LLM in the tool description or response. Agents have no way to know whether an error is retryable, whether they should check job status first, or what the next step should be.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). The create_llamator_run tool has Risk=WRITE (destructive operation that enqueues jobs and waits for completion), but this is not communicated to the MCP client via tool metadata, missing the idempotentHint pattern for job creation.
Nested object schemas lack field-level documentation. The 'tested_model' parameter is typed as an object but contains OpenAI-compatible client config fields (base_url, model, temperature, system_prompts, etc.) that are not described. LLMs cannot determine which fields are required, what formats are accepted, or what constraints apply (e.g. temperature range 0-2 for OpenAI).
get_llamator_run description is vague (61 chars): 'Return aggregated LLAMATOR results for a finished job.' Does not explain when to call it (after job completion?), what fields to expect, or how to detect job progress. Agents may call it prematurely or misinterpret the response format.