Stop shipping agents on vibes. Score every agent output for quality, safety, and cost. MCP server with HTTP ingest and dashboard.
Iris demonstrates solid definition quality with clear, action-oriented tool names and comprehensive descriptions. All 12 tools follow verb_noun naming conventions (log_trace, evaluate_output, get_traces, etc.). Descriptions are detailed and context-rich, averaging ~150-180 characters, well within the 10-1024 character baseline. However, critical gaps exist: no visible input schemas or parameter definitions in the source code provided. The tools are listed in website/public/.well-known/mcp.json but the actual JSON Schema payloads are not shown, preventing verification of parameter types, constraints, and output structures. Tool definitions appear to be self-documenting via description text rather than formal schema, which violates the pattern:tool requirement for structured parameter validation. Composition is excellent, each tool has a single, clear responsibility (log trace, evaluate output, compare runs, etc.) with no conflation of concerns. Error handling guidance is implied in descriptions but not explicitly stated. Security considerations around API keys and authentication are not visible in the provided source.
Did this change make the agent worse? Compares two runs of stored evaluations with an interval — or says the data cannot tell, or that the runs are equivalent within a margin.
How reliably does the agent answer the same question? Groups stored evaluations by case and reports per-case pass rates, flakiness, and an interval that respects repeats.
Remove a deployed custom rule — or, with enabled, disable or re-enable it without removing it — effective on the next evaluate_output call.
Remove one stored trace by id; its spans go with it, and every evaluation linked to it keeps its verdict and loses its text.
Deploy a custom rule that fires on every future evaluate_output call of its bundle — persisted, active immediately, audited.
No input schemas visible in source code. The .mcp.json configuration files and sample Dockerfile show tool names and risk levels, but the actual JSON Schema definitions for input parameters, parameter types, constraints (enums, min/max), and output structures are not provided.
Parameter documentation missing. Descriptions explain what each tool does at a high level, but there is no evidence of per-parameter descriptions explaining what each input parameter controls, what format it expects, or what values are valid. LLMs cannot infer parameter meaning from names alone (e.g., does 'interval' mean seconds, milliseconds, or a time range?).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2026-07-28+ | v2 |
Score an agent output against the deterministic rule bundles: the ship verdict and its basis, every rule result with evidence and uncertainty, and what was not judged.
Re-score every trace in a run under the current rules, into a new run — so a rules change can be compared against the old verdicts instead of overwriting them.
Score an output with an LLM judge on your own provider key: a 0..1 score, a rationale, sub-scores and the spend.
Query stored traces with filters, pagination and sorting; optionally include the dashboard summary in the same response.
The rule inventory: the built-in roster with what each rule needs, the criticality this server applies and its published accuracy, plus every deployed custom rule.
Store one agent execution — input, output, tool calls, spans, cost, latency, token usage — and get the trace_id every later call keys on.
Extract the citations in an output, fetch the sources (opt-in, SSRF-guarded) and ask an LLM judge on your key whether each source supports its claim.
No explicit output schema documentation. Tool descriptions state what they return in prose (e.g., 'gets the trace_id'), but there is no formal JSON Schema defining the response structure, field types, pagination parameters, or how chaining IDs are provided. Response shaping is not verified.
Error handling and recovery guidance not explicit. Descriptions do not state what errors can occur, whether they are retryable, or what the agent should do next. E.g., 'evaluate_output' might fail if the rule bundle is missing, is that retryable? Should the user be asked? No guidance provided.
Destructive operations lack confirmation or dry-run pattern. 'delete_rule' and 'delete_trace' are marked DESTRUCTIVE in risk metadata, but descriptions do not mention dry-run, confirmation steps, or safeguards against accidental deletion.
Pagination and result limits not documented. 'get_traces' and 'list_rules' likely return many items, but there is no mention of pagination (limit, offset, cursor), result caps, or how the LLM should handle large result sets. Unbounded results can exhaust context windows.