In-loop, read-only design review server for coding agents over MCP. Provides asynchronous design reviews for tenant-authorized HTTPS previews with judgment provenance, structured findings, and interactive MCP-Apps panels.
Strong foundational quality with comprehensive input schemas, detailed descriptions, and production-grade error handling guidance. All five tools have explicit input schemas with proper types, constraints, and descriptions. Tool names follow verb_noun patterns (design_review, design_review_get, design_recheck, design_review_cancel, design_review_panel_action). Descriptions are detailed (140-240 chars) and include operational guidance (polling rates, idempotency keys, state-conditional logic). Schemas use JSON Schema with minLength/maxLength/pattern/enum constraints. However, output schemas are not visible in source code (tool-catalog.ts reads them from schemas/mcp-tools.json, which is not provided), so we cannot verify documentation of return structures. Tool annotations are present (Risk field indicates READ_ONLY, WRITE, REVERSIBLE). Error handling is implicit in descriptions (e.g., 'Unchanged targets and exhausted recheck loops are rejected') but explicit error codes/types are not shown in source. The server implements a strict schema-conformance test pattern (tool-catalog.ts validates all payloads against schemas/mcp-tools.json), which is excellent for production quality.
Submit a metered recheck for findings from a completed review after the customer's agent changes the UI. This tool never edits code. Unchanged targets and exhausted recheck loops are rejected without running judgment. Check recheck.provenance before acting: when nothing judged the target every outcome is unjudged with a null confidence. Reuse client_request_id on retries.
Submit an asynchronous, metered design review for a tenant-authorized HTTPS preview. This tool never edits code. Reuse client_request_id on retries, then poll design_review_get no faster than poll_after_ms.
Request best-effort cancellation of a queued or running review job. Terminal jobs keep their existing state. Cancellation does not edit customer systems and consumes no review units.
Get status or a compact, focused, or evidence view for an existing review job. Poll no faster than the returned poll_after_ms. Result reads do not consume review units. Act on review.measurements.violations unconditionally: they are computed from the captured DOM and are true whether or not a model ran. Act on review.findings only when provenance.model_backed is true and coverage.state is full or partial.
Output schemas not visible in source code. tool-catalog.ts loads schemas from external schemas/mcp-tools.json file, which is not provided. Cannot verify documentation of return structures, field types, or whether output is shaped for downstream tool chaining.
design_review_panel_action has a weak name: 'panel_action' is vague and doesn't clearly convey what the tool does. Names like 'apply_finding_fix' or 'dispatch_panel_interaction' would better reflect the actual responsibility of handling reviewer interactions.
Error handling is described in prose within parameter descriptions (e.g., 'Unchanged targets and exhausted recheck loops are rejected without running judgment') but there is no explicit error response schema or error code catalog. LLMs cannot reliably learn what exceptions to expect or how to recover from failures.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 86 | 2025-06-18+ | v2 |
Interactive MCP-Apps panel callback tool. The panel rendered by design_review_get with view: "evidence" calls this tool when the reviewer acts on a finding. apply_fix returns a grounded finding's fix for the host to hand to the coding agent, or human_only for advisory judgment. The pure reducer in panel-interaction.ts decides what happens.
design_review_get's 'view' parameter has valid values ('status', 'summary', 'findings', 'focus', 'evidence') but only 'status' and 'summary' are documented in the description; 'findings', 'focus', and 'evidence' lack guidance on when to use them or what data each returns.
design_review and design_recheck require 'client_request_id' for idempotency but the pattern is only documented in prose; there is no explicit note that these tools are idempotent or guidance on when to reuse vs. generate new IDs.