MCP server and CLI for AI-powered design fidelity — compare Figma designs (or screenshots) against rendered UI. Measured on a golden eval set: 90% mutation recall, zero false positives on clean pairs. Uses your Claude Code or Codex subscription — no API keys.
Kiyas implements 3 well-defined tools with strong descriptions and comprehensive schemas. All tools have clear verb-based names (compare, get_diff_report, list_issues), detailed parameter documentation, and structured output. The compare tool is particularly strong with 16 parameters, each with descriptions and type constraints. Schemas use Zod validation with proper type definitions. Main gap: no output schema documentation in the visible code, tool annotations (readOnlyHint) are missing but appropriate for this read-focused server, and error handling guidance could be more explicit.
Compare a design against a rendered implementation and return discrepancies (measured on kiyas's golden eval set: 90% mutation recall, zero false positives on identical pairs). Provide `figma` (frame URL) or `designImage` (local path or URL of a design screenshot), plus either `target` (a URL) or `component` (natural-language description; kiyas finds it in the codebase). No Figma token needed for `designImage` — if a Figma MCP server is connected, export the frame as an image with its screenshot tool and pass that here. Returns a reportId you can pass to get_diff_report or list_issues.
Fetch a stored kiyas report by reportId. Defaults to JSON; pass format=html to fetch the rendered report. Returns the artifact path and (by default for JSON) inline content.
List discrepancies from a stored kiyas report by reportId, optionally filtered by severity (all | high | medium | low).
Output schemas not formally documented in code or README. Agents cannot plan downstream actions if return types are not explicit. For example, compare returns reportId but the structure and type are inferred, not declared.
Tool annotations (readOnlyHint) missing. While all 3 tools are READ-ONLY (no mutations), declaring this via readOnlyHint in the tool definition would let agents optimize request batching and understand retry semantics.
Error handling lacks recovery guidance. No visible error responses with actionable next steps (e.g., 'reportId not found. Available reports: [...]'). Errors should guide the agent on retryability and alternatives.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 74 | 2026-07-28+ | v2 |
compare tool has complex multi-parameter logic (figma vs designImage, target vs component, model enum, colorScheme auto-detect). Parameter relationships and dependencies documented in descriptions but no mutual exclusion enforcement or validation error messages visible in code.
list_issues lacks pagination info. If a report contains hundreds of issues, return structure is unknown, is there a limit? Offset/cursor parameters? Total count? Without pagination schema, large reports risk context window exhaustion.