MCP server that turns text prompts into premium iOS mobile UI mockups (PNG + HTML). Powered by Claude + Playwright.
Strong tool definitions with clear naming, good descriptions, and complete input schemas. All four tools follow verb-first naming convention (generate_, iterate_, list_, get_). Descriptions range from 120 - 180 chars, within the production baseline (p10=34, p90=392). All parameters have type definitions and descriptions. However, output schemas are not documented, no specification of what fields get_screen or list_screens return, which limits LLM planning. Tool compositions are sound: generate_screen → iterate_screen → get_screen forms a logical chain. No tool annotations (readOnlyHint/destructiveHint) present, which is a missed signal for agent reasoning.
Generate a premium iOS mobile UI mockup from a text brief. Outputs both the screenshot (PNG) and the underlying HTML. Use this when the user asks to design, mock up, or visualize a mobile app screen.
Get details and metadata for a specific screen. Set include_html=true to also return the HTML source.
Refine an existing generated screen based on feedback. Use this for follow-up edits like 'change color', 'add a section', 'make it more spacious'.
List all generated screens, optionally filtered by project name.
Output schemas not documented. No specification of what fields list_screens, get_screen, or generate_screen return. LLMs cannot plan downstream tool calls or extract fields without documented response structure.
Tool annotations missing. No readOnlyHint, destructiveHint, or idempotentHint tags. generate_screen and iterate_screen modify state (write design files), but this is not signaled to the agent. list_screens and get_screen are read-only but lack explicit annotation.
No pagination parameters or limits on list_screens. If a user has hundreds of screens, list_screens will return all of them, potentially bloating the context window. Should accept limit (default 20 - 50) and offset/cursor parameters, and return a total_count.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
Error handling is generic. Handler returns 'Error: {message}' as plain text. No guidance on how the LLM should recover (retry? call search_screens? provide different input?). No categorization of errors as retryable vs. user-fixable vs. fatal.
generate_screen accepts optional 'design_system' parameter as free-form string. No constraints, examples, or enum options. LLM will hallucinate design system specs. Should provide examples or reference a known design system schema.