Design review for the UI your agent just wrote — audits a site's UX and returns a concrete fix for every finding.
Single tool 'audit_url' has a well-structured schema with 8 parameters, all typed and described. Description is clear and action-oriented (94 chars, within 10-1024 baseline). Parameter descriptions are present and specific (e.g., 'Base URL to audit', 'Routes to audit'). However, output schema is not documented in the source, the tool returns audit findings but the response structure is not visible. No error handling guidance documented. Tool naming follows verb_noun pattern ('audit_url') and is unambiguous. Risk classification (READ_ONLY) is correct. Missing: output schema documentation, error recovery guidance, pagination/limit constraints on results.
Audit a site's UX — contrast, tap targets, type scale, copy, scan patterns, resilience — with a concrete fix for every finding.
Output schema not documented. Tool returns audit findings but response structure (fields, types, pagination) is not visible in source. LLMs cannot plan downstream actions or extract specific data without knowing what fields to expect.
No error handling guidance. Tool description does not explain what happens on failure (network timeout, invalid URL, Chrome crash, authentication failure). LLMs have no recovery path.
Result limits not specified. 'routes' parameter accepts an array with no documented max length. A large routes array could produce unbounded output, exhausting context window. No pagination or limit parameter offered.
Parameter 'viewports' is a free-form string ('1280x720,375x667'). Should be an array of objects with width/height fields, or an enum of preset sizes. Free-form strings invite parsing errors and hallucinated formats.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 58 | 2026-07-28+ | v2 |
Parameter 'parallel' lacks bounds. Description says 'default: auto-scaled to cores' but no min/max specified. LLMs could pass 1000, causing resource exhaustion. Should document valid range (e.g., 1-16).