Your AI can already see the screen — it just can't measure it. uisight is an MCP server for web/responsive UIs: live mobile+desktop sessions, a measurement engine (contrast, touch targets, theme drift) that reports findings as text, and a panel you and your agent share.
uisight defines 13 tools with complete input schemas and descriptions present for all tools and parameters. However, the server exhibits significant gaps in production quality: (1) Naming violates verb_noun conventions, tools like 'inspect', 'marks', 'state' are nouns or vague verbs rather than action-oriented names; (2) Descriptions are terse (most 10-40 chars), failing to explain WHEN to use each tool or what downstream actions are enabled; (3) Output schemas are completely undocumented, no structured field descriptions for what inspect(), screenshot(), or marks() return; (4) Error handling is entirely absent, no guidance on how LLMs should recover from failures; (5) Tool composition issues, goto, click, type, role, login are bare UI drivers with no context about session state or dependencies. The server is STDIO-only, capping protocol readiness at 50. Input schemas are well-formed (all params typed, enums for session), but descriptions read like bare API docs, not LLM-optimized guidance. Average tool quality is pulled down by missing output documentation and lack of error recovery patterns.
Audit back button behavior and history stack functionality
Click an element on the page using a selector or text match
Navigate to a URL in the active session
Measure the current screen: contrast, touch targets, covered controls, theme drift, and other accessibility/UI issues
Audit keyboard navigation and input handling in the current session
Extract all links from the current page within a given root domain
Authenticate to the application using configured credentials
Non-action verb naming: 'inspect', 'marks', 'state' are nouns or unclear verbs; LLMs expect verb_noun structure (get_*, measure_*, capture_*)
Output schemas completely missing: tools like inspect(), screenshot(), marks(), state() return complex structured data but no field documentation is visible. LLMs cannot extract required fields or plan downstream calls.
Terse descriptions (10-40 chars): 'Authenticate to the application' (44 chars), 'Type text into the focused input field' (42 chars). Missing WHEN-to-use guidance and prerequisites. Descriptions should explain user intent, when to prefer this tool vs alternatives, and what errors to expect.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 57 | 2026-07-28+ | v2 |
Retrieve marked/pinned notes from the live panel along with their associated screenshots
Audit offline functionality and service worker behavior
Switch to a different user role in the application (requires authenticated session)
Capture a screenshot of the current session at a specific scale
Retrieve the current state of all sessions, URLs, theme, and error status
Type text into the focused input field
No error handling or recovery guidance: tools have no documented error conditions, retryability hints, or recovery steps. If inspect() fails, LLM has no guidance on what to try next.
Ambiguous tool composition: bare UI drivers (click, type, goto, role) lack context about session state dependencies. No documentation on whether role() requires login() first, or how session state is managed across calls.
Parameter 'root' in links() lacks validation constraints: is it a full URL, domain name, or regex pattern? No enum or pattern description forces LLM to guess.