An MCP server that evaluates the user experience and interface of web applications by performing specific tasks and analyzing interaction flow using browser automation and AI agents.
Server defines 2 tools with visible schemas and descriptions. web_eval_agent has a well-structured schema with 3 parameters (url, task, headless_browser), each with descriptions. setup_browser_state has minimal schema (1 optional parameter). Descriptions exist but lack depth for production quality, neither explains error scenarios, recovery paths, or return value structure. No output schema is documented for either tool. Schema parameter type coverage is incomplete: url and task lack explicit 'required' array declarations in the visible schema block. No parameter validation guidance (e.g., URL format constraints, task length limits) appears in descriptions. Tool names use verb_noun (web_eval_agent, setup_browser_state) following conventions, but 'web_eval_agent' is somewhat generic; 'evaluate_webpage_ux' or 'assess_webpage_usability' would be clearer about what it returns. Error handling in source code (tool_handlers.py) validates required arguments but returns only a basic error string, no recovery guidance, no error categories. The log server integration is visible but not documented in tool descriptions, leaving LLMs unaware of the side effect (opening a dashboard). Missing: documented output schemas, actionable error messages, parameter format constraints, idempotency guarantees, and permission/scope declarations.
Sets up and saves browser state for future use. This tool should only be called in one scenario: 1. The user explicitly requests to set up browser state/authentication Launches a non-headless browser for user interaction, allows login/authentication, and saves the browser state (cookies, local storage, etc.) to a local file.
Evaluate the user experience / interface of a web application. This tool allows the AI to assess the quality of user experience and interface design of a web application by performing specific tasks and analyzing the interaction flow. Before this tool is used, the web application should already be running locally on a port.
No output schemas documented for either tool. LLMs cannot determine what fields to expect, forcing unstructured text parsing and downstream failures.
Error messages lack recovery guidance. When arguments are missing or invalid, responses say 'Error: ...' but do not suggest next steps (e.g., 'Try using http://localhost:PORT or https://...'). Agents cannot self-correct.
Parameter descriptions lack format constraints. 'url' accepts 'localhost URL' but also silently prepends 'https://' if protocol is missing, behavior not documented. LLMs may pass malformed URLs expecting the tool to reject them rather than auto-correct.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 23 | - | v1 |
No permission or scope declarations. Tools interact with the local browser and filesystem (save browser state to disk) but do not declare what permissions an agent needs (e.g., 'read:filesystem', 'write:filesystem', 'launch:browser'). Audit and least-privilege configuration impossible.
Side effects not documented. 'web_eval_agent' starts a log server and opens a dashboard as a side effect (visible in tool_handlers.py: start_log_server(), open_log_dashboard()) but the tool description does not mention this. LLMs unaware of non-idempotent behavior and may call it multiple times expecting no side effects.
No idempotency guarantees. 'web_eval_agent' performs browser interactions that may not be repeatable (takes screenshots, evaluates UI state). Agents retry on ambiguous failures, tool should declare whether repeated calls with same inputs produce same outputs, or state what side effects occur on retry.
Unused/unclear parameter: 'headless_browser' in web_eval_agent. Description says 'Whether to hide the browser window' and defaults to False, but tool code passes headless=headless (inverted logic?). Description does not explain performance implications or when to use headless vs non-headless.