MCP Server for visual UI testing and browser automation
This MCP server provides 5 tools for browser automation and visual testing. Tool definitions have mixed quality. All tools have descriptions and input schemas are partially defined, but there are significant gaps in schema completeness, parameter descriptions, and output documentation. The tools are well-intentioned but lack the rigor expected of production-grade agent tools. Tool descriptions exist but are generic and lack the LLM-optimized clarity needed for agent tool selection (baseline 50-200 chars; most here are 100-150 chars with generic language). Parameter descriptions are present but inconsistent, some parameters lack detail on constraints, valid ranges, or dependencies. Error handling and recovery guidance are minimal. This server sits at the C/D boundary: noticeable gaps in schema completeness, output documentation missing, and parameter validation/guidance could be much stronger.
Comprehensive accessibility testing with WCAG audits, color contrast analysis, and keyboard navigation testing
Comprehensive browser monitoring for console logs, network requests, JavaScript errors, and performance metrics
Comprehensive form handling for web automation including field population, submission, validation, and file uploads
Comprehensive user journey simulation with multi-step execution, recording, validation, and optimization
Locate web elements using Playwright selector syntax with full support for CSS, XPath, text selectors, shadow DOM, iframes, and compound selectors
Missing output schema documentation. No tool documents what fields or structure it returns. LLMs cannot plan downstream actions or extract correct data without knowing response structure.
Vague parameter descriptions for shared parameters across tools. For example, 'html' parameter is described identically in multiple tools ('HTML content to set for testing (optional)') but its interaction with page state, DOM replacement behavior, and persistence is never clarified. Parameters like 'url', 'timeout', 'selector' would benefit from format specifications and constraint examples.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 34 | - | v1 |
journey_simulator uses a 'steps' parameter accepting a JSON string of objects. This is extremely error-prone, LLMs frequently produce invalid JSON or miss required fields. The parameter lacks detailed examples, strict validation hints, or enumeration of all allowed step actions and their required/optional properties. This violates the pattern of constrained input and actionable error guidance.
form_handler's 'data' parameter is an untyped object ('type:object', no properties schema). The description mentions it supports 'placeholder text or field identifiers as keys for React forms' but provides no enumeration, validation, or guidance on what fields exist. An LLM has no way to know which keys are valid without calling detect_fields first, forcing an extra round-trip.
Enum descriptions lack clarity on when each option applies. For example, locate_element's 'type' enum lists ['css','xpath','text','aria','data'] but does not explain the difference or when to use each. This forces LLMs to guess or make multiple calls.
No error handling or recovery guidance. The tools are silent on what errors can occur, how to detect them, or what to do next. None of these tools provide that, no categorization of retryable vs. fatal errors, no actionable error messages, no recovery hints.
browser_monitor tool has overlapping parameter definitions. For example, 'consoleFilter' is a complex nested object, but the tool also accepts 'type' and 'textPattern' as top-level parameters for get_console_logs. This creates ambiguity, which parameters apply to which action? The tool description does not clarify these dependencies.
Implicit state management across tools. Multiple tools accept 'url' and 'html' parameters and appear to initialize page state, but the semantics are never documented. Does passing 'html' replace the current page? Does it persist across calls? Does passing 'url' navigate away and lose previous state? This breaks the agent's ability to reason about state consistency.