E2E test runner using Chrome Pool (browserless/chrome) with parallel execution, visual verification, and AI-powered test generation
This MCP server defines 14 E2E testing tools with mostly complete schemas and descriptions. Strengths: all tools have non-empty descriptions (avg ~150 chars), comprehensive input schemas with type info and enums where appropriate, and clear risk categorization (READ_ONLY vs WRITE). Weaknesses: descriptions are somewhat verbose and tutorial-heavy (especially e2e_create_test), parameter descriptions lack validation constraints and format guidance, no explicit output schema documentation (responses inferred from context), error handling is not documented within tool definitions, and some tools conflate multiple concerns (e.g., e2e_dashboard 'start/stop/status' could be two tools). Naming is strong (verb_noun convention), but some tools exhibit responsibility overload. Average tool definition quality: 62/100.
Analyze a live URL to detect interactive elements, forms, buttons, and page structure. Returns JSON-LD structured data, Aria landmarks, form fields, and clickable elements.
Create a reusable test module. Modules are parameterized action sequences that can be invoked by multiple tests via {"$use": "module-name", "params": {...}}. Useful for login flows, common workflows, and repeated patterns.
Create a new E2E test JSON file. Prefer built-in actions over evaluate — more robust and readable. Full catalog: the e2e-testing skill / references/action-types.md. Action cheat-sheet: - Click: click (by text), click_regex, click_menu_item, click_option, click_chip, click_icon, click_in_context; in a dialog use click with scope:"dialog" (+ last/visible). - Select (MUI): select_combobox (open+optional filter+pick), select, focus_autocomplete. - Assert text: assert_text (present), assert_no_text (absent), assert_text_in (scoped regex), assert_element_text, assert_matches. - Assert elements (selector, NOT text): assert_visible, assert_not_visible, assert_count, assert_attribute, assert_input_value, assert_url. - Nav/wait: goto, navigate (SPA), wait {text|selector|gone|value(ms)}, wait_network_idle. - Form: type, type_react (React inputs; optional blur/waitAfter), clear, press. Field rules: assert_text/assert_no_text use "text" (whole page); assert_visible/assert_not_visible use "selector"; for text absence use assert_no_text. Use evaluate only for computed styles, complex logic, GraphQL (window.__e2eGql), or app state. Modules: { "$use": "module-name", "params": {...} } references reusable modules in e2e/modules/ (they compose). Run e2e_list to see available modules.
Start or stop the dashboard web server. The dashboard provides a UI for running tests, viewing results, analyzing stability, and managing modules.
No output schema documentation. Tools have input schemas but responses are inferred rather than formally documented. LLMs cannot plan downstream operations without knowing what fields to expect. For example, e2e_run returns test results, but the exact structure (fields, types, pagination) is not specified in the tool definition.
e2e_create_test description is 500+ characters and reads as a tutorial (action cheat-sheet, field rules, modules guidance). Should be concise (50-200 chars), with a separate 'detailed guide' or link. Current length wastes tokens and buries the core action ('Create a new E2E test JSON file').
Parameter descriptions lack validation constraints. Examples: 'suite' description doesn't specify format (alphanumeric, hyphens?), 'concurrency' has no min/max range, 'verificationStrictness' enum is documented but 'moderate' is marked 'default' in text but not in schema. Add format specs, ranges, and explicit defaults.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | <=2025-11-25 | v2 |
Fetch an issue from GitHub, Jira, Linear, or other platforms and generate a test prompt. Detects the platform from the URL and fetches issue details + suggested test actions.
Query test stability insights from the learning database. Queries include: summary (overview), flaky (flakey tests), errors (recurring error patterns), selectors (CSS selector stability), pages (page health), api (API call reliability), trends (performance/stability trends), history (test-by-test timeline), insights (per-run analysis).
List all available E2E test suites with their test names and counts.
Analyze test modules for extraction candidates and unused modules. Identifies duplicated 3-8 action sequences that should be refactored into reusable modules.
Fetch network request/response logs for a specific test run. Shows all HTTP requests, responses, headers, timing, and network errors.
Get the status of the Chrome pool (browserless/chrome). Shows availability, running sessions, capacity, and queued requests.
Run E2E browser tests. Specify "all" to run every suite, "suite" for a specific suite, or "file" for a JSON file path. Returns structured results with pass/fail status, duration, and error details.
Take a screenshot of a URL (or capture a visual state). Optionally register the hash for future verification. Returns image path and computed visual hash.
Manage test variables (e.g. auth tokens, user IDs, URLs). Set, get, list, or delete variables for use in test hooks or actions.
Use Claude API to generate an end-to-end test for an issue. Fetches issue details, analyzes the target page, and generates a complete test JSON with actions and assertions.
e2e_dashboard tool conflates three responsibilities: start, stop, and status. Per pattern:tool, each tool should do one thing. Split into e2e_start_dashboard, e2e_stop_dashboard, and e2e_get_dashboard_status for clarity and composability.
No error handling guidance in tool descriptions. If e2e_run fails (e.g., Chrome pool unavailable, test syntax error), what should the LLM do next? Add error classification (retryable, user-fixable, fatal) and recovery hints.
e2e_issue and e2e_verify_issue both fetch and generate tests from issues. Naming does not clearly distinguish: e2e_issue appears to 'fetch' while e2e_verify_issue 'generates'. Rename to e2e_fetch_issue and e2e_generate_test_from_issue for clarity.