Code-aware browser testing for AI coding agents. Reads your codebase so tests use real routes and field names, runs them in Playwright, remembers what broke, and reports what your last change fixed or regressed. No LLM calls inside. MCP server for Claude Code/Cursor/Windsurf, or standalone CLI.
vibe-testing is a specialized browser testing tool with 7 tools offering reasonable descriptions and functional schemas. However, there are significant quality gaps: (1) Parameter descriptions are frequently missing or vague, violating the pattern that every parameter needs explicit documentation; (2) Output schemas are not formally documented, the code shows input schemas but return types are inferred from narrative descriptions only; (3) Error handling guidance is absent, tools do not indicate what errors are retryable, how to recover, or what to do next; (4) Some descriptions lack actionable detail for LLM selection (e.g., 'explore_page' says 'tries every element' but doesn't explain when to call it vs scan_page_elements); (5) The scenario structure in execute_scenario is deeply nested and complex, making it difficult for LLMs to construct valid payloads without trial-and-error. The tool set is well-intentioned and covers a coherent domain (codebase scanning → element discovery → scenario execution → reporting), but falls short of production-grade definition quality due to incomplete parameter documentation and missing output schema declarations.
Execute a single test scenario (a sequence of navigate/fill/click/assert steps) and return detailed results with step-by-step logs, screenshots after each state-changing step, API errors observed, and the final page state. The editor LLM can construct scenarios based on scan_codebase output or create custom ones.
Perform a full interactive exploration of a page: discover all elements, click buttons, fill inputs, test tabs, observe API calls, and report what happened. Returns detailed interaction outcomes, API observations, and screenshots. This is the "senior tester" mode — it tries every element and reports what works and what breaks.
Generate an HTML report of test results, coverage gaps, and recommendations. The report summarizes all scenarios executed in this session, shows which ones passed/failed, highlights untested areas, and suggests next test scenarios.
Establish an authenticated browser session by executing a login scenario. Returns the post-login URL, token state, and a screenshot. Uses saved credentials from previous runs if available, or accepts provided credentials.
Execute all generated test scenarios (from scan_codebase) in sequence and return aggregated results. This is the main entry point for automated testing — it scans, logs in if needed, runs all scenarios, and generates a report.
Output schemas not formally documented. Tools describe return values in prose (e.g., 'Returns a ProductModel with routes, behaviours, coverage map, gaps, and generated test scenarios') but do not provide structured JSON Schema definitions for the response objects. LLMs cannot reliably parse what fields to expect or plan downstream operations.
Parameter descriptions missing or generic. 'login' tool has parameters 'email' and 'password' with brief descriptions but no guidance on format or when to use saved credentials. 'execute_scenario' has a deeply nested 'scenario' object with 23 sub-fields (id, name, route, steps[], expected_outcome, etc.) but only 1-2 lines per parameter. LLMs cannot construct valid payloads without extensive trial-and-error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
Analyze a project's codebase to understand its structure, routes, forms, components, existing tests, and coverage gaps. Returns a ProductModel with routes, behaviours, coverage map, gaps, and generated test scenarios. Call this first before any testing.
Navigate to a specific page and discover all interactive elements (buttons, links, inputs, selectors, checkboxes, tabs). Returns a structured list of elements with their types, text, selectors, and disabled state. Also returns a screenshot of the page. Use this to understand what's on a page before deciding what to test.
No error handling guidance. None of the 7 tools document what errors they can return, whether errors are retryable, or how to recover (e.g., 'If login fails due to invalid credentials, try again with correct email/password'; 'If scenario execution times out, try increasing timeout or simplifying the scenario'). Agents have no actionable recovery path when a call fails.
Ambiguous tool selection guidance. 'scan_page_elements' and 'explore_page' have overlapping purposes (both discover elements on a page), but descriptions do not clearly state when to call one vs the other. The scan returns 'a structured list of elements'; explore performs 'full interactive exploration' and 'tries every element'. An LLM may pick the wrong tool or call both unnecessarily.
execute_scenario has an excessively complex nested input schema. The 'scenario' parameter contains 8 top-level fields, and 'steps' is an array of objects with 9 fields each (action, selector, value, url, timeout, description + enum constraints on action). No examples provided. LLM must reason about valid JSON structure without scaffolding, significantly increasing error rate.
Missing return field documentation. Tools like 'scan_codebase' claim to return 'a ProductModel with routes, behaviours, coverage map, gaps, and generated test scenarios' but do not specify: are 'gaps' an array of strings or objects? Do 'routes' contain nested sub-routes? Is 'coverage' a percentage or a map? Without this structure, agents cannot reliably extract or chain results.