A Model Context Protocol server that provides browser automation and web testing capabilities through Playwright.
The server defines 2 tools with explicit schemas and descriptions. browser_action has a well-structured, complex input schema with proper type definitions, enums, and nested objects (Viewport, CaptureOptions). compare_screenshots has a simpler but complete schema. Both tool descriptions are present and exceed 20 characters. However, there are several quality gaps: (1) Tool naming could be more verb-focused, browser_action is generic and doesn't clearly convey the primary intent (automation, testing, or navigation); (2) Error handling in the code shows basic validation but error responses to the LLM are not documented in tool descriptions, no guidance for recovery strategies; (3) Parameter descriptions are mostly good, but some lack format/constraint details (e.g., 'mask_selectors' array items are strings but no pattern or length guidance); (4) Output schemas are not formally documented in the tool definitions, responses are mentioned implicitly in descriptions but not structured as a formal response schema; (5) Security: the code includes secrets-detection regex patterns but no documentation of what happens to captured console/network data that might contain credentials. Baseline: avg tool description ~194 chars; both tools exceed 50 chars. Avg params per tool ~4; browser_action has 2 top-level params (actions array, viewport, capture), compare_screenshots has 2. Most parameters have descriptions.
Execute a browser action or test scenario. Supports a sequence of actions including navigation, clicking, form filling, text validation, screenshot capture, and more.
Compare two screenshots (baseline and current) and generate a diff image showing visual differences.
Tool naming lacks action-verb clarity. 'browser_action' is generic and does not immediately convey intent (automation, testing, network troubleshooting). LLMs cannot infer from name alone whether to use it for navigation, form filling, or validation.
Output schemas are not formally documented. Tool descriptions mention return values (screenshot artifacts, DOM text, network errors, console logs) but no structured JSON response schema is provided. LLMs cannot plan downstream tool calls or extract fields reliably.
Error handling guidance missing from tool descriptions. Code includes validation (required fields, timeout constraints) but tool descriptions do not explain what errors can occur, when they are retryable, or what the LLM should do next (e.g., 'If selector not found, try wait_for_selector first').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Sensitive data handling not documented. Code captures console, network, and DOM, which may include passwords, tokens, or auth headers. Descriptions do not mention sanitization, data retention, or what happens to captured secrets. The regex patterns in code (SECRET_KEY_RE, AUTH_HEADER_RE) suggest intent to filter, but tool contract doesn't specify this.
Parameter dependency documentation incomplete. ScenarioStep.validate_required_fields enforces action-specific field requirements in code (e.g., goto requires 'url', click requires 'selector'), but these constraints are not documented in parameter descriptions. LLMs must guess which fields apply to which actions.
Mask_selectors parameter lacks constraint documentation. Described as 'Additional selectors to mask in screenshot steps' but no guidance on: CSS selector format, max array length, what 'masking' means visually, or whether invalid selectors cause tool failure.