MCP server for parallel browser automation across multiple providers.
The server defines 26 tools with reasonable naming and descriptions. Most tools have explicit schemas with parameter types and descriptions. However, there are consistent gaps: (1) descriptions are often generic and lack detail about when/why to use each tool or what it returns; (2) many tools lack output schema documentation; (3) error handling is minimal, most tools return generic errors with no recovery guidance; (4) several tools have overlapping functionality without clear distinction (e.g., browser_snapshot vs browser_get_page_structure, browser_click vs browser_mouse_click_xy); (5) no tool annotations (readOnlyHint, destructiveHint, idempotentHint) are visible despite many tools being write operations. The tool naming follows verb_noun convention well (browser_*, start_session, close_session), which is a strength. Parameter descriptions are present but often minimal (e.g., 'Numeric session ID' repeated 26 times without context on how to obtain a sessionId). Output schemas are not documented, the code returns JSON/text results but the agent has no way to know what fields to expect.
Click an element by selector.
Query element presence, count, and state without waiting.
Drag from one element to another.
Run JavaScript in the page context.
Fill a field with text.
Fill multiple form fields.
Generate locator suggestions for an element.
Missing output schema documentation. No tool documents what fields/types it returns. Agents cannot plan downstream calls or extract specific data.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). 18+ write tools (browser_click, browser_fill, close_session, etc.) lack destructiveHint, preventing the agent from understanding irreversibility and safe-retry boundaries.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 66 | <=2025-11-25 | v2 |
Return a readable page structure summary.
Go back in browser history.
Hover over an element by selector.
Press a keyboard key or key chord.
Type text into the active page.
Click at viewport coordinates.
Drag the mouse from one point to another.
Move the mouse to viewport coordinates.
Navigate the session page to a URL.
Capture a screenshot of the current page.
Select options in a select element.
Return a structured page snapshot.
Upload files to a file input.
Wait for an element to match a given state.
Wait for a specified duration.
Close all active browser sessions.
Close a browser session by numeric session ID.
List all active browser sessions.
Start a browser session using the configured provider.
Weak error handling. Code catches errors and returns generic messages (e.g., 'Tool failed: <message>'). No recovery guidance, no actionable next steps, no categorization of retryable vs fatal errors.
Overlapping tool functionality without clear distinctions. browser_snapshot and browser_get_page_structure both return page structure; browser_click and browser_mouse_click_xy both click; browser_drag and browser_mouse_drag both drag. LLM will waste reasoning cycles choosing between them.
Minimal parameter descriptions. 'Numeric session ID' is repeated verbatim across all 20+ browser_* tools. No guidance on how to obtain a sessionId, whether it persists across calls, or what happens if invalid.
Tool descriptions lack WHEN/WHY guidance. 'Navigate the session page to a URL' says WHAT but not when to use this vs browser_go_back or browser_evaluate. 'Capture a screenshot' doesn't explain when to call it vs browser_snapshot or browser_get_page_structure.
No confirmation/dry-run pattern for destructive operations. close_session and close_all_sessions have no safety mechanism, an agent in a retry loop could accidentally close all sessions without warning.
Bare numeric sessionId parameters. Tool descriptions say 'Numeric session ID' but don't explain it's returned by start_session or how to list active sessions. Implies agent must track IDs manually, brittle and error-prone.