Full-featured browser automation for AI agents with Set of Marks (SoM) screenshots, developer tools, session recording, and multi-tab support.
BrowserControl has clear action verb names and generally well-structured tool definitions, but suffers from inconsistent parameter documentation, missing schema constraints (enums), and lack of explicit output schema documentation. 20 of 25 tools have descriptions (80%), but parameter coverage is uneven, many parameters lack constraint information (format, range, allowed values). No tool annotations (readOnlyHint/destructiveHint) are visible despite the risk field being documented in the server metadata. Error handling is basic (try/except with RuntimeError) without actionable recovery guidance. The tool set is well-composed (single responsibility per tool, good naming for browser automation domain), but lacks the LLM-optimization needed for A-grade definitions.
Check or uncheck a checkbox by element ID.
Click on an element by its ID number shown in the screenshot.
Click at specific x,y coordinates.
Get browser console logs (errors, warnings, info, log messages).
Get captured network requests (API calls, resources, etc.).
Get the page content as markdown text.
Get JavaScript errors that occurred on the page.
Missing explicit output schemas for 6 tools (get_page_content, get_page_info, get_console_logs, get_network_requests, get_page_errors). Tools return unstructured strings (e.g., 'Console Logs:\n[LEVEL] text (location)') instead of structured JSON. LLMs cannot reliably extract or chain data from free-text responses.
Parameter constraints missing for string enums. 'direction' in scroll should be declared as enum(['up', 'down', 'left', 'right']), not bare string. 'amount' in scroll accepts multiple formats ("small", "page", "500px") but no enum constraint. 'key' in press_key accepts arbitrary strings with no validation. LLMs will hallucinate invalid values.
No tool annotations despite clear risk classifications. Tools marked WRITE risk (navigate_to, click, type_text, etc.) lack destructiveHint annotations. READ_ONLY tools lack readOnlyHint. Idempotent tools (screenshot, get_text) lack idempotentHint. This prevents LLMs from understanding side-effect severity and retry safety.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Get current page URL and title.
Get the text content of an element by its ID.
Navigate back to the previous page.
Navigate forward to the next page.
Hover over an element by its ID number.
Inspect an element to get its computed styles, dimensions, and properties.
Navigate to a URL. Returns an annotated screenshot with numbered interactive elements.
Press a keyboard key.
Refresh the current page.
Execute JavaScript code in the browser console and return the result. Useful for debugging, inspecting variables, or manipulating the page.
Execute JavaScript and return the result.
Take a screenshot of the page.
Scroll the page.
Scroll to bring an element into view.
Select an option from a dropdown by element ID.
Type text into an input element by its ID number.
Upload a file to a file input element.
Wait for a specified time (useful for pages with animations or loading).
Insufficient error guidance. Error responses are generic RuntimeError with minimal context. E.g., 'Get console logs failed: [exception]' does not tell the LLM whether to retry, check browser state, or escalate. No recovery hints (e.g., 'Try screenshot() first to ensure page is loaded').
Parameter descriptions lack constraint details. 'seconds' in wait has no min/max. 'num_requests' in get_network_requests defaults to None but should specify range (1 - 100). 'x', 'y' in click_at lack viewport bounds. LLMs will pass out-of-range values without validation feedback.
Pagination missing from list-like tools. get_console_logs, get_network_requests, and get_page_errors all internally cap results (50, 30, 20 logs respectively) but expose no pagination mechanism (limit, offset, next_cursor) to the LLM. If more than 50 console logs exist, the agent cannot retrieve them without losing data.
No confirmation or dry-run for destructive operations. Tools like upload_file (file_path parameter) and run_javascript should support a dry_run flag or explicit confirmation to prevent accidental data destruction. Agents can be tricked via prompt injection into destructive actions.