MCP server exposing SeleniumBase Pure CDP Mode as tools for MCP clients. Controls web browsers through Chrome DevTools Protocol (CDP) without WebDriver, with CAPTCHA-solving support and stealth capabilities.
The SeleniumBase MCP server defines 22 browser automation tools with comprehensive descriptions but exhibits inconsistent schema completeness and several naming/parameter issues. Tools have reasonable descriptions (avg ~150-200 chars, within baseline), but parameter documentation varies significantly. Input schemas are present for all tools but lack detailed type constraints (enums, ranges, patterns). Many parameters accept free-form strings where enums would prevent hallucination. Error handling guidance is minimal, most tools lack recovery hints. The server lacks tool annotations (readOnlyHint, destructiveHint) despite clear risk classifications. Notable strengths: clear naming (verb_noun convention mostly followed), stateless design. Notable weaknesses: no parameter validation examples, limited enum usage, missing idempotency guidance, no output schema documentation in descriptions.
Verify an expected condition and treat failure as an assertion error.
Perform an immediate, non-waiting state check of a condition.
Click an element on the page.
Close the active browser session and clean up resources.
Execute custom JavaScript on the page and return the result.
Discover and inspect multiple elements matching a selector, returning structured data.
Set focus on an element for positioning and visual focus.
Missing enum constraints on 'action' parameters. Tools like open_url, manage_history, manage_cookies, manage_storage, manage_tabs, and check_if_condition accept free-form action strings (e.g., 'open', 'back', 'forward', 'refresh') but do not declare them as enums. LLMs will hallucinate invalid actions like 'navigate', 'prev', 'reload' instead of using the exact valid values. This violates the constrained-input pattern.
No tool annotations for risk indication. All 22 tools are missing destructiveHint and readOnlyHint annotations. The schema shows WRITE vs READ_ONLY risk classifications, but these are not exposed to the MCP client via tool annotations. Agents cannot determine which tools are safe to retry or which require confirmation without reading raw risk labels.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
Get attributes from an element or all elements matching a selector.
Extract text or HTML content from the page or from specific elements.
Get browser metadata: current URL, page title, and page state.
Perform hover, click-and-hold, or drag-and-drop actions on elements.
Get, clear, save, or load cookies.
Read the browser's navigation history, clear history, or manage browser history state.
Get, clear, save, or load localStorage/sessionStorage.
List open tabs, switch tabs, or close tabs.
Navigate to a URL or perform browser history operations.
Save page output as a PNG, PDF, or HTML file.
Select an option from a dropdown or select element.
Click the checkbox of a CAPTCHA on the page.
Launch a persistent SeleniumBase Pure CDP Mode browser session. Call this before using browser interaction tools such as open_url, get_content, click_element, type_text, or find_elements. The same browser session remains active across subsequent MCP tool calls until close_browser is called or the server process exits. Pure CDP Mode controls the browser through the Chrome DevTools Protocol (CDP), not WebDriver.
Type text into an input field or element.
Wait for a condition to become true with optional timeout.
Missing output schema documentation. No tool description documents what fields/structure is returned. LLMs cannot plan downstream tool calls or extract specific response fields without knowing the output schema. E.g., find_elements should document it returns [{selector_match, tag, text, attributes, position}, ...], get_page_info should document {url, title, ready_state}.
Error handling lacks recovery guidance. Tool descriptions do not explain what to do if a call fails (e.g., 'If element not found, try wait_for_condition first' or 'If proxy fails, check proxy syntax SERVER:PORT or USER:PASS@SERVER:PORT'). Agents receive no actionable guidance on retries or alternatives.
Parameter ranges and formats undocumented. Numeric parameters like 'tab_index' and 'limit' lack minimum/maximum constraints. String parameters like 'selector', 'proxy', 'browser_executable_path' lack format hints. LLMs may pass negative indices, excessive limits, or malformed paths without guidance.
Conditional parameter dependencies not documented. E.g., hover_action's 'target_selector' is only valid when action='drag_to', but this dependency is not stated. LLMs may pass target_selector with action='hover' causing silent failures or ambiguous behavior.
No idempotency guidance. Tools like start_browser, click_element, type_text, and execute_script do not indicate whether repeated calls with identical parameters are idempotent. If an agent retries due to ambiguous failure, it may duplicate clicks, type text twice, or re-execute scripts with side effects.
Mutually exclusive parameters not documented. E.g., start_browser's 'browser_executable_path' and 'use_chromium' are mutually exclusive, but no description warns against combining them. Similarly, 'incognito' and 'guest' conflict, but the description only says 'Do not combine incognito=True' without explaining the consequence.
Check_type parameter accepts free-form strings with no enum constraint. Valid values are 'visible', 'enabled', 'exists', 'text_visible', 'has_value', but LLMs may pass 'is_visible', 'disabled', 'visible_text', 'contains_text' without validation. Should be an enum.
No confirmation pattern for destructive operations. Tools like close_browser, manage_cookies (clear), manage_storage (clear), and manage_tabs (close) are irreversible but offer no dry-run or confirmation step. An agent mistake (e.g., 'clear all cookies') cannot be undone.