A Model Context Protocol server for browser automation using Selenium WebDriver, enabling web testing, scraping, and interaction through MCP tools
42 tools with consistent naming convention (verb-first: createSession, closeSession, navigateTo, etc.) and all having descriptions and parameter schemas. However, descriptions are terse (20-80 chars), parameter descriptions lack depth, output schemas are undocumented, and error handling is not visible in the schema definitions. Tools follow a clean verb_noun pattern but lack the LLM-optimized richness needed for confident tool selection. Per-tool analysis: average naming=88, average description=45, average schema=65, average overall=63.
Adds a cookie to the current session
Captures a screenshot of the current page and returns it as base64-encoded PNG
Clears the text content of a web element (typically input fields)
Deletes all cookies from the current session
Clears the network activity log
Clicks on a web element identified by locator strategy and value
Closes an active WebDriver session and releases associated resources
Creates a new WebDriver session with specified browser type and optional device type for mobile automation
Output schemas undocumented, no visible return type documentation for any of the 42 tools. LLMs cannot plan downstream calls or extract fields without knowing what fields each tool returns.
Descriptions are terse (mostly 40-60 chars), below the LLM-optimized baseline of 50-200 chars. Many lack context for WHEN to use the tool or any prerequisites. E.g. 'Gets the current URL' tells an LLM nothing about when to call it vs alternative tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Deletes a cookie from the current session
Double-clicks on a web element
Drags a source element and drops it on a target element
Executes asynchronous JavaScript code and waits for its completion
Executes arbitrary JavaScript in the context of the current page
Finds a web element using specified locator strategy and value
Finds multiple web elements using specified locator strategy
Gets an attribute value from a web element
Retrieves browser console log messages (requires Chrome DevTools Protocol enabled)
Retrieves all cookies from the current session
Gets the current URL of the active page in the session
Retrieves network activity log (requires Chrome DevTools Protocol enabled)
Gets the HTML source code of the current page
Gets the title of the current page
Gets the HTML tag name of a web element
Gets the visible text content of a web element
Gets the current size of the browser window
Navigates back to the previous page in browser history
Navigates forward to the next page in browser history
Checks if a web element is displayed (visible) on the page
Checks if a web element is enabled (not disabled)
Checks if a web element is selected (for checkboxes, radio buttons, options)
Holds down a keyboard key
Releases a held keyboard key
Maximizes the browser window
Moves the mouse cursor to a web element
Navigates to a specified URL in the active session
Presses a keyboard key (e.g., ENTER, ESCAPE, TAB, etc.)
Refreshes the current page
Releases a held keyboard key
Right-clicks (context click) on a web element
Sends keyboard input to a web element
Sets the size of the browser window
Submits a form element
Parameter descriptions are minimal, most are a single phrase (e.g. 'UUID of the session'). Missing validation ranges, format constraints, and when to use alternative parameters. E.g. locatorStrategy param lacks guidance on which strategy to choose for different scenarios.
No enum constraints on free-form string parameters. locatorStrategy accepts (id, xpath, css, name, tag, link, partial_link, class) but schema shows type:string with no enum. LLMs may hallucinate unsupported values.
Error handling guidance missing, no visible recovery hints in descriptions. If findElement fails, LLM doesn't know whether to retry with a different locator strategy or call a discovery tool.
No idempotency guidance. Tools like createSession, addCookie, and executeScript should clarify whether repeated calls with same input are safe to retry or risk side effects (duplicate sessions, duplicate cookies).
Appium mobile automation (deviceType parameter) is unimplemented, throws UnsupportedOperationException with a placeholder message. Schema includes the parameter but tool doesn't work.
No pagination support documented. Tools like findElements and getNetworkLog may return large result sets but lack limit, offset, or cursor parameters. LLM context could be exhausted by a single response.
browserType parameter in createSession lacks enum constraint. Schema shows type:string but doesn't list (chrome, firefox, edge, safari) as valid enum values, inviting LLM hallucination.
No tool annotations visible (readOnlyHint, destructiveHint, idempotentHint). Risk tags are present in the provided metadata (READ_ONLY, WRITE) but not reflected in tool schema or capability announcements.