MCP server for browser automation using Selenium WebDriver. Provides tools to control browsers, interact with web pages, perform automation tasks, and extract data from websites.
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
Browser Automation MCP exhibits significant quality gaps across naming, descriptions, and schema completeness. While tool names follow verb conventions (navigate_to_url, click_element, take_screenshot), most parameter descriptions are minimal or missing required constraint details. The server exposes 26 tools with basic input schemas but lacks documented output schemas, error handling guidance, and parameter validation constraints. Parameter descriptions are present but frequently generic (e.g., 'Selector type: css (default) or xpath' without explaining when to use each). No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite having WRITE operations. Schema validation rules, format constraints, and error recovery guidance are absent across all tools. The analyzer tool (find_page_context) hints at RAG integration but lacks documentation on expected response formats or pagination. Critically, tools modifying browser state (click_element, type_text, execute_javascript, scroll_page) lack confirmation patterns or dry-run support, increasing risk of unintended side effects in agent loops.
Analyze the current page structure and extract all interactive elements, forms, text content, and detected popups. Results are stored in RAG engine for smart element selection.
clear_session_contextwritesource verified48/100
Clear the session context and RAG engine data
click_elementwritesource verified67/100
Click an element on the page by CSS selector or XPath
close_popupswrite50/100
Close all detected popups, modals, and overlays on the page
execute_javascriptwritesource verified58/100
Execute arbitrary JavaScript code in the current page context
fill_formwritesource verified58/100
Fill multiple form fields at once
find_page_contextread onlysource verified55/100
Query RAG engine to find page context by search query. Uses semantic search to find relevant elements, content, and information from the last analyzed page.
No documented output schemas across all 26 tools. LLMs cannot plan downstream tool calls or extract chaining IDs (e.g., element_id, selector_path) without knowing what fields are returned.
Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint) for all WRITE operations. Tools like click_element, execute_javascript, type_text, scroll_page, fill_form, submit_form, close_popups, and track_action_result operate on mutable browser state but lack annotations to signal side effects to the LLM.
Add documented output schemas to all 26 tools. For analyze_current_page and find_page_context, specify: return type (array of objects vs. single object), element structure (id, selector, text, type, position, confidence), and field types. Example: {elements: [{id: string, selector: string, text: string, type: enum(button|input|link|...), x: number, y: number, confidence: number}]}.
Annotate all WRITE operations with destructiveHint and idempotentHint. Set destructiveHint=true for: click_element, execute_javascript, scroll_page, select_dropdown, fill_form, submit_form, close_popups, clear_session_context. Set idempotentHint=true for: type_text (if clear_first=true), scroll_page (idempotent by position), wait_for_element. Set readOnlyHint=true for: navigate_to_url (changes browser state but is not destructive; reconsider), all get_* and analyze_* tools.
Add confirmation/dry-run support for high-risk operations. For submit_form and execute_javascript, add optional 'dry_run' parameter (boolean, default=false). When dry_run=true, return what would happen without executing. For click_element, return a preview of the element and ask for user confirmation before executing.
Document parameter constraints in descriptions using formal syntax: 'selector (CSS selector: ^[a-zA-Z][a-zA-Z0-9_:.*\[\]() -]*$, max 500 chars). Examples: "#submit-btn", ".form-input", "[data-qa=login]". Use for elements on the current page.'. Add similar constraints for XPath, JavaScript, URL format.
Add structured error handling to all tools. Document expected error cases: (1) timeout (selector not found within wait_for_element timeout), (2) selector not found (no matching element), (3) element not clickable (covered by another element, disabled, or hidden), (4) invalid input (malformed selector, negative amount), (5) network error (if apply_action involves a navigation). For each error, state: is it retryable? Should the LLM ask the user? What should the next step be?
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 29 points across a rubric change (v1 → v2)
50/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
50
<=2025-11-25
v2
2026-03-09
F
21
-
v1
get_browser_statusread only50/100
Get the current status of the browser
get_current_urlread onlysource verified58/100
Get the current page URL
get_detected_popupsread onlysource verified52/100
Get list of popups, modals, and overlays automatically detected and their close buttons from the last page analysis.
get_element_textread onlysource verified62/100
Get text content from an element
get_page_sourceread onlysource verified60/100
Get the raw HTML source of the current page
get_page_titleread only50/100
Get the current page title
get_rendered_htmlread onlysource verified57/100
Get the rendered HTML of the current page (after JavaScript execution)
Destructive operations (click_element, execute_javascript, submit_form, clear_session_context) lack confirmation or dry-run support. Agents in retry loops or with ambiguous intent can trigger unintended clicks, form submissions, or data erasure without user confirmation.
Minimal parameter descriptions lack constraint details. 'Selector type: css (default) or xpath' is present but does not explain when to choose each, or what a valid CSS selector looks like. No regex patterns, length limits, or format guidance for CSS/XPath selectors.
No error handling guidance. Tools do not document what happens on timeout, selector not found, element not clickable, or JavaScript execution failure. LLMs cannot self-correct on failures or know whether to retry.
RAG-based tools (analyze_current_page, find_page_context, get_smart_element_selector) reference a RAG engine for semantic search but do not document: (1) what data structures the RAG indexes, (2) expected response format from semantic search, (3) how elements are stored/retrieved, (4) whether pagination is supported for large result sets.
fill_form accepts a generic 'fields' object (dictionary of field_selector: value pairs) without documenting: (1) how selectors map to form field types (input, textarea, select, checkbox), (2) expected format for different field types (e.g., date format), (3) validation errors if a selector does not match any form field, (4) whether it can handle file uploads or rich text fields.
Session management tools (track_action_result, get_session_progress, clear_session_context) lack documentation on: (1) session lifetime and scope (per-browser, per-user, global), (2) what constitutes an 'action' in action history, (3) whether sessions persist across stop_browser/start_browser cycles, (4) maximum history size or retention policy.
No pagination support documented for tools returning lists (analyze_current_page, find_page_context, get_detected_popups, get_session_progress). Large result sets could exhaust token budgets or exceed context windows. Missing 'top_k' parameter on find_page_context hint at intent but other list tools lack explicit limits.
execute_javascript accepts arbitrary JavaScript with no sandboxing, execution scope, timeout, or return value constraints documented. LLMs could accidentally break page state, trigger client-side errors, or hang the browser with infinite loops.
execute_javascript
Enhance RAG tool documentation. For analyze_current_page: document what data is indexed (element selectors, text content, type, position, visibility) and the expected response schema. For find_page_context: clarify that semantic search returns ranked elements with confidence scores; document confidence threshold and how to interpret scores. For get_smart_element_selector: explain that it returns multiple selector variants (CSS, XPath, data-* attributes) ranked by uniqueness and robustness.
For fill_form, break the generic 'fields' parameter into explicit field-type variants: 'fields_text' (input[type=text], textarea), 'fields_select' (select elements with option values), 'fields_checkbox' (checkbox elements with true/false), 'fields_radio' (radio button groups with option values). Return per-field success/failure so partial fills are clear. Document which field types support which value formats.
Document session lifecycle and scope. State: sessions are per-browser-instance; clearing one browser session does not affect others. Action history persists in memory until clear_session_context is called or the browser stops. Recommend calling track_action_result after each significant action (navigation, form submission, click) to build a recovery trail.
Add pagination (limit, offset or cursor) to list-returning tools. For analyze_current_page, add 'max_elements' parameter (default 50, range 1-500) to cap interactive elements returned. For get_session_progress, add 'limit' (default 20) and 'offset' (default 0). Return 'total_count' and 'has_more' boolean in response.
Restrict execute_javascript with constraints: add 'timeout_ms' parameter (default 5000, max 30000) to prevent hangs. Document: execution context is the current page (can access window, document, DOM); return value must be JSON-serializable (no functions, DOM nodes); undefined/null returns are converted to empty string. Add error handling for syntax errors, timeout, and non-serializable returns.
Add tool chaining hints to descriptions. E.g., navigate_to_url: 'After navigating, call analyze_current_page to discover interactive elements.' find_page_context: 'Returns elements with selectors; pass selector to click_element, type_text, or get_element_text.' get_smart_element_selector: 'Returns CSS/XPath selectors for detected elements; use with click_element or type_text.'
Define idempotence clearly for each tool. click_element is NOT idempotent (repeated clicks trigger multiple actions). type_text with clear_first=true is idempotent (same final state regardless of repetition). scroll_page is idempotent (scrolling to position 500px twice lands at the same place). Document this so agents know which tools are safe to retry blindly.