Stealth browser automation via MCP. Camoufox (custom Firefox) with zero Chrome DevTools Protocol exposure and real OS-level mouse/keyboard input via PyAutoGUI — undetectable by bot detection. Passes Cloudflare, CreepJS, BrowserScan, Pixelscan, and major bot detectors.
This server provides 30 tools with reasonable coverage for browser automation. Naming is mostly verb-based and clear (goto, click, get_text, etc.). However, critical gaps exist: many tools lack explicit input schemas in the source code (only inferred from parameter lists), descriptions are present but terse and often lack context about when/why to use each tool, and there is no documented output schema for any tool. The run_script meta-tool is well-described with control flow documentation, but individual action steps (goto, click, fill, etc.) have minimal parameter descriptions and no indication of expected return values. Error handling is not visible in the source. Security considerations (e.g., secrets, injection) are not documented at the tool level.
Calibrate OS-level mouse coordinates for system_click. Must be called before using system_click. Available within run_script action steps.
Check a checkbox. Available within run_script action steps.
Click element by CSS selector or XPath. Fast and reliable. Available within run_script action steps.
Detect documented challenge widgets and conservative generic challenge cues in the current page. Read-only: never clicks, solves, or enters frames. Returns `absent`, `present`, or `unknown`, plus bounded vendor, confidence, location, and sanitised frame-path evidence.
Execute JavaScript and return result. Available within run_script action steps.
Set input field value instantly (clears first, no keystrokes). Available within run_script action steps.
No documented output schemas for any tool. LLMs cannot predict what fields to expect in responses, breaking tool chaining and forcing context-wasting discovery. E.g., get_element returns 'tag, text, attributes, bounding box, visibility' but exact JSON structure is undocumented.
Parameter descriptions are minimal (under 50 chars) and lack context. E.g., 'get_element' accepts a 'selector' param but does not specify whether it accepts only CSS selectors or also XPath, or what happens if multiple elements match. Many params (e.g., in check, uncheck, eval) have no descriptions at all in the source.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | 2026-07-28+ | v2 |
Get selected computed CSS properties for one element. Available within run_script action steps.
Get one element's tag, text, attributes, bounding box, and visibility. Available within run_script action steps.
Get a bounded list of matching element summaries. Available within run_script action steps.
Get full HTML source of the current page. Available within run_script action steps.
Find all interactive elements (buttons, links, inputs). Returns list of elements with x, y, width, height, text, selector. Available within run_script action steps.
Get current URL, title, ready state, viewport, document, and scroll data. Available within run_script action steps.
Get visible text from current page (truncated to 10,000 chars). Available within run_script action steps.
Report whether configured virtual camera and microphone sources are active for getUserMedia(), plus dynamic-source state. Available within run_script action steps.
Navigate to a URL. Available within run_script action steps.
Click at absolute screen coordinates (or current position). Available within run_script action steps.
Press a keyboard key. Available within run_script action steps.
Reload current page. Available within run_script action steps.
Run multiple browser actions as a single atomic script. All steps execute sequentially on the SAME browser instance under a single request lock. This is the only safe way to do multi-step workflows in a load-balanced cluster. Each step is a dict with "action" and its parameters. Add "output_id" to any action step to collect its result in the response outputs dict. Steps can also use control nodes: {"if": {"condition": ..., "then": [...]}}, {"repeat": {"count": 1, "steps": [...]}}, and {"while": {"condition": ..., "max_iterations": 1, "steps": [...]}}. Conditions support element, text, URL-glob, JavaScript-boolean, and prior-output checks.
Take and save a screenshot to disk. Available within run_script action steps.
Take a screenshot of the browser viewport. Available within run_script action steps.
Select option(s) from a select element. Available within run_script action steps.
In opt-in dynamic-media mode, select an existing contained file for future and already-acquired virtual tracks. Available within run_script action steps.
Pause execution for a specified duration. Available within run_script action steps.
Click at viewport coordinates using OS-level mouse. Undetectable but requires calibrate first. Available within run_script action steps.
Type into element (keystroke simulation). Available within run_script action steps.
Uncheck a checkbox. Available within run_script action steps.
In opt-in dynamic-media mode, store a bounded base64 audio or video file under VIRTUAL_MEDIA_DIR and optionally activate it. Available within run_script action steps.
Wait for an element to appear. Available within run_script action steps.
Wait for the page to navigate. Available within run_script action steps.
No error handling guidance. Tools do not indicate what errors might occur, how to recover, or whether an error is retryable. E.g., goto() fails if network is unavailable, but the LLM has no guidance on retry strategy or fallback actions.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in source. LLMs cannot distinguish safe (read-only) tools from risky (destructive) ones, increasing risk of accidental data corruption via unintended tool composition.
Ambiguous naming for similar tools. 'click', 'system_click', 'mouse_click' all perform clicks but with different implementations; the descriptions do not make the trade-offs clear (speed vs. detectability vs. OS-level interaction). LLMs will struggle to pick the right variant.
Tool descriptions lack WHEN/WHY context. E.g., 'eval' says it 'Execute JavaScript' but does not explain when to use it vs. built-in get_text/get_html/get_element tools. Many tools exist for the same conceptual goal (getting element info: get_element, get_elements, get_interactive_elements, eval), but the docs do not guide selection.
Underspecified parameter constraints. E.g., 'wait_until' accepts string values but enums are not declared in the source (description lists 'domcontentloaded', 'load', 'networkidle' as free text, not enums). 'limit' in get_elements is described as '1-100, default 20' but no min/max constraint is visible in the schema.