MCP server for browser automation with Camoufox anti-detection capabilities
This MCP server has 22 tools with reasonable coverage of browser automation tasks. Most tools have descriptions and input schemas present, but quality is inconsistent. Many descriptions are adequate but lack the LLM-optimization guidance required by patterns. Parameter descriptions vary widely, some are detailed, others are minimal. Several tools combine multiple concerns or lack clear output documentation. Schema definitions are present and properly typed, but no tool annotations (readOnlyHint, destructiveHint, idempotentHint) are declared. Error handling is referenced in the assessment (errorReporting=true) but not visible in the source code snippets provided. Overall, the server demonstrates basic quality but falls short of production standards in description depth, output schema documentation, and composition clarity.
Perform click on a web page element.
Close the browser. This will terminate the browser instance and clean up all resources. A new browser will be launched on the next navigation command.
Configure proxy before launching browser. Call BEFORE browser_navigate. If browser running, call browser_close first.
Returns all console messages.
Perform drag and drop between two elements.
Evaluate JavaScript expression on page or element.
Upload one or multiple files. This tool should be called when a file chooser dialog appears. If paths is omitted, the file chooser is cancelled.
Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint). Tools like browser_navigate, browser_click, browser_type are marked with Risk labels (WRITE, READ_ONLY, REVERSIBLE) in the input but no corresponding MCP tool annotations are declared. This prevents agents from understanding state-modification semantics and safety implications without extensive trial-and-error.
Descriptions lack LLM-optimization depth. Many tool descriptions (e.g., browser_hover: 'Hover over an element on the page.', browser_resize: 'Resize the browser window.') are under 50 characters and do not explain WHEN to use them, what prerequisites exist, or what downstream steps might follow. Baseline for A+ tools is 50 - 200 chars with clear intent guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Fill multiple form fields.
Handle a dialog (alert, confirm, prompt, beforeunload).
Hover over an element on the page.
Install Camoufox browser. Call this if you get an error about the browser not being installed. This will download and install the Camoufox browser binary.
Navigate to a URL.
Go back to the previous page.
Returns all network requests since loading the page.
Press a key on the keyboard.
Resize the browser window.
Select an option in a dropdown.
Capture accessibility snapshot of the current page. This provides a structured view of the page content that is better for understanding page structure than screenshots.
List, create, close, or select a browser tab.
Take a screenshot of the current page or element.
Type text into an editable element.
Wait for text to appear, disappear, or a specified time to pass.
No documented output schemas for any tool. While input schemas are well-defined with types and descriptions, there is no visible documentation of what each tool returns (return type, fields, structure). This forces LLMs to guess what fields are available for downstream tool calls, increasing chaining errors.
Composition concern: browser_snapshot returns 'a structured view of the page content' but no description of what that structure contains or how agents should parse it. Similarly, browser_console_messages and browser_network_requests return arrays but offer no guidance on result limits, pagination, or field mappings. Large unfiltered results could blow context windows.
browser_click, browser_type, browser_drag, browser_select_option require both 'element' (human-readable description) and 'ref' (exact reference) parameters. No description clarifies the relationship or when to use which. Does the agent need to call browser_snapshot first to get 'ref'? This undocumented dependency invites errors.
browser_evaluate accepts either '() => { code }' or '(element) => { code }' function syntax but description does not explain how to distinguish or what happens if both are provided. This is an undocumented parameter relationship that could lead to misuse.
browser_wait_for has three parameters (text, text_gone, time) that appear mutually exclusive or dependent, but no description clarifies their interaction. Can all three be provided? Must one be set? This ambiguity forces agents to guess.
browser_tabs action parameter uses enum ['list', 'new', 'close', 'select'] but 'index' is optional and its behavior depends on action. Description does not state: 'index is required for close/select, ignored for list/new' or vice versa. Undocumented parameter dependencies.
browser_fill_form uses a complex nested object structure for 'fields' array (each with name, ref, type, value) but no examples or clarification of how 'type' maps to 'value' interpretation. E.g., for type='checkbox', is value 'true'/'false' or '1'/'0'? This ambiguity invites failures.
browser_file_upload description states 'If paths is omitted, file chooser is cancelled' but does not clarify whether omitting 'paths' (setting it to null or empty array) are equivalent, or if one is required. This invites ambiguous agent behavior.