Production-grade AI agent framework with RAG, memory, tools, and multi-model support. Includes tool management, guardrails, security policies, and toolkit system for aggregating tools with shared dependencies.
Strong browser automation toolkit with consistent naming and clear descriptions. All 28 tools follow verb_noun convention (browser_*). Descriptions are well-written and action-oriented (avg ~120 chars, well within 10-1024 range). Input schemas are complete with type definitions for all parameters. However, output schemas are entirely absent, no documentation of what these tools return, forcing LLMs to guess at response structure. Error handling guidance is minimal. Tool composition is excellent (single responsibility, chainable refs). Security is solid (no secrets exposed). Missing tool annotations (readOnlyHint/destructiveHint) despite clear risk stratification. No evidence of pagination for snapshot/source tools that could return large data.
Clear the contents of an input field or textarea.
Click an element using a ref (e.g. "e1") or CSS selector (e.g. "button#submit"). Use browser_snapshot() first to get element refs.
Click the first element whose visible text contains 'text'. More reliable than selectors on dynamic sites. Optionally restrict to a tag: browser_click_by_text("Sign in", "button")
Click an element only if it is currently visible. Safe to call on conditionally-shown elements like popups, cookie banners, etc.
Drag an element to another element using native drag-and-drop. Use for reordering lists, sliders, Kanban boards, and canvas apps.
Execute JavaScript in the page context and return the result. Example: browser_execute_js("document.title")
Output schemas completely missing, no documentation of return types for any tool. LLMs cannot predict response structure or chain tools effectively without knowing what fields are returned.
No tool annotations (readOnlyHint/destructiveHint/idempotentHint) despite clear risk stratification in metadata. browser_click, browser_type, browser_drag are marked WRITE but lack destructiveHint. Read-only tools like browser_get_text lack readOnlyHint.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Fill multiple form fields at once. Each field: {ref, type, value}. Checkboxes use setChecked(), others use fill(). Example: browser_fill_form([{"ref":"e1","type":"text","value":"John"}, {"ref":"e2","type":"checkbox","value":true}])
Return the value of an HTML attribute on an element. Example: browser_get_attribute("e1", "href") or browser_get_attribute("a.logo", "href")
Return situational snapshot: URL, title, scroll position, viewport size. Call this to understand page context before deciding your next action.
Return the full page HTML source (capped at 20 000 chars).
Return the visible text content of an element. Use a ref (e.g. "e1") from the last snapshot, or a CSS selector (e.g. "h1", "#main"). Defaults to the entire page body.
Return the current page title.
Return the current page URL.
Navigate to the previous page in browser history.
Navigate forward in browser history.
Hover the mouse over an element. Reveals dropdown menus, tooltips, and hover-triggered content.
Check if an element is currently visible on the page. Returns "true" or "false".
Navigate to a URL. REQUIRED: url (str) — must be a full URL including the scheme, e.g. 'https://example.com'. Returns the final URL and page title.
Press a keyboard key on the currently focused element (no selector needed). Use for Enter, Tab, Escape, Backspace, ArrowDown, ArrowUp, etc. Example: browser_press_key("Enter"), browser_press_key("Escape")
Send keystrokes to a specific element (requires ref or CSS selector). Use "Enter" for Enter, "Tab" for Tab. Example: browser_press_keys("e1", "Enter") If you don't have a selector, use browser_press_key(key) instead.
Reload the current page.
Take a screenshot of the current page and save it to a file. Returns the file path. Use this to visually inspect the page.
Select an option from a <select> dropdown by its visible text. Example: browser_select_option("e3", "United States")
Set files on a file input element. Example: browser_set_input_files("e5", ["/path/to/file.pdf"])
Set an element's value directly — works for sliders and range inputs. Example: browser_set_value("e4", "75")
Return a role-based accessibility view of the page with element refs. Each interactive element gets a ref like [ref=e1] that you can use with other browser tools (e.g. browser_click("e1"), browser_type("e2", "hello")). This is the BEST way to understand page structure. Use it BEFORE interacting.
Clear an input field and type text into it. Use a ref (e.g. "e2") or CSS selector. Example: browser_type("e2", "hello@example.com")
Type text with human-like 75 ms delays between keystrokes. Use on sensitive form fields to avoid bot-detection triggers.
browser_get_source documents a 20,000 char cap but offers no pagination mechanism or continuation hint. Large pages truncated without guidance on how to retrieve remaining content.
browser_fill_form parameter 'fields' is an array but individual field object schema is only shown in description text, not formally in the schema. Type 'checkbox' is mentioned in description but field type constraint is not formalized.
No error handling guidance in any tool description. E.g., browser_click does not explain what happens if element is not found, if it's disabled, or if navigation is blocked.