Browser automation where an LLM plans and Jev (Typesafe System One) makes every per-step decision. Library, CLI and MCP server.
jev-browser demonstrates solid definition quality with comprehensive schemas, detailed descriptions, and clear action semantics. All 8 tools are explicitly registered with input schemas and descriptions. Tool names follow verb_noun convention and clearly indicate their action. Descriptions are substantial (100-300+ chars) and include context about decision models, status values, and usage patterns. Parameter descriptions are present and detailed. However, there are gaps in output schema documentation and some parameter constraints could be more formally specified. Error categorization and recovery guidance are implicit but not formalized.
Perform one action on an element number from the latest browser_snapshot (or from browser_do candidates). No decision model involved. Numbers are matched to the current page; if the element is gone, take a new snapshot. Confirm/prompt dialogs the action opens are dismissed unless accept_dialog is true.
Ask a yes/no question about the current page. Returns the probability of yes (≥0.85 reliable yes, ≤0.15 reliable no, in between: look yourself with browser_snapshot).
Ask which of several options is true of the current page. Options must be provided (Jev cannot generate text). Returns the choice and a probability per option.
Close the browser session. The next call starts a fresh one.
Work toward ONE observable outcome on the current page. A fast decision model (Jev) picks each element/action/value; it cannot write text or plan, so: Write one outcome per call ("Log in", "Add the Backpack to the cart", "Open the Pull requests tab"); split ordered sub-tasks into separate calls. Put every string to type, option to pick or file path to upload in `values`, with meaningful keys ({email, password}). Make open-ended goals measurable ("until at least 3 new results are shown"). Statuses: done | likely_done (Jev is unsure the goal is met: verify with browser_check or browser_snapshot before moving on) | needs_login (sign-in wall and no credentials given: ask the user to log in, e.g. with JEV_BROWSER_HEADED=1 and JEV_BROWSER_PROFILE, or pass credentials in values) | needs_confirmation (next click looks irreversible: re-call with allow_irreversible=true only if the user wants it) | error (page shows an error) | blocked | stuck | ambiguous (see candidates; use browser_act) | max_actions. After steps with side effects, use browser_check to confirm nothing unintended changed.
Output schemas not explicitly documented in tool definitions. While tools return structured results (evident from bench/context-cost.mjs showing JSON output with status, url, title, actions, done_score, etc.), the MCP tool registration does not include outputSchema declarations. This forces LLMs to infer the output structure from experimentation.
browser_act: 'action' parameter is an enum with 10 values, but the description does not list all valid actions or explain what each does. LLMs must infer semantics from names alone (e.g., does 'select' select a checkbox or a dropdown option?). Descriptions for each action type would improve clarity.
browser_do: status values ('done', 'likely_done', 'needs_login', 'needs_confirmation', 'error', 'blocked', 'stuck', 'ambiguous', 'max_actions') are documented in the description as prose, not as an enum constraint or structured list in the schema. This makes it harder for tools calling browser_do to parse and react to status values programmatically.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | 2026-07-28+ | v2 |
Navigate the browser session to a URL and wait until the page settles. Starts the browser on first use.
Screenshot of the current page (viewport unless full_page).
Compact view of the current page: visible text and numbered interactive elements. Use it to take over when browser_do is ambiguous or stuck, then act with browser_act.
Parameter 'values' in browser_do is typed as 'object' with no schema for its contents. Description says 'Named strings Jev may type/select/upload' but does not specify valid keys or value types. This allows arbitrary objects and leaves LLMs to guess what keys are expected.
Error handling is implicit. Tools return status values and optional 'info' field in responses, but there is no formal error categorization or recovery guidance in the schema. Tools like browser_do list status values that signal failures (error, blocked, stuck) but do not prescribe what the LLM should do next.
browser_act: 'element' parameter is required for most actions but is not required for scroll/back/page-level press_key. The schema does not declare this conditional requirement. Description documents it, but a schema constraint (oneOf, conditional schema) would be more precise.