The server defines 16 tools with reasonable naming (verb-first convention) and structured schemas using Zod. However, descriptions are terse (averaging ~40-80 chars, below the production baseline of 194 chars), lack context on when/why to use each tool, and omit key details like side effects and error recovery guidance. All tools have input schemas with typed parameters and descriptions, which is a strength. Parameter descriptions are adequate but minimal. Output schemas are not explicitly documented, responses are hardcoded as 'text' or 'image' content types without structured field documentation. No error handling guidance is present; failed tool calls would return generic errors. The tool set is well-composed (each does one thing, good verb-noun naming), but the descriptions need substantial expansion to meet production LLM optimization baselines.
Tools (16)
backwritesource verified68/100
Go back in browser history.
clickwritesource verified75/100
Move the pointer to (x, y) and click.
double_clickwritesource verified72/100
Move the pointer to (x, y) and double-click.
dragwritesource verified73/100
Drag along a path of points. The first point is mouse-down, the last is mouse-up.
forwardwritesource verified68/100
Go forward in browser history.
get_current_urlread onlysource verified70/100
Get the current browser URL.
get_dimensionsread onlysource verified72/100
Get the screen or viewport dimensions as [width, height].
Tool descriptions are below production baseline. Average ~55 chars vs. baseline 194 chars. Descriptions lack context on when/why to use tools, prerequisites, side effects, and recovery guidance.
No output schemas documented. Tools return hardcoded 'text' or 'image' content types, but LLMs cannot see what fields or structure to expect in responses. This forces agents to infer structure and plan downstream calls blindly.
Expand tool descriptions from ~40-80 chars to 100-200 chars. Include: WHAT the tool does, WHEN to use it (vs. similar tools), what state it modifies, and any side effects. Example: 'Click on a screen element at (x, y). Use after a screenshot to interact with detected UI. Specify button (left/middle/right). Returns success or error if coordinates are out of bounds or element not found.'
Document output schemas explicitly. For 'screenshot', document: 'Returns base64-encoded PNG image data (image/png MIME type). Dimensions match get_dimensions(). Use with screenshot_region for partial captures.' For 'click', document: 'Returns {success: bool, error?: string}. On failure (coordinates out of range, element unresponsive), error message indicates cause.'
Add error recovery hints to tool descriptions. E.g., 'click' → '...If click fails due to stale coordinates, call screenshot() to verify current UI state before retrying.' This guides agent recovery.
Document side effects explicitly. Prefix descriptions of state-modifying tools (click, type, keypress, goto, back, forward, drag) with 'This modifies the active application state. ' so LLMs know these are not idempotent.
Add constraints to coordinate parameters. E.g., 'x: number (0 to screen width). Y coordinate must be within screen bounds; use get_dimensions() first if needed.' This prevents out-of-range errors.
For 'screenshot_region', document coordinate ordering: 'x1, y1 is the top-left corner; x2, y2 is the bottom-right corner. Coordinates must form a valid bounding box (x1 < x2, y1 < y2).'
Score history
Overall score trend
↑ 11 points across a rubric change (v1 → v2)
61/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
61
2026-07-28+
v2
2026-03-09
D
50
-
v1
72/100
Get the environment type: "linux", "mac", "windows", or "browser".
gotowritesource verified71/100
Navigate the browser to a URL.
keypresswritesource verified73/100
Press a key combination. For combos like Ctrl+C, pass ["ctrl", "c"].
movewritesource verified70/100
Move the pointer to (x, y) without clicking.
screenshotread onlysource verified68/100
Take a screenshot and return it as a base64-encoded PNG.
screenshot_regionread onlysource verified72/100
Take a screenshot of a specific region and return it as a base64-encoded PNG.
scrollwritesource verified74/100
Move the pointer to (x, y) and scroll by the given deltas.
typewritesource verified67/100
Type the given text as keyboard input.
waitread onlysource verified68/100
Wait for a duration in milliseconds.
No error handling guidance. Tools have no error-recovery hints. When a click fails (e.g., element not found), LLMs receive no guidance on what caused it or what to try next.
Many tools lack context on side effects. Tools like 'click', 'type', 'keypress', 'goto', 'back', 'forward', 'drag' are state-modifying, but descriptions do not explicitly state this. LLMs need to know which calls are safe to retry.
Parameters lack context on constraints and expected input ranges. For example, 'x' and 'y' in 'click' have no bounds documented (e.g., 'Must be within screen dimensions'). 'button' enum is good, but others could benefit from min/max or format hints.
clickdouble_clickscrollmovedragscreenshot_region
Add 'typing' context to 'type' tool. E.g., 'Type text as keyboard input into the active text field. Use after clicking on an input to focus it. Special characters (newline, tab) must be passed as \n, \t.'
Document 'wait' usage. E.g., 'Pause execution for N milliseconds. Use between rapid actions to allow UI to render, animations to complete, or network requests to settle (e.g., wait 500ms after goto before taking a screenshot).'
For browser tools (goto, back, forward, get_current_url), clarify environment detection. E.g., 'goto only works in browser environment (get_environment returns "browser"). In native/desktop environments, desktop automation tools must be used instead.'
Add parameter dependencies to descriptions. E.g., 'scroll' → 'scroll_x and scroll_y: horizontal/vertical deltas in pixels. Pass 0 for either to skip that direction. At least one must be non-zero.'
Separate concerns where possible. 'screenshot' and 'screenshot_region' are well-separated; consider whether 'drag' could be split into 'drag_from_to' with clearer semantics (start_x, start_y, end_x, end_y) vs. an array of points.