Static source inference · medium confidence · detected: stateful session
Deprecated protocol patterns detected
Summary
Computer Use MCP exposes a single, complex tool 'computer_call' with a well-structured union-based input schema covering 10 action types (click, double_click, triple_click, drag, keypress, move, screenshot, scroll, type, wait). The tool definition is explicit and visible in src/index.ts with comprehensive parameter validation. The schema uses JSON Schema oneOf patterns correctly to enforce mutually exclusive action types. However, the tool description is brief (72 chars) and lacks explicit state-modification language, recovery guidance, and composition hints. The server correctly implements HTTP transport (StreamableHTTPServerTransport) and modern request-per-session patterns. Major gaps: no tool annotations (readOnlyHint/destructiveHint/idempotentHint), no pagination/result limiting on screenshot output, no error classification guidance, and missing per-action timeout/retry semantics.
Tools (1)
computer_callwritesource verified71/100
Execute a computer-use action (click, type, screenshot, drag, scroll, keypress, move, wait, double_click, triple_click) on a virtual browser
Tool description lacks explicit state-modification language and action scope clarity. Current description (72 chars) does not answer: What are the side effects? When should this tool be used vs. others? Are actions idempotent or destructive? This forces LLMs to infer intent from action types alone.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) declared in tool definition. The 'computer_call' tool performs both read-only actions (screenshot, move) and destructive/side-effect actions (click, type, drag, keypress). Without annotations, LLMs cannot reason about safety, rollback, or retry semantics for each action type.
Screenshot action returns a Buffer without documented schema or size constraints. No pagination/result limiting guidance for large screenshots. If an agent calls screenshot repeatedly in a loop, token consumption is unbounded. Baseline guidance: document max image size, format, and offer optional quality/resolution parameters.
computer_call
Recommendations
Expand tool description from 72 to 150-200 chars: clarify that this tool controls a virtual browser via actions like clicking, typing, and taking screenshots. Add: 'Side effects: click/type/drag modify the browser state; screenshot is read-only. Idempotent: screenshot. Non-idempotent: click, type, keypress, drag (may have cumulative effects). Use screenshot between actions to verify state changes.' This enables LLMs to reason about retry safety.
Add tool annotations to the tool definition: declare readOnlyHint=true for screenshot and move actions; destructiveHint=true (or at least warn) for click, type, keypress, drag (they modify browser state and are non-idempotent). If official MCP SDK supports tool.annotations, integrate it; otherwise, document in description and error responses.
Document error classification in description: 'Common errors and recovery: (1) Timeout: browser unresponsive; retry or take a screenshot to check state. (2) Invalid coordinates: x/y outside viewport; take screenshot to see current layout. (3) Element not found: the target element may have moved or been removed; call screenshot to recheck. (4) Browser crashed: restart the session (outside this tool's scope).'
Expand 'button' parameter description: 'Mouse button to click: left (default, left-click), right (context menu), wheel (scroll wheel, rarely used), back/forward (browser navigation buttons). Most use cases require left; right opens context menus.'
For scroll action, clarify null coordinates: 'x and y: coordinates of the scroll anchor (optional; null = scroll at viewport center). scroll_x and scroll_y: distance to scroll in pixels (positive = right/down, negative = left/up). Example: {type: "scroll", x: null, y: null, scroll_x: 0, scroll_y: 500} scrolls down 500px from center.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Stateful initialize / Mcp-Session-Id (removed; protocol is stateless) - make each request self-contained
No error classification or recovery guidance in tool description. LLMs do not know which errors are retryable (network timeout, browser crash) vs. user-fixable (invalid coordinates, element not found) vs. fatal (sandbox exhaustion). This forces blind retries and wastes agent compute.
Parameter 'button' in click action uses enum ['left', 'right', 'wheel', 'back', 'forward'] with minimal description ('Mouse button to click'). Description does not clarify what 'wheel' does (scroll?), what 'back'/'forward' mean in a virtual browser context, or browser compatibility. This invites hallucinated or invalid button values.
Scroll action parameters allow null coordinates ('x': {'type': ['number', 'null']}) without clear documentation of what null means or when it should be used. This ambiguity invites LLMs to pass null incorrectly and encounter silent failures.
Wait action has no parameters (beyond type) but no description of duration, default behavior, or semantics. Does 'wait' sleep for a fixed time, or does it wait for a condition? How long? This is under-specified.
No batch/concurrent action support documented. If an agent wants to perform multiple actions (e.g., type text, then click submit), it must issue sequential computer_call invocations. This wastes latency and invites mid-chain failures. Consider adding support for action sequences or batch mode.
computer_call
For wait action, add a duration parameter: 'duration_ms: how long to wait in milliseconds (default 1000, range 100 - 30000). Use this to let async operations complete (e.g., page load, animation finish) before the next action.'
Consider adding a batch_actions variant or multi-action mode: allow the tool to accept an array of actions, execute them sequentially, and return a single screenshot at the end. This reduces round-trips, latency, and mid-chain failure risk. Document: 'For efficiency, batch related actions (e.g., type text + click submit) into a single call; the server executes them atomically and returns the final screenshot.'
Add per-action timeout and retry guidance in error responses: when an action times out, suggest calling screenshot first to check browser state, then retrying if the state hasn't changed. Return clear, actionable error messages like 'Click at (500, 300) timed out: browser may be unresponsive. Call screenshot() to verify state before retrying.'
Document screenshot output format explicitly: 'Returns a base64-encoded PNG or JPEG image (format controlled by server config). Default quality: 80%; max resolution: 1920×1440. To capture high-detail regions, zoom the browser or call multiple screenshots of adjacent regions.'
Add composition hints: 'Typical workflow: (1) screenshot → inspect current state. (2) identify target element coordinates. (3) computer_call with action (click, type, etc.). (4) screenshot → verify change. Repeat until goal achieved. Note: actions are not idempotent; if you need to recover from a mistake, you may need to navigate back or reload the page.'
Validate input ranges early and return actionable errors: if x/y coordinates are outside the viewport, return 'Coordinates (2500, 3000) are outside viewport (max 1920×1440). Take a screenshot to see current layout.' instead of a silent or generic error.
Document rate limits and cooldowns: 'To avoid overwhelming the browser, insert small delays (100 - 200ms) between rapid actions. Use wait action for explicit pauses. Sustained action rates >10/sec may cause browser instability.'
Consider adding a 'dry_run' or 'validate' flag (if feasible) to let agents preview action effects without executing: 'dry_run: true → describe what the action would do without executing it.' This prevents accidental clicks on destructive buttons.