Local MCP server that gives any AI agent safe cross-OS desktop control — the fallback execution layer for when APIs, CLIs, and direct integrations aren't available. Works with any tool-calling model (Claude, GPT, Gemini, Llama) on Windows, macOS, and Linux.
clawdcursor demonstrates strong definition quality across 28 tools with comprehensive schemas and descriptions. All tools have explicit descriptions and input schemas. Most tools follow verb_noun naming conventions (read_screen, list_windows, invoke_element, compile_ui). Descriptions are detailed and action-oriented, averaging ~150-250 characters and providing clear guidance on WHEN to use each tool. However, some tools lack sufficient parameter-level documentation, and a few parameter descriptions are overly terse. Output schemas are largely undocumented, tools define inputs well but rarely document what fields are returned. Error handling guidance is present but inconsistent across tools. Security considerations (credentials, sanitization) are not explicitly documented in tool descriptions. Overall, the server demonstrates mature tool design with room for improvement in schema completeness and error documentation.
Collapse an accessibility element.
Expand an accessibility element.
Get the value of an accessibility element.
Select an accessibility element.
Toggle an accessibility element.
Run SEVERAL known next actions in ONE call instead of one per turn (saves round-trips). steps = JSON array of {"name","args","precheck"?}. Each step is a normal tool call (e.g. {"name":"type","args":{"text":"hi"}}). Optional "precheck" is a precondition re-checked against live state before the step: {"window":"notepad"} (that window must be focused) or {"element":"Send"} (that a11y element must exist). An `expect` assertion array inside a step's args is verified after that step (DEVIATION halts the batch). The batch STOPS at the first failed precondition, safety stop, step error, or DEVIATION and returns a per-step trace so you continue from real state. el_NN refs are only safe up to the first screen-changing step — after that the screen moved; target later steps by name. Use it for deterministic stretches (open -> focus -> type -> save). Do NOT put perception-only reads or terminal tools (done/give_up) in a batch.
Output schemas are not documented for any tool. While input schemas are comprehensive, tools do not declare what fields they return, forcing LLMs to infer return structure and preventing reliable downstream tool chaining.
Accessibility-related tools (a11y_expand, a11y_collapse, a11y_toggle, a11y_select, a11y_get_value) have minimal descriptions (~65 chars) that lack guidance on WHEN to use them instead of invoke_element. LLMs cannot distinguish which tool is appropriate without better context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 68 | 2026-07-28+ | v2 |
Click an element in browser. Converted to cdp_click on MCP surface.
Connect to a browser for debugging/automation. Converted to cdp_connect on MCP surface.
Navigate browser to a URL. Converted to navigate_browser on MCP surface.
Read page context from browser (DOM listing path). Converted to cdp_page_context on MCP surface.
Type text in browser. Converted to cdp_type on MCP surface.
Click at screen coordinates. Converted to mouse_click on MCP surface.
Compile a UI map with targeted perception for a specific purpose.
Drag from one coordinate to another. Converted to mouse_drag on MCP surface.
Find an action button matching an intent.
Find an input field matching a purpose.
Focus an accessibility element by name.
Get the state of an accessibility element.
Click/activate a UI element by its accessibility name. MORE RELIABLE than coord clicks — use this when the snapshot shows a named target.
Press a key combination. Converted to key_press on MCP surface.
List visible top-level windows with title, process, and bounds. Useful when the active window is wrong or missing.
START HERE — cheapest perception. Read the accessibility tree of the focused window: buttons, inputs, text elements with coordinates. The snapshot is auto-attached each turn; call this again only when you expect the screen changed since the last turn. If the tree is empty, escalate to read_text (OCR) next, then screenshot only as a last resort.
Use OCR to read text from the screen. Converted to ocr_read_screen on MCP surface.
Capture a screenshot. Converted to desktop_screenshot on MCP surface.
Scroll in a direction at a coordinate. Converted to mouse_scroll on MCP surface.
Set the value of an accessibility field by name.
Type text. Converted to type_text on MCP surface.
Verify assertions about the current screen state.
browser_connect and browser_read descriptions are vague. browser_connect says 'Connect to a browser for debugging/automation' but does not explain how this affects subsequent browser calls or whether it is required. browser_read says 'Read page context from browser (DOM listing path)' without clarifying what 'DOM listing path' means.
screenshot, read_text, and browser_read have no input parameters documented in their schemas. While these are read-only tools, the absence of any schema (even an empty properties object with description) is inconsistent with best practices.
Error handling guidance is absent from tool descriptions. No tool explains what happens on failure, whether errors are retryable, or what the LLM should do next. For example, invoke_element does not say what error occurs if the element is not found or how to recover.
verify tool description lacks detail on what assertion types are accepted and what DEVIATION means. The batch tool documentation mentions assertions but verify does not explain the assertion schema or available operations.