Wayland MCP server providing tools for mouse control, keyboard input, screenshot capture and analysis, and action chaining
Wayland MCP has 7 tools with moderate quality issues. Tool names follow verb_noun convention (move_mouse, click_mouse, etc.) which is good. However, descriptions are present but inconsistent in quality and depth. Input schemas are visible and properly typed (integer, string, boolean), but several tools lack comprehensive parameter descriptions or output schema documentation. Error handling returns basic success/error dicts but lacks actionable recovery guidance. The server does not implement tool annotations (readOnlyHint, destructiveHint) despite having clearly destructive and read-only tools. capture_screenshot lacks an input schema definition (filename has a default but no explicit schema shown). No per-tool output schemas are documented for LLMs to plan downstream calls.
Analyze screenshot using VLM.
Capture screenshot with measurement rulers.
Simulate left mouse click at current position.
Compare two images using VLM.
Perform drag operation between coordinates.
Move mouse to specified screen coordinates.
Scroll vertically (positive=up, negative=down). Each unit represents one notch on the scroll wheel (120 units = high-definition scroll). Typical values range from 2-3 to 5-10 for normal scrolling.
Missing tool annotations (readOnlyHint, destructiveHint). Tools like move_mouse, click_mouse, drag_mouse, scroll_mouse are WRITE operations that modify system state, but no @destructiveHint annotation declares this. compare_images, analyze_screenshot, capture_screenshot are READ operations but lack @readOnlyHint. LLMs cannot distinguish safe operations from risky ones without these hints.
Incomplete output schema documentation. Tools return dicts with {'success': bool, 'error': str} but this structure is not formally documented in tool schemas. analyze_screenshot returns a string directly with no type annotation in the schema. LLMs cannot plan downstream operations without knowing return types and fields.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 9 | - | v1 |
click_mouse has NO input parameters but lacks explicit schema definition. The @mcp.tool() decorator shows Input: {} but no JSON Schema is visible. This violates the schema visibility requirement, if a tool accepts no params, the schema must still be formally present.
Error messages lack recovery guidance. All tools return {'success': False, 'error': str(e)} with raw exception text. Per pattern:recovery-guide, errors should be actionable: 'Failed to click: device not found. Verify Wayland is running and /dev/input is accessible.' Users/agents cannot act on bare exception strings.
capture_screenshot description is too brief (36 chars: 'Capture screenshot with measurement rulers.'). Below recommended 50-200 char range. Lacks explanation of when to call, what the output contains, or how 'measurement rulers' are used.
compare_images and analyze_screenshot descriptions are generic and lack context. 'Compare two images using VLM' and 'Analyze screenshot using VLM' do not explain the return format, when to use each tool, or what fields the VLM output contains. LLMs cannot plan tool selection without these details.
No pagination or result limits. If screenshot analysis returns free-text responses, there is no mechanism to truncate large outputs or warn the agent about context size. Large VLM responses could exhaust context window.