Model Context Protocol server for Wayland desktop automation with VLM analysis, mouse/keyboard control, and action chaining
Wayland MCP provides 7 tools with inconsistent quality across dimensions. Naming is action-verb based and mostly clear (move_mouse, click_mouse, drag_mouse, scroll_mouse, capture_screenshot, compare_images, analyze_screenshot). However, critical gaps exist: (1) Several tools lack input schema visibility or complete parameter definitions, compare_images and analyze_screenshot show parameters but no explicit type constraints visible in source; (2) Description quality varies widely, some are clear and actionable (scroll_mouse: 161 chars with specifics on units), others are minimal (click_mouse: 61 chars, no context on what happens); (3) No output schemas documented anywhere, LLMs cannot know what fields to expect from tools; (4) Error handling is absent, no guidance on what to do if a tool fails; (5) Security concerns: tools directly control desktop input/output without documented permission gates or audit logging. The server implements desktop automation tooling (a legitimate use case), but the definitions lack LLM-readiness patterns. Average tool score: 52.
Analyze a screenshot with a VLM. Requires a configured VLM provider.
Capture the screen, with measurement rulers drawn on the result.
Simulate left mouse click at current position.
Compare two images using a VLM. Requires a configured VLM provider.
Perform drag operation between coordinates.
Move mouse to specified screen coordinates.
Scroll vertically (positive=up, negative=down). Note: Each unit represents one notch on the scroll wheel (120 units = high-definition scroll). Typical values range from 2-3 to 5-10 for normal scrolling.
click_mouse has empty input schema (no parameters declared) but should document side effects and cursor state expectations.
No output schemas documented for any tool. LLMs cannot plan multi-step chains without knowing what fields are returned.
click_mouse description is only 61 characters and lacks context: 'Simulate left mouse click at current position.' Does not explain whether this requires mouse to be positioned first, what happens on click, or error conditions.
compare_images and analyze_screenshot require VLM provider configuration but lack error guidance. Descriptions do not explain failure modes (missing images, API key not set, timeout) or recovery paths.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 53 | <=2025-11-25 | v2 |
No permission gates or audit logging declared. Desktop input/output control (move_mouse, click_mouse, etc.) is inherently destructive and sensitive but lacks documentation of scope declarations or audit trails.
capture_screenshot returns filename and boolean include_mouse as input parameters but output schema is not documented. LLMs cannot verify success or retrieve the image path.
compare_images and analyze_screenshot accept string paths to images but do not document validation, file size limits, supported formats, or access control (path traversal risk).