An MCP server that enables Claude to control a desktop environment via Docker containers, supporting screenshots, mouse/keyboard input, window management, file operations, and multi-container environments
This MCP server exposes 32 desktop automation and container management tools. While tool names follow a clear verb_noun convention (computer_*) and all tools have descriptions, there are significant gaps in schema completeness and parameter documentation. Many tools lack explicitly visible input schema definitions in the provided source. Descriptions range from adequate to generic. Output schemas are not documented. Error handling is minimal. The codebase appears incomplete in the provided snippet, critical schema and validation code may exist in unshown portions, but based on the evidence provided, quality is below production baseline.
Execute a shell command in the container. Can run any command accessible in the environment (bash, curl, git, python, etc.).
Create a new container environment with specified configuration.
List all available container environments with their configurations.
Start a stopped container environment.
Stop a running container environment.
Append text to a file in the workspace. Creates the file if it doesn't exist.
Delete a file from the workspace.
List files and directories in a workspace directory.
Output schemas completely undocumented. LLMs cannot infer what fields are returned from tools like computer_screenshot, computer_window_list, computer_env_list, or computer_file_list. This forces agents to guess the response structure and breaks chaining, a downstream tool that needs window_id cannot reliably extract it if the schema is unknown.
Many parameter descriptions lack constraint details. For example, computer_key_press accepts 'key' but does not explicitly document valid key names, modifiers, or format (is it 'Return', 'Enter', 'ctrl+a', or 'Ctrl+A'?). This forces LLMs to guess or hallucinate key names, causing failed keypresses.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 17 | - | v1 |
Create a directory in the workspace.
Read the contents of a file from the workspace. Returns up to 16,000 characters. For binary files, returns a message indicating the file is binary.
Write or overwrite a file in the workspace.
Press a key or key combination. Supports single keys, modifier+key, and key names.
Execute a named keyboard shortcut (e.g., 'copy', 'paste', 'undo', 'save'). Supported shortcuts: copy, cut, paste, undo, redo, select_all, delete_line, save, save_as, open, new_file, print, find, find_replace, find_next, new_tab, close_tab, reopen_tab, next_tab, prev_tab, refresh, hard_refresh, address_bar, back, forward, close_window, fullscreen, switch_window, terminal_copy, terminal_paste, zoom_in, zoom_out, zoom_reset.
Click the mouse button (left, right, or middle) at the current cursor position.
Press and hold a mouse button.
Drag the mouse from (x1,y1) to (x2,y2). Clicks down, moves, then releases.
Move the mouse cursor to the specified coordinates.
Scroll the mouse wheel up or down.
Release a held mouse button.
Take a screenshot of the desktop. Use this frequently to observe the current state. Finds windows by title substring or exact window ID. Returns the screenshot in base64 JPEG format.
Record an action within an active session with a description.
End an active session and return all recorded actions.
Start a session to record a series of actions with descriptions. Useful for capturing workflows.
Type text character-by-character into the focused window. Supports printable ASCII and common symbols.
Activate a window (bring to front and focus) by window ID or title substring.
Close a window by window ID or title substring.
Focus a window by window ID or title substring.
List all visible application windows with their positions, sizes, and names.
Maximize a window by window ID or title substring.
Minimize a window by window ID or title substring.
Move a window to a new position.
Resize a window to new dimensions.
computer_key_shortcut accepts a 'shortcut' parameter, but the valid shortcut names are hardcoded in a SHORTCUTS map (copy, cut, paste, undo, redo, etc.). This map is not exposed in the tool description as an enum, the description merely lists supported shortcuts as prose. LLMs cannot parse prose lists reliably; they need formal enums. Result: agents may pass unsupported shortcut names.
computer_mouse_click, computer_mouse_down, computer_mouse_up all accept a 'button' parameter described as 'Mouse button: left (1), right (3), or middle (2)'. This mixes enum values (left/right/middle) with numeric codes (1/3/2) in the description but does not use a formal enum schema. This ambiguity causes LLMs to pass invalid values like 'button: 1' or 'button: LEFT'.
Destructive tools (computer_file_delete, computer_window_close, computer_cmd_exec) lack dry-run, confirmation, or warning patterns. computer_file_delete description does not warn that it is irreversible or suggest a preview step. computer_cmd_exec accepts arbitrary shell commands with no validation or sandbox hints. Agents can accidentally delete files or execute dangerous commands.
Error handling is absent from tool descriptions. Tools offer no guidance on what to do if a window is not found, a file does not exist, a command times out, or a container fails to restart. Agents have no recovery path, they receive a generic error and must guess next steps.
Tool descriptions are generic and lack WHEN-to-use context. For example, computer_mouse_move says 'Move the mouse cursor to the specified coordinates' but does not explain when to use this vs computer_mouse_drag, or mention that coordinates must be display pixels, not API pixels (scaling is applied but not documented in the tool desc).
Session recording tools (computer_session_start, computer_session_action, computer_session_end) lack documented output structure. What fields does computer_session_end return? Just action strings, or also timestamps, results, or metadata? Without schema docs, agents cannot parse session data.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. The server declares risk levels (READ_ONLY, WRITE, DESTRUCTIVE, IRREVERSIBLE) but does not map these to MCP protocol tool annotations. This prevents clients from enforcing safety policies (e.g., blocking irreversible operations without approval).
Container parameter 'container' is optional on all tools and defaults to DEFAULT_CONTAINER, but this default behavior is not documented in any tool description. If an agent does not know about the default, it may fail to specify a container for a multi-container setup, causing the action to silently run on the wrong container.