Open source AI computer — MCP server with Ubuntu sandbox, browser, and desktop automation. Give any AI a real computer.
Strong tool coverage (33 tools) with consistent naming patterns and most parameters well-described. All tools have descriptions and input schemas. However, output schemas are not documented in the tool definitions, error handling guidance is minimal, and composition dependencies between tools could be better specified. The server demonstrates good fundamentals but lacks the production-grade polish of A-tier tools. Tool names follow verb_noun pattern consistently (vm_create, fs_read, browser_click). Most parameters include descriptions, but several lack type constraints (e.g., 'label' in vm_create is free-form string; snapshot operations lack validation feedback). Error messages are likely raw backend responses rather than LLM-optimized recovery guidance.
Navigate back in browser history.
Click an element by its ref number (from browser_snapshot). Use this to interact with buttons, links, and other clickable elements.
Execute arbitrary JavaScript in the browser context and return the result.
Navigate forward in browser history.
Press keyboard keys (e.g. Enter, Escape, Tab, ArrowDown). Supports single keys and key sequences.
Navigate to a URL in the browser.
Take a full page screenshot (scrolls to capture entire page).
Scroll the page (or an element) by a delta, or to a specific element ref.
Output schemas not documented. Tools return results but the response structure is not defined in tool registration. LLMs cannot infer what fields to expect or plan downstream tool calls.
Error handling and recovery guidance missing. No documented error cases or actionable recovery hints in tool descriptions. Errors will likely be raw backend responses that don't guide the LLM on corrective actions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 71 | 2026-07-28+ | v2 |
Set the browser viewport size.
Take a screenshot of the current browser tab and return an annotated image with numbered overlays for interactive elements. Use for vision-based interaction: get the ref number of the button/link you want to click, then pass it to browser_click.
Type text into a focused input field (or via ref).
Copy text to the sandbox's clipboard (xclip).
Get text from the sandbox's clipboard.
Search codebase content using ripgrep. Much faster and smarter than fs_search — supports regex, file type filters, context lines, and higher result limits.
Move mouse to (x, y) and click. Use for desktop UI elements outside the browser.
Press desktop keyboard keys (xdotool key).
Take a screenshot of the desktop (VNC framebuffer) without browser markup.
Type text on the desktop (uses xdotool type).
Run a shell command inside the sandbox. Returns combined stdout+stderr. Use for any CLI work: git, npm, pip, apt, curl, etc.
Replace a string in a file. Fails if old_string is not found. Set replace_all=true to replace every occurrence.
List files in a directory (ls -la, or recursive find with depth 3).
Read a file inside the sandbox.
grep -rn for a pattern in a directory.
Write content to a file inside the sandbox (creates parent dirs as needed).
Delete a saved snapshot by label. This frees disk space and means the next vm_create with this label will start fresh.
List all saved snapshots. Each snapshot can be resumed by calling vm_create with the same label and use_snapshot=true.
Create a new sandbox (isolated Ubuntu mini-computer with shell, Chrome + CDP, VNC, xfce4). IMPORTANT: Before creating, check if the user has saved snapshots by looking at the `available_snapshots` field in the response. If snapshots exist, ask the user whether they want to resume from a snapshot or start fresh. Set `use_snapshot: true` to resume. Backend is auto-detected: Firecracker microVM if KVM available, Docker otherwise. Returns the sandbox id used by all other tools, plus a noVNC URL the human can open to watch what the AI is doing. IMPORTANT: Always present vnc_url exactly as returned — never modify the URL path, filename, or query parameters.
Destroy a sandbox and free its ports. By default the VM filesystem is saved as a snapshot so recreating with the same label resumes where you left off. Set save_snapshot to false to destroy without saving (e.g. when the VM contains sensitive data the user does not want persisted). Use vm_reset to also delete any previously saved snapshot.
List active sandboxes created in this MCP session.
Rename a running VM by changing its label. The label is used for snapshot naming — renaming does NOT migrate existing snapshots (they stay under the old label).
Reset a VM to a clean state. Destroys the current VM AND deletes its saved snapshot, so the next vm_create with the same label starts fresh from the base image.
Restart a sandbox. Stops and starts the container — all files, databases, and installed packages are preserved. Only processes are restarted.
Get resource usage (CPU, memory, disk, uptime, top processes) and port mappings for a sandbox. Use to diagnose OOM, high CPU, or check available disk space.
Parameter constraints are underspecified. 'label' in vm_create accepts free-form strings with no validation hints. 'command' in exec and 'pattern' in searches lack regex or format guidance. LLMs cannot validate input and will pass invalid values.
Browser automation tools lack return value documentation. browser_click, browser_type, browser_key, browser_scroll do not document whether they return success/failure status, error details, or the new DOM state. Agents cannot validate actions completed.
Tool composition dependencies not documented. vm_create returns a 'sandbox id' used by all other tools, but this critical dependency is not explicitly stated in other tool descriptions. fs_edit's old_string matching and regex behavior are ambiguous.
Destructive operations lack confirmation mechanism. vm_reset, vm_destroy, and snapshot_delete are irreversible but provide no dry-run, confirmation_required, or idempotent_hint. Agents risk deleting user data without review.
Desktop/VNC coordinate-based tools (desktop_click, desktop_key) lack guidance on screenshot interpretation. No hints on how to obtain x,y coordinates from a desktop_screenshot or how to locate UI elements by visual inspection.