Local LLM Agent Sandbox with code execution, desktop control, and TCP forwarding capabilities
EdgeBox MCP server exposes 16 tools with high-risk operations (code execution, file deletion, desktop control, VNC access). While tool names follow verb_noun conventions (execute_python, read_file, delete_directory), descriptions lack the specificity required for LLM-safe operation of destructive tools. Most critical: parameter descriptions are minimal or absent, input schemas lack validation constraints (enums, patterns, ranges), and error handling does not guide recovery. The server is suitable for sandboxed local operation but does not meet production agent safety standards. No output schemas are documented. Several tools (execute_python, execute_javascript, execute_bash, delete_file, delete_directory, open_url, click, type_text) perform irreversible or high-impact actions without confirmation mechanisms or permission gates.
Click the mouse at the current position or specified coordinates
Create a new directory
Delete a directory and its contents
Delete a file
Execute bash commands in isolated environment (stateless - suitable for system administration and file operations)
Execute JavaScript code in isolated environment (stateless - suitable for browser automation and Node.js tasks)
Execute Python code in isolated environment (stateless - suitable for calculations and analysis)
Get VNC URL for remote desktop access to the sandbox container
Destructive operations (delete_file, delete_directory) lack confirmation or dry-run support. No recovery guidance in error responses.
Code execution tools (execute_python, execute_javascript, execute_bash) lack output schema documentation and error recovery guidance. Descriptions do not specify timeout, resource limits, or sandbox constraints.
Desktop automation tools (click, type_text, press_key, move_mouse) lack descriptions explaining coordinates, coordinate system origin, or screen resolution assumptions. Parameter 'button' in click has no enum constraint (left|right|middle).
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 14 | - | v1 |
List files in a directory with size and modification time information
Move mouse to specific coordinates
Open a URL in the desktop browser
Press a keyboard key
Read the contents of a file
Take a screenshot of the current desktop and return as base64
Type text on the keyboard
Write contents to a file (creates or overwrites)
No input validation documented. Parameters accept free-form strings with no ranges, patterns, or constraints. E.g., path parameters lack validation against path traversal; file content size is unbounded.
No output schemas documented for any tool. LLMs cannot predict return structure, plan chaining, or extract needed fields (e.g., what does execute_python return, stdout only, or stdout+stderr+exit_code?)
Permission checks missing. Any agent can call delete_file, delete_directory, execute_bash without verification. No scope declarations (read:filesystem, write:filesystem, delete:filesystem, execute:code).
Error handling does not guide recovery. No classification of errors (retryable, user-fixable, fatal). No actionable messages for common failures (file not found, permission denied, command timeout, invalid syntax).