MCP server providing Computer Use tools (bash, file operations, web automation, sub-agent delegation) via HTTP/Streamable HTTP transport with Docker container isolation per chat session
Open Computer Use has 5 tools with mixed quality. Naming follows verb_noun conventions (bash, view, create_file, str_replace, sub_agent). All tools have descriptions that explain purpose and risk level. However, parameter schemas are incomplete: bash and view lack documented parameter types/formats beyond names. create_file and str_replace have minimal schema detail. sub_agent has the most complete schema but still lacks constraints on model/cli enum values. Descriptions are generally adequate (100-200 chars) and mention irreversibility/side effects, which is good for risk communication. Output schemas are not documented, LLMs cannot infer what structure to expect from these tools. Error handling guidance is absent; tools do not indicate retry eligibility or recovery paths. The bash tool description mentions 'Streamable HTTP' for long-running commands, which is excellent for protocol support, but implementation details are not visible in source. Overall, this is a functional tool set with decent naming and risk awareness, but falls short of production-grade due to missing schema formalism, undocumented outputs, and limited error guidance.
Execute bash commands in isolated Docker container. Returns command output (stdout/stderr), exit code, and execution time. Long-running commands stream progress via Streamable HTTP.
Create new files with specified content. Supports text, JSON, binary, and base64-encoded data. Fails if file already exists.
Edit files via text replacement. Finds old_str within the file and replaces with new_str. Supports multi-line edits and regex patterns.
Delegate coding tasks to an autonomous sub-agent (Claude Code, Codex, or OpenCode). Supports multi-turn agentic loops with configurable model, max turns, and timeout. Returns structured result with code output, artifacts, and execution logs.
View files and directories. Returns file contents (text/JSON/binary), directory listings, or error if path not found. Respects container sandbox boundaries.
No documented output schemas for any tool. LLMs cannot infer what fields to expect, forcing them to guess at response structure and breaking tool chaining.
Parameter schemas lack type definitions and constraints. 'command', 'path', 'file_text', 'binary_data', 'old_str', 'new_str' are all strings but have no length limits, format hints, or validation rules. This invites invalid inputs from LLMs.
sub_agent's 'cli' parameter should be an enum ('claude'|'codex'|'opencode') but is documented as a free-form string. LLMs may pass invalid CLI names.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 62 | <=2025-11-25 | v2 |
No error handling guidance. Tools do not document recoverable vs fatal errors, do not suggest recovery actions, and do not guide LLMs on retry eligibility. E.g., bash should document exit code semantics, timeout behavior, and which errors are retryable.
create_file and str_replace lack idempotency guidance. create_file 'fails if exists' is good, but str_replace does not document whether repeated calls are safe or if 'old_str not found' is an error.
sub_agent delegates to external CLIs but lacks timeout and max_turns bounds documentation. An LLM could pass max_turns=10000, causing runaway sub-agent loops.
view and bash operate on container filesystem but lack documentation of sandbox boundaries, accessible paths, or what 'respects container sandbox' means. Can LLMs access /etc/passwd? /proc? /root?