A code execution sandbox MCP server that provides a Docker-based environment for safely executing code with file I/O capabilities
The server defines 7 tools with explicit schemas and descriptions visible in main.go. All tools have descriptions (range: 80-200 chars) and input schemas with type definitions. However, there are significant gaps in parameter quality, output schema documentation, and error handling patterns. Tool naming is generally clear (verb_noun pattern), but schemas lack depth, no enum constraints on docker image selection, no pagination support despite file transfer operations, and critically missing output schemas for all tools. The descriptions are adequate for discovery but lack LLM-optimization (e.g., no 'WHEN to use' guidance, no dependency hints between tools, no error recovery guidance). Error handling is minimal, no documentation of retryable vs fatal errors, no recovery suggestions. The server is STDIO-only, which is a hard transport limitation. Overall, this is a functional but mediocre MCP server typical of community tools.
Copy a single file to the sandboxed filesystem. Transfers a local file to the specified container.
Copy a single file from the sandboxed filesystem to the local filesystem. Transfers a file from the specified container to the local system.
Copy a directory to the sandboxed filesystem. Transfers a local directory and its contents to the specified container.
Execute commands in the sandboxed environment. Runs one or more shell commands in the specified container and returns the output.
Initialize a new compute environment for code execution. Creates a container based on the specified Docker image or defaults to a slim debian image with Python. Returns a container_id that can be used with other tools to interact with this environment.
Stop and remove a running container sandbox. Gracefully stops the specified container and removes it along with its volumes.
No output schemas documented for any tool. LLMs cannot infer what fields to expect in responses, making downstream tool chaining impossible and forcing wasteful discovery calls.
Missing enum constraint on 'image' parameter in sandbox_initialize. The description says '(e.g., python:3.12-slim-bookworm)' but accepts free-form strings, inviting hallucinated image names. Should provide a constrained list of supported base images.
Parameter descriptions lack format/constraint details. E.g., 'file_name' in write_file_sandbox has no guidance on allowed characters, length limits, or path traversal restrictions. LLMs will pass arbitrary strings; the tool must validate or describe constraints in the description.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Write a file to the sandboxed filesystem. Creates a file with the specified content in the container.
No error handling or recovery guidance in any tool description. Tools like sandbox_exec are marked WRITE but lack documentation of what happens on failure, whether operations are retryable, or what the LLM should do next. No dry-run or confirmation pattern for destructive operations (sandbox_stop).
copy_file, copy_project, copy_file_from_sandbox lack pagination or size limit constraints. No documentation of max file size, max directory depth, or streaming behavior for large transfers. An LLM could attempt to copy a 10GB dataset, hanging the agent.
Tool descriptions do not clarify dependencies or sequence. E.g., copy_project requires sandbox_initialize first, but this is not stated in copy_project's description. An LLM could invoke tools in the wrong order and waste context on errors.
sandbox_initialize has a default image ('python:3.12-slim-bookworm') but no way for an LLM to discover what other images are available. The description offers no guidance on how to choose images or what pre-installed tools each has.