MCP server that runs commands in sandbox containers (Docker or Kubernetes) with support for artifact collection, streaming output via Server-Sent Events, and configurable execution policies.
Three tools defined with strong schemas and detailed descriptions. run_sandbox has excellent parameter documentation with type constraints (enum for as_user, minLength/maxLength for identifier), though output schema is implicit (SSE + polling pattern). delete_sandbox and restart_sandbox are simpler but lack output documentation. Tool names follow verb_noun convention (run_, delete_, restart_). Descriptions are LLM-optimized and include procedural context (SSE vs polling fallback, artifact handling, user permission levels). No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite destructive operations. Error handling is not explicitly documented in descriptions.
Stop and remove the sandbox container for an identifier.
Recreate an empty sandbox container for an identifier (delete + create).
Run commands in a sandbox keyed by identifier. Returns a run_id plus events_url (SSE) and status_url. To stream output, connect to events_url and read until event run_finished or run_failed. If SSE is unavailable, poll status_url until state is completed/failed. Set options.await_completion=true to include a final run summary inline (including captured stdout/stderr, subject to truncation). Command form: provide exactly one of shell (string) or argv (string[]). For compatibility with some clients, shell may be omitted/empty/false when using argv. Use options.as_user="root" for administrative operations (e.g. apt install); use "sandbox" for non-privileged commands. Artifacts are copied from /artifacts after completion. Networking and other sandbox policy are configurable server-side.
Missing tool annotations for destructive operations. delete_sandbox and restart_sandbox lack destructiveHint annotation to signal to LLMs that these operations are irreversible and should trigger confirmation patterns.
Incomplete output schema documentation. run_sandbox returns run_id, events_url, and status_url but these are described only in text, not in a formal output schema block. LLMs cannot reliably extract or compose with undocumented return types.
delete_sandbox and restart_sandbox descriptions are under 50 characters (bare statements with minimal context).
Error handling guidance missing. Descriptions do not indicate what errors are retryable (timeout), user-fixable (invalid identifier format), or fatal (container runtime failure). LLMs cannot plan recovery without error classification.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 0 | 2025-11-25+ | v1 |
Confirmation or dry-run pattern missing for destructive tools. run_sandbox with allow_failure and delete_sandbox are high-consequence operations but lack a confirmation step. Agents can accidentally destroy sandboxes without safeguards.
Timeout parameter is optional on individual commands but lacks clear guidance on what happens if timeout_ms elapses. Does it kill the process? Return partial output? Retry automatically?