MCP server that executes code securely in isolated Docker containers. Supports 12 languages: Python, JavaScript, TypeScript, Ruby, Bash, Zsh, Fish, Java, Clojure, Kotlin, Groovy, Scala. Provides code execution, syntax validation, and session reset capabilities.
Code Sandbox MCP provides 3 tools with complete input schemas and descriptions. All tools follow verb_noun naming conventions (execute_code, validate_code, reset_session). Descriptions are present and substantive (107 - 159 chars), though they could better address WHEN to use each tool and dependency hints. Input schemas are well-formed with proper types and enums. However, tool responses lack documented output schemas, critical for helping LLMs understand what data flows between tools. Error handling is basic: the server declares errorReporting=true but provides no examples of recovery-guiding error messages. Parameter descriptions are concise but could better specify constraints (e.g., max code size, timeout behavior). The tools follow good composition, each does one thing, but lack per-tool risk annotations (destructiveHint, idempotentHint) that would help agents reason about side effects.
Execute code securely in isolated Docker containers. Supports 12 languages: Python, JavaScript, TypeScript, Ruby, Bash, Zsh, Fish, Java, Clojure, Kotlin, Groovy, Scala. Network access enabled for package installation and API calls.
Reset the execution session, clearing any state from previous code executions.
Validate code syntax for supported programming languages without execution.
Output schemas not documented. LLMs cannot predict what execute_code returns (stdout, stderr, exit code, execution time?), limiting ability to chain tools or extract results.
Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint). reset_session modifies state but lacks destructiveHint; validate_code is read-only but lacks readOnlyHint. Agents cannot reason about side effects without explicit hints.
Parameter descriptions lack constraint details. 'code' parameter has no length limit, timeout, or resource warnings. 'filename' lacks format guidance (extension required?). LLMs will pass unbounded input.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
No dependency hints in descriptions. Users reading tool descriptions won't know that validate_code should be called before execute_code for safety. Missing WHEN guidance.
reset_session description is vague ('Reset the execution session'). What state is cleared? Can it be called mid-execution? Are partial results lost? Unclear when to use it.