MCP server that executes Python code in isolated Docker containers with resource limits, security constraints, and sandboxing
Single tool 'run_code' with clear verb-noun naming and decent schema structure. Tool description is adequate (89 chars) but lacks critical safety guidance on irreversibility. Input parameters have types and basic descriptions, but lack validation constraints (e.g., max_code_bytes is enforced in code but not documented in parameter description). Output schema is inferred from code (model_dump()) rather than explicitly documented. Error handling is present (ToolError for empty/oversized code) but recovery guidance is minimal. No tool annotations despite IRREVERSIBLE risk classification.
Execute Python code in an isolated Docker container with configurable resource limits
Tool description (89 chars) does not explicitly state the tool is IRREVERSIBLE or warn about the destructive consequences of running arbitrary code. LLMs need clear state-modification language: 'WARNING: Executing code may alter container state permanently. This operation cannot be undone.'
Parameter 'code' lacks documented constraints despite enforcement in implementation. Description says 'Must not be empty and must not exceed max_code_bytes limit' but does not specify the actual byte limit value. LLMs cannot reason about constraints they cannot see in the schema.
Output schema is not explicitly documented. Code calls result.model_dump() but the JobResult Pydantic model definition is not visible in the provided source. LLMs cannot plan chained calls without knowing what fields are returned (status, stdout, stderr, exit_code, duration_ms, truncated, error_type, error_message).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No tool annotations despite high-risk operation. Tool should declare 'destructiveHint: true' to signal to clients that this operation modifies state. The IRREVERSIBLE risk classification in the rubric indicates this tool requires explicit safety signaling.
Error messages in code (e.g., 'INVALID_INPUT: code is too large') provide recovery guidance but the pattern is inconsistent. Some errors (EXECUTOR_BUSY, EXECUTOR_UNAVAILABLE) include retry guidance; others do not. Standardize error responses per pattern:recovery-guide.