A sandbox environment for executing code in multiple programming languages using Docker containers. Provides secure code execution with resource limits.
Single tool 'execute_code_in_sandbox' has decent naming and reasonable descriptions, but exhibits critical security and error handling gaps. The tool name follows verb-noun convention (execute_code) and the description is bilingual and present (~90 chars). Schema is properly defined with typed parameters (language, code, version) and required/optional markers. However, the implementation lacks error recovery guidance, input validation details, output schema documentation, and has a severe security concern: the tool accepts arbitrary code and executes it in Docker containers without documented sandboxing guarantees or resource limits. The source code shows a panic() on factory creation failure (line in sandboxHandler), which violates error handling patterns. No tool annotations (readOnlyHint/destructiveHint) are used despite this being a WRITE-risk tool. Description does not warn about execution risks or timeout behavior.
在沙盒环境执行代码 | Execute the code in a sandbox environment
Missing output schema documentation. Tool implementation calls sb.Execute() and returns stdout/stderr, but the tool definition provides no schema for what fields the response contains, their types, or how they should be interpreted.
Catastrophic error handling: sandboxHandler calls panic(err) on factory creation failure, which crashes the server instead of returning a proper error response. This violates the pattern for recovery guidance.
Tool description lacks safety and timeout warnings. Description states 'Execute the code in a sandbox environment' but does not document: execution timeout, resource limits, supported languages, or failure modes. LLM cannot infer sandbox guarantees or what to do if code hangs.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 41 | - | v1 |
No tool annotations despite being a destructive/stateful operation. Tool name and description do not carry destructiveHint, readOnlyHint, or idempotentHint. LLM cannot reason about whether this tool can be safely retried.
Parameter 'language' description lacks constraint documentation. Description says 'Programming language' but does not enumerate supported languages or validation rules. Code calls configManager.GetLanguageConfig(language) but LLM has no way to know valid options.
Error responses in code use generic fmt.Errorf() with stack traces ('failed to execute in sandbox: %w'). These do not provide recovery guidance (e.g., 'Timeout exceeded. Try reducing code complexity.' or 'Unsupported language. Try: python, go, php'). Pattern recovery-guide requires actionable error messages.
Parameter 'code' description does not document format constraints, max length, or character restrictions. Description is just 'The code to be executed'. LLM has no guidance on payload size limits or whether binary code is allowed.
Output schema not documented in source or tool registration. Tool returns mcp.CallToolResult containing stdout/stderr, but no schema is visible in the tool definition. LLM cannot know what fields to expect in the response.