MCP server for code sandbox execution using E2B, supporting Python code execution, file uploads, and execution history tracking in Jupyter notebook format
The server has 5 tools with significant quality gaps. Tool names follow verb_noun convention reasonably well (initialize_sandbox, execute_code, upload_file, get_execution_history, close_sandbox). However, descriptions have major inconsistencies: some are in Chinese, others in English, creating cognitive friction. Parameter descriptions are present but often incomplete or poorly localized. Critical issues: (1) Global state management (Active_Sandboxes, Execution_DB dicts) violates stateless request handling; (2) execute_code returns a TypedDict (CodeOutput) without formal schema documentation; (3) upload_file parameter file_list has type 'array' but no item type specification; (4) No input validation, error categorization, or recovery guidance; (5) Resource endpoint get_execution_history is defined as a resource, not a tool, yet is listed in tool inventory, creating classification confusion. The server is functional but falls short of production-grade tooling patterns.
所有代码运行完毕之后,关闭已有的沙箱环境
Use this to execute python code and do data analysis or calculation
获取指定session的所有代码和代码执行历史记录,以Jupyter Meta Data形式返回
创建一个新的沙箱环境用于代码执行
上传本地文件到当前正在执行的沙箱中
Global state management violates stateless request handling, Active_Sandboxes and Execution_DB are module-level dicts that persist across requests, creating session coupling and potential race conditions in concurrent environments.
Output schemas not formally documented in tool definitions. execute_code returns TypedDict (CodeOutput) visible only in return annotation, not in the tool schema. upload_file returns string with no schema. get_execution_history returns JSON but no structured schema declaration.
Parameter constraints missing or incomplete. upload_file.file_list is typed as 'array' with no item type (should be array of strings). initialize_sandbox.timeout has no min/max bounds. No validation of session_id format or existence until runtime.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Descriptions are mixed Chinese/English, incomplete, and lack recovery guidance. No tool description explains when to call it vs similar tools, what permissions are needed, or how to handle failures. Error messages are plain strings, not categorized (retryable, user-fixable, fatal).
Critical security issues not documented: (1) execute_code accepts arbitrary Python with no input sanitization, prompt injection risk. (2) upload_file accepts 'absolute paths' with no path traversal protection. (3) No authentication/authorization checks. (4) E2B API key injected from environment but no note on secret scoping.
Resource vs tool classification confusion: get_execution_history is registered as a @mcp.resource but included in tool inventory. Resources and tools are distinct MCP primitives, mixing them creates LLM selection ambiguity.
No error categorization or recovery guidance. Exceptions caught and logged but error responses don't tell LLM what to do next (retry? ask user? fatal?). No distinction between transient failures (API timeout) and permanent ones (session not found).
Parameter descriptions are sparse or ambiguous. session_id appears in 4 tools but no description explains its format, lifetime, or how agents should obtain it. file_list description says 'each is absolute path' but doesn't specify if relative paths are rejected or what happens to symlinks.