A secure Python code execution service that provides a self-hosted alternative to OpenAI's Code Interpreter
CodeBox-AI defines a single tool, execute_code, with a reasonable structure but significant quality gaps. The tool name is action-verb-based (positive), the description is present but generic (194 chars is within baseline, but lacks actionable guidance on when/why to use), and the input schema is visible with type definitions. However, the schema lacks descriptions for individual parameters, error handling guidance is minimal, and output schema is only partially documented. The tool's destructive nature (code execution risk) is acknowledged but not reflected in security patterns like confirmation steps or dry-run support. Resources are declared but the session resource endpoint snippet is incomplete in the provided source.
Execute Python code and return a list of TextContent or ImageContent objects
Parameter descriptions missing. 'code' and 'dependencies' parameters lack descriptions in the schema; LLMs cannot infer constraints or format expectations.
Output schema not fully documented. Tool returns List[TextContent | ImageContent] but downstream agents don't know what fields ImageContent carries or how to chain results.
No error handling guidance for LLM recovery. Errors are returned as plain TextContent with no classification (retryable vs. fatal) or recovery steps.
Destructive tool lacks confirmation/dry-run pattern. execute_code runs arbitrary Python in a container with potential side effects (file writes, network calls, dependency installs) but offers no dry-run or confirmation step.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Tool description lacks actionable guidance. States WHAT (execute Python) but not WHEN (e.g., 'use for data analysis, plotting, test execution') or prerequisites (e.g., 'package installation may take 10 - 30s').
No input validation constraints exposed. Parameters accept free-form strings with no length limits, character restrictions, or format patterns documented for the LLM.
Session management incomplete and exposed to LLM. Code creates/deletes sessions internally but offers no session resource fully; get_session_info endpoint is truncated in source. Session lifecycle should be opaque or fully exposed with clear semantics.
No security audit trail in tool. Code execution is a sensitive operation; tool logs errors but does not provide traceable caller info, execution time, resource usage, or outcome for compliance.