CodeAct implementation with Docker and MCP - an agent framework that executes code and integrates with MCP servers
CodeAkt exposes only one tool (exec) with a highly problematic definition. The tool name lacks a clear action verb, the description is generic and incomplete, the schema is present but severely underspecified, and the tool implements an irreversible, dangerous operation (arbitrary code execution) without any safety guardrails, confirmation mechanisms, or error guidance. The risk annotation 'IRREVERSIBLE' is noted but not actionable to an LLM. No parameter descriptions exist. Output schema is entirely undocumented. This tool violates multiple critical patterns around naming, description clarity, parameter documentation, and error handling.
Execute Python code with access to tools and get execution results
Tool name 'exec' lacks action verb. Should be 'execute_python_code' or 'run_python'. Single-word names are ambiguous and do not convey intent to LLMs.
Tool description is only 95 characters and lacks critical context: What happens when called? When should it be used vs alternatives? What does it return? Does it have side effects? No LLM guidance on safe/unsafe usage.
Input parameters lack descriptions entirely. 'code', 'tool_names', 'tool_server_port', 'interpreter_id', 'session_id' are undefined. An LLM cannot determine what values to pass or why these parameters exist.
No output schema documented. Callers do not know what fields to expect from execution results, making it impossible to chain this tool with others or extract relevant data.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Tool executes arbitrary Python code with access to external tool servers, a fundamentally dangerous capability with no confirmation, dry-run, or resource limit. No error recovery guidance provided. An LLM making a mistake here can break systems or leak secrets.
Risk annotation 'IRREVERSIBLE' is noted in metadata but never surfaces in the description or error handling. LLMs cannot read metadata fields, the description must explicitly state this is dangerous and when confirmation is needed.
No validation or error guidance. If the provided 'code' is invalid Python, or 'tool_server_port' is unreachable, the LLM receives no actionable error message and cannot self-correct.
Parameter 'tool_server_port' exposes infrastructure details (a port number). This should be abstracted: either require a tool server URL, or resolve it internally. Forcing the agent to supply raw port numbers is fragile and unmaintainable.
Parameter 'interpreter_id' and 'session_id' lack description. Are they opaque tokens? Do agents need to track them across calls? Can they be omitted? Undefined.