Execute Python code in an isolated environment with Playwright and headless browser support for web scraping
Single tool with adequate naming and description, but moderate schema quality issues and missing error guidance patterns. The tool name 'execute-python' clearly conveys action. Description is reasonably detailed (222 chars, within 10-1024 baseline) and explains WHAT (execute Python), WHEN (real-time info, no internal source), and PREREQUISITES (Playwright available). However, the schema lacks formal constraints and error responses lack recovery guidance. Output schema is not documented. No input validation hints or field documentation in code. Error messages are generic ('Failed to create temp dir', 'Execution failed') rather than actionable.
Execute Python code in an isolated environment. Playwright and headless browser are available for web scraping. Use this tool when you need real-time information, don't have the information internally and no other tools can provide this information. Only output printed to stdout or stderr is returned so ALWAYS use print statements! Please note all code is run in an ephemeral container so modules and code do NOT persist!
Missing output schema documentation. Tool returns text via mcp.NewToolResultText() but no schema is documented for what structure or fields downstream tools should expect.
Error responses lack actionable recovery guidance. Errors like 'Failed to create temp dir: [error]' and 'Execution failed: [error]' do not tell the agent what to do next or how to self-correct.
Parameters lack formal constraints. 'code' is a free-form string with no maximum length, timeout, or security bounds specified. 'modules' accepts comma-separated strings but does not validate module names or prevent injection attacks.
No input validation or sanitization visible in tool implementation. Command construction uses string concatenation (strings.Join(shArgs, ' ')) without escaping, risking command injection if 'code' or 'modules' contain shell metacharacters.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Destructive/high-risk tool lacks confirmation or dry-run capability. Executing arbitrary Python code is inherently dangerous; tool should support a confirmation step or preview mode before execution.
No timeout enforcement visible. External docker execution could hang indefinitely, blocking the entire agent. No explicit timeout is set on cmd.Output().
Parameter descriptions lack format constraints. 'code' description does not specify max length, timeout, or security restrictions. 'modules' does not document the expected format (comma-separated, no spaces, module names only, etc.).