Execute Python code in Docker containers with optional package requirements
The server implements a single tool 'run_python' for executing Python code in Docker containers. While the tool has a description and a properly-structured input schema, the definition quality suffers from critical gaps: (1) the parameter 'code' lacks a description in the schema (only type:string is present), (2) the tool description is generic and does not explain when to use it or what it returns, (3) there is no documented output schema, the response structure is ad-hoc text concatenation with no guaranteed fields, (4) no error recovery guidance is provided to the LLM, and (5) the tool combines multiple concerns (building Docker images, running code, cleaning up) without clear separation. The naming 'run_python' is reasonable (starts with verb), but the implementation itself is risky, code execution is inherently irreversible and the lack of output structure, timeout documentation, or input validation guidance creates high operational risk. The schema visible in code shows 'code' parameter but the actual schema returned by list_tools() does not include descriptions for the 'code' parameter, a critical omission per the rubric.
Execute a Python command (mock implementation).
Parameter 'code' lacks description in inputSchema. The schema defines type but not what the parameter does, constraints, or format expectations.
Tool description is too generic (38 characters: 'Execute a Python command (mock implementation).'). Rubric baseline for tool descriptions is 194 chars (p10=34, p90=392). This description does not explain WHEN to use the tool, WHAT happens to the executed code, what the return value structure is, or whether execution is safe to retry. Does not meet the 'WHAT, WHEN, returns' requirement.
Output schema is not documented. The tool returns TextContent with a free-form string concatenation (Docker build info, code echo, output). There is no declared return structure (no documented fields, types, or format). LLMs cannot reason about downstream tool chains when output is unstructured. The response mixes status messages ('Docker build and run complete'), input echoes ('Executed code:...'), and actual output ('Output:...') in a single text blob.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 19 | - | v1 |
No error recovery guidance. When the tool fails (Docker build error, Python error, missing Docker), it returns error text but does not tell the LLM what to do next. 'Docker build failed: ensure Docker daemon is running' or 'Python code exited with error, check stderr.' A raw error message provides no actionable guidance.
Irreversible operation (code execution) lacks confirmation or dry-run pattern. The tool definition does not mention timeout, resource limits, or whether execution has side effects. Code execution is inherently dangerous (can delete files, consume resources, hang indefinitely). No mention of timeout constraints in the tool description, yet src/types.py defines an optional 'timeout' parameter that is not in the actual inputSchema returned by list_tools().
Parameter validation lacks clear constraints. The 'code' parameter accepts any string, but the implementation validates it ('non-empty string'). The schema does not document minLength or other constraints. The implementation also references 'requirements' parameter (a list) in code, but the schema returned by list_tools() does not include it, only 'code' is declared. This mismatch between schema and implementation is dangerous.
Schema incompleteness: src/types.py defines RunPythonRequestParams with 'code' (str) and 'timeout' (int|None=30) parameters, but the list_tools() implementation only declares 'code' in inputSchema. The 'requirements' parameter is used in call_tool() but not declared anywhere. This mismatch means the LLM will not know these parameters exist and may not be able to set timeouts or install requirements.
No input validation feedback. When code is empty or requirements is invalid, the tool returns an error string. However, errors like 'requirements must be a list of strings' are returned as-is to the LLM with no guidance on retry or correction.
Tool combines multiple concerns: Docker image building, code execution, and cleanup. Per the composition pattern, tools should do exactly one thing. This tool orchestrates a multi-step Docker workflow, making it harder to test, debug, and reason about failure modes. Consider splitting into separate tools or at least documenting internal steps.