A secure Python code execution sandbox for MCP with Docker-based isolation, supporting OMOP CDM schema management, Synthea data loading, and analytics
OMCP Python Sandbox has significant definition quality issues. While 14 tools are defined, many lack proper schema details, parameter descriptions are sparse or missing, and output schemas are not documented. The server exposes infrastructure concerns (Docker IDs, file paths) rather than user-facing abstractions. Several tools have duplicate or conflicting parameter definitions (e.g., execute_python_code has both 'python_code' and 'code' parameters). Tool naming is inconsistent (Get_information_Schema uses PascalCase; others use snake_case). Descriptions range from adequate (create_sandbox, execute_python_code) to minimal (ping, query_duckdb). Error handling guidance is absent from nearly all tools. The server mixes domain concerns (OMOP/Synthea-specific tools) with generic infrastructure (sandbox management, query execution), making discovery and composition unclear. No tool describes what downstream tools it chains to or what IDs are needed next.
Return available schemas and tables. Compatible with agents expecting Get_information_Schema().
Execute an arbitrary read-only SQL query against DuckDB (if DB_PATH is set) or PostgreSQL otherwise. Compatible with agents expecting Select_Query().
Perform analytics on OMOP data using pandas and LLM-friendly output. Args: sandbox_id: The sandbox to execute in. analysis_type: Type of analysis ('basic', 'demographics', 'conditions').
Create OMOP CDM schema in PostgreSQL database from within the sandbox. Args: sandbox_id: The sandbox to execute in. Returns: Dict with 'output' and 'exit_code' or 'error'.
Create a new Python sandbox environment. This tool creates a new Docker container that will serve as an isolated Python execution environment. The container is configured with: - No network access (security) - Memory and CPU limits - Auto-removal when stopped - Enhanced security options (read-only, dropped capabilities) - User isolation (sandboxuser) - Temporary filesystem mounts Args: timeout: Optional timeout for the sandbox in seconds (default: 300) Returns: Dict containing: - success: Boolean indicating if creation was successful - sandbox_id: Unique identifier for the created sandbox - created_at: ISO timestamp of creation - last_used: ISO timestamp of last usage - error: Error message if creation failed
Duplicate and conflicting parameters: execute_python_code defines both 'python_code' AND 'code' in input schema, creating ambiguity about which parameter to use. LLMs will be confused about which is canonical.
Naming inconsistency: Get_information_Schema and Select_Query use PascalCase, violating the snake_case convention of all other tools (create_sandbox, list_sandboxes, etc.). This inconsistency will confuse LLMs and agent tooling.
Output schemas not formally documented: 14 tools have descriptions that mention return structure in prose only (e.g., 'Dict containing success, output, error'). No formal JSON Schema definitions for outputs are visible in the tool registration. LLMs cannot plan downstream tool calls without knowing what fields they'll receive.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 11 | - | v1 |
Execute Python code in a secure sandbox environment. Args: sandbox_id: The unique identifier of the sandbox to execute code in code: The Python code to execute (must be non-empty string) timeout: Optional execution timeout in seconds (default: 30) Returns: Dict containing: - success: Boolean indicating if execution was successful - output: The stdout output from code execution - error: The stderr output or error message - exit_code: The exit code from the Python process
Install a Python package in a sandbox. Args: sandbox_id: The unique identifier of the sandbox to install the package in package: The package name and version timeout: Optional installation timeout in seconds (default: 60) Returns: Dict containing: - success: Boolean indicating if installation was successful - output: Installation output - error: Installation error or stderr output - exit_code: The exit code
List all active Python sandboxes. Args: include_inactive: Whether to include inactive sandboxes (default: False) Returns: Dict containing: - success: Boolean indicating if listing was successful - sandboxes: List of sandbox information dictionaries - count: Number of sandboxes in the list - error: Error message if listing failed
Perform LLM-friendly dataframe operations on OMOP data. Args: sandbox_id: The sandbox to execute in. operation: Natural language description of the operation. table_name: Target OMOP table.
Load Synthea CSV files into PostgreSQL OMOP database from within the sandbox. Args: sandbox_id: The sandbox to execute in. csv_directory: Directory containing Synthea CSV files.
ping
Run a SQL query against the DuckDB file and return the results. Args: sql: The SQL query to run. Returns: Dict with 'success', 'columns', 'result', and 'error' keys.
Directly query an OMOP table without sandbox overhead (Fast Path). This tool runs directly on the server, bypassing the Docker container creation process. It is significantly faster for read-only data retrieval. Args: table_name: Name of the OMOP table (e.g., 'person', 'visit_occurrence') limit: Maximum number of rows to return (default: 100, max: 1000) columns: List of columns to select (default: all) where: Optional raw SQL WHERE clause (disabled by default) filters: Optional column/value filters (equality only)
Remove a Python sandbox. Args: sandbox_id: The unique identifier of the sandbox to remove force: Whether to force removal of active sandboxes (default: False) Returns: Dict containing: - success: Boolean indicating if removal was successful - message: Success message or error description - error: Error message if removal failed
No error handling guidance: Tools like remove_sandbox (destructive), execute_python_code (arbitrary execution), and install_package (external network) lack error recovery instructions. Descriptions do not tell LLMs whether failures are retryable, require user intervention, or are fatal.
No security warnings or input validation guidance: execute_python_code accepts arbitrary Python code; llm_dataframe_operation accepts free-form natural language for SQL generation; query_duckdb and Select_Query accept raw SQL. No mention of injection risks, timeout protections, or validation strategy in descriptions.
Tool naming does not reflect user intent: Several tools expose infrastructure abstractions (sandbox_id, Docker IDs) rather than domain concepts. 'execute_python_code' is infrastructure; 'analyze_omop_data' is domain. LLMs must reason about sandboxes as an implementation detail, not as user-facing abstractions.
No tool chaining metadata: Tools do not indicate what downstream tools they chain to or what IDs are required. E.g., create_sandbox returns sandbox_id, but the description doesn't say 'use this with execute_python_code' or 'required for all sandbox operations'.
Minimal descriptions for several tools: ping (1 char), Get_information_Schema (16 chars), query_duckdb (2 lines), load_synthea_to_postgres (2 lines), create_omop_schema (3 lines). These are well below the 50-200 char baseline for LLM-optimized descriptions and lack context on WHEN to use the tool.
Free-form string parameters without constraints: 'operation' in llm_dataframe_operation, 'sql' in query_duckdb and Select_Query, and 'python_code' in execute_python_code accept arbitrary text with no enum, pattern, or length constraints. This invites injection attacks and hallucinated values from LLMs.
No idempotency or confirmation pattern for destructive operations: remove_sandbox is marked DESTRUCTIVE but has no dry-run, confirmation, or idempotency guarantees. Agents cannot safely retry if removal partially fails.
Duplicate tool functionality: query_duckdb, query_omop_table, and Select_Query all perform SQL queries. The LLM must reason about which to use based on implicit backend differences (DuckDB vs PostgreSQL vs OMOP) rather than explicit, differentiated tool names.
Missing parameter bounds and validation rules: timeout parameters (create_sandbox, execute_python_code, install_package) have no min/max values. limit in query_omop_table notes a max of 1000 in description but no formal schema constraint. chunk_size in load_synthea_to_postgres has no bounds. LLMs may pass absurd values.