Run arbitrary JavaScript inside disposable Docker containers and install npm dependencies on the fly.
This server presents significant definition quality gaps across most tools. While tool names generally follow verb_noun conventions (sandbox_initialize, sandbox_exec, run_js), descriptions are often lengthy and lack the focused, LLM-optimized structure needed for reliable tool selection. Parameter schemas are partially visible but incomplete, most tools lack explicit parameter descriptions in the schema definitions. Output schemas are not documented anywhere in the source. Error handling is absent from tool descriptions, leaving agents without guidance on recovery paths. The server mixes high-risk operations (Docker container execution, arbitrary JavaScript) with missing security context and permission declarations. Only 2 of 7 tools have descriptions that meet the 50-200 character sweet spot; most are either too long or lack critical context. No tool documents what fields are returned or how outputs chain to other tools.
Generate text using Google Gemini. Provide a prompt and optional model name.
Given an array of npm package names (and optional versions), fetch whether each package ships its own TypeScript definitions or has a corresponding @types/… package, and return the raw .d.ts text. Useful whenwhen you're about to run a Node.js script against an unfamiliar dependency and want to inspect what APIs and types it exposes.
Install npm dependencies and run JavaScript code inside a running sandbox container. After running, you must manually stop the sandbox to free resources. The code must be valid ESModules (import/export syntax). Best for complex workflows where you want to reuse the environment across multiple executions. When reading and writing from the Node.js processes, you always need to read from and write to the "./files" directory to ensure persistence on the mounted volume.
Run a JavaScript snippet in a temporary disposable container with optional npm dependencies, then automatically clean up. The code must be valid ESModules (import/export syntax). Ideal for simple one-shot executions without maintaining a sandbox or managing cleanup manually. When reading and writing from the Node.js processes, you always need to read from and write to the "./files" directory to ensure persistence on the mounted volume. This includes images (e.g., PNG, JPEG) and other files (e.g., text, JSON, binaries). Example: ```js import fs from "fs/promises"; await fs.writeFile("./files/hello.txt", "Hello world!"); console.log("Saved ./files/hello.txt"); ```
ALL 7 tools have ZERO visible input schemas in server.ts. Schemas are imported as argSchema from separate tool files (initialize.ts, exec.ts, etc.), but the actual Zod schema definitions are not present in the provided source code. Unable to verify parameter types, constraints, enums, or descriptions. This violates the critical rule: 'If a tool has NO visible schema, schema score MUST be 0.'
Output schemas are entirely undocumented. No tool description states what fields are returned, their types, or structure. Agents cannot plan downstream calls or extract necessary IDs (e.g., sandbox_id, container_id) without guessing. This blocks tool composition and forces agents to explore trial-and-error.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 35 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Execute one or more shell commands inside a running sandbox container. Requires a sandbox initialized beforehand.
Start a new isolated Docker container running Node.js. Used to set up a sandbox session for multiple commands and scripts.
Terminate and remove a running sandbox container. Should be called after finishing work in a sandbox initialized with sandbox_initialize.
Critical security context missing: (1) No permission declarations (e.g., 'requires Docker write access'). (2) No error handling guidance, agents don't know what to do if a container fails to initialize. (3) No mention of rate limits, timeouts, or resource constraints. (4) Tools accept arbitrary shell commands and JavaScript code with no input validation described. Agents need explicit warnings about injection risks.
Tool descriptions are poorly optimized for LLM selection. sandbox_exec: 'Execute one or more shell commands...' (88 chars, minimal context). run_js: lengthy multi-line description (400+ chars) that buries key details about ESModules syntax, file persistence, and manual cleanup requirements in prose. Descriptions should state WHAT (1 line), WHEN (preconditions), and WHY (when to choose over alternatives).
Parameter descriptions are absent or vague in user-facing text. For run_js and run_js_ephemeral, the description mentions reading/writing to './files' directory, but there is no documentation of what parameters these tools accept (code, dependencies, etc.), their types, or constraints. Agents cannot determine valid input without reading implementation.
No error handling guidance. Descriptions do not explain failure modes or recovery strategies. If sandbox_initialize fails (e.g., Docker daemon unreachable, resource exhausted), agents have no recovery path. Critical pattern: errors must answer 'What to do next?', either retry, ask user, or give up.
Tool composition is fragile due to missing chaining IDs in output documentation. sandbox_initialize presumably returns a container_id or sandbox_id, but this is not stated in the description. Agents cannot confidently pass that ID to sandbox_exec, sandbox_stop, or run_js without guessing field names (container_id vs sandbox_id vs id).
ai_generate is a generic utility tool that belongs in a different MCP server (if at all). Mixing 'generate text via Gemini' with Docker sandbox tools violates single responsibility. This tool should be removed or moved to a dedicated AI provider server.
Tool naming uses underscores (sandbox_initialize, run_js_ephemeral) which is correct, but sandbox_initialize and run_js_ephemeral lack clear distinction in purpose. The description for run_js_ephemeral mentions it 'automatically clean[s] up' but the 'ephemeral' suffix is technical jargon, agents need explicit guidance: 'Use this for one-shot runs that don't need to persist state. Use run_js for multi-step workflows that reuse the environment.'