Run arbitrary JavaScript inside disposable Docker containers and install npm dependencies on the fly.
This server provides 7 tools with generally adequate descriptions and schemas, but has several moderate issues that prevent a higher score. Tool naming is mostly clear and action-oriented (sandbox_initialize, sandbox_exec, run_js, sandbox_stop, run_js_ephemeral, get_dependency_types, search_npm_packages). Descriptions are present and reasonably detailed (average ~250 chars), which is good. However, several parameters lack descriptions, and the server has critical security and composition concerns. Most tools have proper input schemas with type definitions, though some parameter descriptions are thin. Error handling guidance is minimal, most tools do not explain recovery paths or categorize errors. The tools function as a cohesive set for Docker-based Node.js sandboxing, but descriptions could be more LLM-optimized and parameter validation more thorough.
Given an array of npm package names (and optional versions), fetch whether each package ships its own TypeScript definitions or has a corresponding @types/… package, and return the raw .d.ts text. Useful whenwhen you're about to run a Node.js script against an unfamiliar dependency and want to inspect what APIs and types it exposes.
Install npm dependencies and run JavaScript code inside a running sandbox container. After running, you must manually stop the sandbox to free resources. The code must be valid ESModules (import/export syntax). Best for complex workflows where you want to reuse the environment across multiple executions. When reading and writing from the Node.js processes, you always need to read from and write to the "./files" directory to ensure persistence on the mounted volume.
Run a JavaScript snippet in a temporary disposable container with optional npm dependencies, then automatically clean up. The code must be valid ESModules (import/export syntax). Ideal for simple one-shot executions without maintaining a sandbox or managing cleanup manually. When reading and writing from the Node.js processes, you always need to read from and write to the "./files" directory to ensure persistence on the mounted volume. This includes images (e.g., PNG, JPEG) and other files (e.g., text, JSON, binaries). Example: ```js import fs from "fs/promises"; await fs.writeFile("./files/hello.txt", "Hello world!"); console.log("Saved ./files/hello.txt"); ```
Execute one or more shell commands inside a running sandbox container. Requires a sandbox initialized beforehand.
Critical security gap: No input sanitization documented or visible for tools that execute arbitrary code (run_js, run_js_ephemeral, sandbox_exec). These tools accept code/command strings from the LLM without any validation, lint, or sandbox verification, risking code injection attacks.
Output schemas are not documented for any tool. LLMs cannot plan downstream composition without knowing what fields are returned. E.g., what does sandbox_initialize return? A container_id? Are there constraints or metadata? What does run_js_ephemeral return, stdout, exit code, file list?
Error handling is undocumented and provides no recovery guidance. No tool description explains what happens on failure, how to retry, or what the LLM should do next. E.g., if sandbox_exec times out, should the agent retry? Call sandbox_stop and reinitialize? Create a new container?
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 28 | - | v1 |
Start a new isolated Docker container running Node.js. Used to set up a sandbox session for multiple commands and scripts.
Terminate and remove a running sandbox container. Should be called after finishing work in a sandbox initialized with sandbox_initialize.
Search for npm packages by a search term and get their name, description, and a README snippet.
Destructive operations (sandbox_stop) lack confirmation or dry-run support. An agent mistake or retry loop could tear down active sandboxes without warning. No undo or warning mechanism exists.
Parameter descriptions are sparse or missing for several tools. 'commands' in sandbox_exec has a description but no format guidance (JSON array? Shell syntax?). 'dependencies' in run_js is just 'NPM dependencies to install', no format (package@version or just package name?).
Composition and state management are unclear. run_js requires manual cleanup via sandbox_stop. The description says 'you must manually stop the sandbox' but does not explain what happens if you don't, whether resources leak, or how to detect active sandboxes. This breaks idempotency and complicates agent planning.
Tool naming ambiguity: 'run_js' vs 'run_js_ephemeral' differ only in lifecycle. An LLM may struggle to choose correctly, no clear guidance on when to prefer one over the other (e.g., 'Use run_js for multi-step workflows; run_js_ephemeral for single commands'). Consider renaming to 'run_js_persistent' or providing explicit comparison.
No rate limiting or resource quotas documented. Tools can spawn unlimited Docker containers or execute unbounded Node.js scripts. An agent in a retry loop or prompt injection attack could exhaust system resources.