An MCP server that provides tools to read, edit and run notebooks on a Jupyter server or a code sandbox.
The Jupyter MCP server provides 4 sandbox management tools with basic schemas and descriptions. However, the definitions have significant gaps: descriptions are minimal (mostly 6-12 words), many parameters lack descriptions or have vague ones, output schemas are not documented, and error handling guidance is absent. Tool naming follows verb-noun conventions (launch_, list_, terminate_, use_) but the descriptions are too terse to guide LLM selection properly. Parameters like 'environment', 'server_url', 'kernel_id', 'proxy_token' lack context about their purpose or format. The server exposes authentication tokens as parameters ('token', 'proxy_token'), violating the secret-injection pattern. No evidence of pagination, output field documentation, or recovery guidance in error cases.
Launch a code sandbox and register it for later use.
List all launched sandboxes.
Terminate one launched sandbox.
Select or clear the active sandbox used by execute_code.
Token and authentication parameters exposed as tool inputs ('token', 'proxy_token'). This violates secret-injection pattern: credentials must be injected server-side, not passed as parameters. Agent logs will contain secrets.
Tool descriptions are extremely terse (6-12 words). Provide complete context: WHAT the tool does, WHEN to use it vs. similar tools, WHAT it returns. Current descriptions ('Launch a code sandbox...', 'List all launched sandboxes') lack actionable detail for LLM selection.
Many parameters lack descriptions or have single-word descriptions. Examples: 'environment' (no context on valid values or format), 'kernel_id' (unclear when/how to obtain), 'server_url' (no format hint), 'channels_url' (no purpose stated).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 60 | 1.10.1+ | v1 |
No output schemas documented. LLMs need to know: what fields does launch_sandbox return? Is there a sandbox_id? What about list_sandboxes, does it return a count, pagination tokens, or just names? Without output documentation, agents cannot plan downstream calls.
No error handling guidance. What happens if launch_sandbox fails? Can the agent retry? Should it try a different variant? What if the sandbox already exists? Errors must guide recovery: 'Sandbox already exists, use use_sandbox() to select it, or terminate_sandbox() first.'
Parameter 'environment' and 'environment_version' lack enum constraints or format guidance. Agents will hallucinate invalid values. Provide explicit valid options: environment ∈ ['python', 'nodejs', 'rust', ...] or reference how to discover available values.
Parameter 'variant' has a default ('eval') but other parameters like 'timeout' (default 60) lack guidance on units, bounds, or rationale. Document: timeout in seconds, range 1-3600, default 60 suitable for interactive code.