The server defines 6 tools with clear names, structured descriptions, and detailed parameter schemas. All tools follow verb-first naming (execute, read_output, drain_output, interrupt, status, restart). Descriptions are comprehensive (250-800 chars), explaining WHEN to use each tool and what to expect. Input parameters are well-typed with Pydantic annotations and clear descriptions. However, output schemas are not formally documented in the tool definitions, the descriptions explain output structure in prose, but JSON Schema output type definitions are absent. Error handling guidance is embedded in descriptions rather than formalized. Security is handled server-side (kernel configuration), not exposed in tool parameters.
Consume all pending execution output and outcomes in a single request. Returns all unread results within an 8 MiB JSON budget, removing each consumed outcome. Repeats until the response says "No pending output or outcomes". Results not returned in the budget remain unread; do not rely on reception for completed work. Each result includes a [metadata] text block containing JSON execution_id, status (running/succeeded/failed/cancelled), truncated (output dropped), and error (null or {type, message}). Metadata is followed by text/image content blocks. Use to clear the queue before submitting more code, or to retrieve all outcomes at the end of a session. No individual execution_id is needed.
Run code in the persistent Jupyter kernel for calculations, analysis, or images. Use source and libraries supported by the configured kernel. Variables, imports, and functions survive calls. Jupyter display data containing PNG or JPEG images is returned as image content. Returns content only: a [metadata] text block containing JSON execution_id, status (running/succeeded/failed/cancelled), truncated (output dropped since the previous read), and error (null or {type, message}), then text/images. Returned output is consumed. If running, call read_output with that ID and a positive wait_seconds; do not resubmit code. A final outcome removes the ID, so no follow-up read is needed. Code errors may leave partial state. If busy, read or interrupt the active execution before submitting more code. Interactive stdin is disabled. wait_seconds defaults to 10 and never stops code.
Request that the kernel stop executing the active code. Sends an interrupt signal via the kernelspec's interrupt mode (e.g., CTRL+C on Unix). Returns the execution_id that was targeted, if any. Interrupt is a request, not a confirmed stop; only a later read confirms the outcome. The kernel may ignore the request or complete before it arrives. Only one execution at a time is active; confirm the outcome by reading or waiting for drain_output. Returns {execution_id, interrupt_sent}: execution_id is the ID targeted or null if no code was active; interrupt_sent is true if the signal was sent.
Output schemas not formally documented. Tool descriptions explain output structure in prose (execution_id, status, truncated, error, content blocks), but JSON Schema definitions for return types are absent from tool registration. LLMs cannot programmatically parse expected fields.
No formal error classification in tool schemas. Error guidance ('Invalid requests, busy kernels, and consumed/expired IDs are MCP tool errors') is embedded in INSTRUCTIONS prose, not structured error codes or recovery hints in ToolError responses. Agents cannot programmatically distinguish retryable vs fatal errors.
drain_output has empty input schema (no parameters), but description mentions an '8 MiB JSON budget' and repeat-until semantics. No formal pagination or response-size documentation in the schema itself. LLMs cannot know whether to inspect response metadata for continuation signals.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 74 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Retrieve and consume one pending execution's unread output and current outcome. Use the ID returned by execute or listed in status.executions. Defaults to an immediate read; use wait_seconds=10 to wait for unfinished work. Returns content only: a [metadata] text block containing JSON execution_id, status (running/succeeded/failed/cancelled), truncated (output dropped since the previous read), and error (null or {type, message}), then text/images. If running, keep reading the same ID. A final outcome removes the record, including silent completions; do not read that ID again. No code is run.
Shut down the kernel and start a fresh one. Clears all in-memory state and aborts the active execution, if any. The configured Jupyter executable and kernel name are unchanged; the initial working directory is restored. Use to recover from kernel failures, unresponsive code, or resource exhaustion. Restart clears variables and files created during the session but restores disk state from the initial directory. Returns the updated status (state, error, configuration) after restart completes. A successful restart sets state=ready and error=null. Failure sets state=unavailable with an error; do not retry submit or read on an unavailable kernel—call restart again or inspect the error.
Inspect the kernel's lifecycle state, execution backlog, and availability. Returns a structured summary without executing code. Includes kernel configuration (Jupyter executable, kernel name, working directory), lifecycle state (ready/busy/unavailable/etc.), the currently active execution_id if any, a list of pending executions with unread output block counts, and a cumulative unread output count. Use to recover after a lost response or to monitor pending work without reading all output. A non-null error indicates a kernel failure; restart can recover it. state=ready means code submission is allowed. state=busy means read or interrupt the active execution first. Pending executions include both running work and unconsumed completed outcomes (within retention limits). Output blocks are counted after adjacent stream merging; metadata blocks are excluded.
Tool descriptions are lengthy (250-800 chars) and comprehensive for human reading, but exceed the 200-char sweet spot for LLM token efficiency. Descriptions could be condensed and delegate detailed semantics to INSTRUCTIONS or inline error messages.