Expose VS Code Jupyter notebooks as MCP tools: read/edit/run cells and capture outputs (Cursor, Claude Code, Windsurf).
The server demonstrates solid tool naming conventions and consistent schema definitions across all 15 tools. Most tools have descriptions (12/15 ≥20 chars), and parameter schemas are well-structured with types and descriptions. However, there are critical gaps: (1) most output schemas are undocumented, tools like notebook_list_open, notebook_list_cells document what they *do* but not their response structure, forcing LLMs to infer output format; (2) error handling guidance is absent, no tool description mentions what to do on failure, recovery paths, or actionable error messages; (3) some parameter descriptions lack depth, especially around 'response_format' which appears 15 times identically without context on when to use markdown vs JSON; (4) no tool has destructive/readonly/idempotent annotations despite clear risk classifications (7 tools marked WRITE/DESTRUCTIVE); (5) the 'notebook_uri' parameter is repeated on 13 tools with identical description, creating redundancy that could be abstracted into a shared pattern or global context. The server is competent but not production-grade, descriptions are functional but not LLM-optimized for selection and reasoning.
Insert multiple cells into the notebook at once.
Clear the outputs of a specific cell.
Delete a cell from the notebook.
Edit the content of an existing cell by index.
Get the full source code of a specific cell. Args: - index (number): Cell index (0-based) - response_format ('markdown' | 'json'): Output format (default: 'markdown')
Get the outputs of a specific cell (text, errors, images). Args: - index (number): Cell index (0-based) - response_format ('markdown' | 'json'): Output format (default: 'markdown')
Output schemas undocumented for all 15 tools. No tool declares what fields its response contains, response types, or pagination structure. LLMs must infer output format from descriptions, causing hallucinations and downstream tool-chain failures.
No error handling guidance in any tool description. Tools performing destructive operations (delete_cell, edit_cell, run_cell) provide no recovery path or actionable error messages. Agents cannot determine if failures are retryable, user-fixable, or fatal.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 76 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 55 | - | v1 |
Get the current kernel context including variables, imports, and recent execution history. This tool provides Claude with awareness of the notebook's runtime state, enabling more informed code generation. Args: - include_variables (boolean): Include current kernel variables (default: true) - include_history (boolean): Include recent cell execution history (default: true) - max_value_length (number): Max characters for variable value representation (50-500, default: 100) - response_format ('markdown' | 'json'): Output format (default: 'markdown')
Get information about the active notebook's kernel. Returns kernel status, language, and notebook URI. Args: - response_format ('markdown' | 'json'): Output format (default: 'markdown')
Get a structured outline of the notebook showing markdown headers and code cell summaries.
Insert a new cell into the notebook at a specified position.
List all cells in the active notebook with metadata including type, language, content preview, and execution state.
List all open notebooks with their URIs, filenames, cell counts, and which one is currently active (focused).
Move a cell from one position to another within the notebook.
Execute a cell and wait for the result (with timeout).
Search for text within the notebook with configurable case sensitivity and context.
Tool annotations missing. 7 tools are marked with risk classifications (WRITE, DESTRUCTIVE, REVERSIBLE, READ_ONLY) but no tool registers with readOnlyHint, destructiveHint, or idempotentHint annotations in the MCP schema. LLMs cannot see risk classification from tool definitions alone.
Parameter 'response_format' repeated identically on 15 tools with generic description 'Output format: markdown for human-readable or json for structured data.' No guidance on when to choose each format, LLMs cannot reason about this parameter effectively.
Parameter 'notebook_uri' repeated on 13 tools with identical description but no constraint on format (is it a file path, a URI scheme, a notebook name?). This forces LLMs to rely on tool discovery and error messages to understand valid values.
notebook_run_cell and notebook_list_cells descriptions lack detail on runtime behavior. 'Execute a cell and wait for the result (with timeout)', what timeout value? What happens on timeout? notebook_list_cells returns 'metadata', which metadata fields?
No pagination guidance for list tools. notebook_list_open and notebook_list_cells may return hundreds of notebooks or cells, but no limit, offset, or cursor parameters are defined. Large result sets will exhaust context windows.