MCP server that wraps the Open Notebook API
Open-notebook-mcp has complete tool definitions with explicit schemas and descriptions for all 23 tools. However, quality varies significantly across tools. Descriptions are present but often generic (many are 40-80 chars, well below the 194-char baseline). Several parameters lack actionable constraints or format guidance. The tool set shows good naming conventions with clear verb_noun patterns (list_, get_, create_, update_, delete_, search_). However, parameter descriptions are sparse and lack context about constraints, formats, or when to use related tools. No tool annotations (readOnlyHint/destructiveHint) are present despite tools spanning READ_ONLY, WRITE, and DESTRUCTIVE risk levels. Error handling and recovery guidance are absent from all tool descriptions. The response schemas are not documented, forcing LLMs to infer output structure. This is a C-grade server with functional definitions but significant gaps in LLM-optimal description quality and parameter guidance.
Ask a question about your content with detailed control.
Ask a question about your content with simplified interface.
Create a new AI model configuration.
Create a new note.
Create a new notebook.
Create a new source (link, upload, or text).
Delete a model configuration.
Destructive tools (delete_notebook, delete_source, delete_note, delete_model) lack confirmation/dry-run pattern. No description mentions the irreversible consequences or guidance for recovery.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present despite explicit risk classification in metadata (READ_ONLY, WRITE, DESTRUCTIVE). LLMs cannot parse risk from the risk field alone.
Parameter descriptions are sparse and lack actionable constraints. E.g., 'Filter by archived status' (boolean) does not explain what happens when omitted; 'Order by field and direction' lacks examples of valid values (e.g., 'updated desc', 'created asc'). LLMs cannot infer valid formats from these generic descriptions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 27 | - | v1 |
Delete a note.
Delete a notebook.
Delete a source.
Get a specific model by ID.
Get a specific note by ID.
Get a specific notebook by ID.
Get a specific source by ID.
Get all configured AI models.
Get all notebooks with optional filtering and ordering.
Get all notes with optional filtering.
Get all sources with optional filtering.
Search content using vector or text search.
Search tools exposed by this server with progressive detail levels.
Update a note.
Update a notebook.
Update a source.
No output schemas documented in source code. Tool descriptions do not explain what fields are returned or their types. LLMs must infer response structure, risking incorrect downstream tool calls.
Error handling absent from all tool descriptions. No guidance on retryability, user-fixable errors, or recovery paths. E.g., delete_notebook description does not explain what happens if the notebook is not found.
ask_question and ask_simple tools both accept identical parameters and solve the same problem (ask a question). This creates redundancy and forces LLMs to decide between two equivalent tools. Consolidate or document the distinction clearly.
create_source parameter 'type' accepts free-form string with no enum constraint visible in the input schema. Description mentions 'link, upload, or text' but does not enforce them. LLMs may pass invalid types.
Pagination parameters (limit, offset) present in list tools but default values and bounds not documented. No mention of maximum limit or what happens if limit exceeds server max. This invites out-of-bounds requests.
Tool descriptions use passive voice and lack context on WHEN to use each tool vs. alternatives. E.g., 'Get all notebooks' (list_notebooks) vs 'Get a specific notebook' (get_notebook) distinction is clear, but no guidance on when search_capabilities should be called first for discoverability.
Topics parameter in update_source and update_note is an array of strings with no description of what 'topics' are, how they relate to the notebook structure, or validation rules. LLMs cannot infer the intent.