An MCP server that wraps Google NotebookLM, exposing notebooks, sources, and chat as tools.
The NotebookLM MCP server has 6 tools with visible schemas and descriptions. All tools start with action verbs (list_, create_, add_, chat_) which is positive. However, the server has critical weaknesses: (1) Parameter descriptions are universally MISSING or bare (notebook_id, url, text, query have no descriptions beyond the parameter name itself). (2) Output schemas are not documented, return types are inferred from code (list[dict], dict, str) but no field-level documentation exists. (3) No error handling guidance is provided, tools will fail if API calls error but the responses lack recovery instructions. (4) Input validation is absent, parameters are raw strings with no length limits, format constraints, or enum declarations. (5) Tool descriptions are concise but lack the LLM-optimized detail required by the rubric (WHAT, WHEN, prerequisites). The server would struggle in production: an LLM cannot reason about output structure without documented schemas, and cannot correct errors without recovery guidance. The code shows async context manager usage and clean API wrapping, which is positive, but definition quality lags significantly behind production expectations.
Add raw plain text as a source to a NotebookLM notebook.
Add a web URL as a source to a NotebookLM notebook.
Ask a question against a NotebookLM notebook. The AI answers using the current sources in the notebook.
Create a new NotebookLM notebook.
List all NotebookLM notebooks for the authenticated user.
List all sources in a NotebookLM notebook.
Parameter descriptions are missing or empty. 'notebook_id', 'url', 'text', 'query' parameters in the input schema have no descriptions beyond the parameter name itself. LLMs cannot determine parameter purpose, valid values, or format constraints without explicit descriptions.
Output schemas are completely undocumented. Return types are inferred from code (list[dict], dict, str) but no field-level schema is provided. For example, list_notebooks returns [{'id': ..., 'title': ...}] but the schema for 'id' and 'title' fields is not declared. LLMs cannot plan downstream tool calls or extract data without knowing output field names and types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
No error handling or recovery guidance. Tools will fail if the NotebookLM API returns errors (e.g. invalid notebook_id, network failure, authentication failure) but no error messages explain what went wrong or what the LLM should do next. Per the rubric, error responses must tell the agent what to do: 'Notebook not found. Call list_notebooks() to see available notebooks.'
No input validation or constraints. Parameters like 'url' and 'text' accept arbitrary strings with no length limits, format validation, or enum constraints. For example, add_source_url does not verify the URL is valid HTTP(S) before passing it to the API. Invalid input should be rejected early with actionable error messages.
Tool descriptions lack LLM-optimization context. Descriptions like 'List all sources in a NotebookLM notebook' omit WHEN to call the tool, what it returns, and dependencies. Per the rubric, descriptions should answer: What does it do? When should the LLM call it? What does it return? Incomplete docstrings force LLMs to guess.
list_sources and list_notebooks do not support pagination. If a user has hundreds of notebooks or sources, all are returned in a single response, which could exhaust the context window. Tools returning lists should accept page/offset and limit parameters and return a total count.
Inconsistent output field extraction. add_source_url and add_source_text use getattr(source, 'id', str(source)) to handle missing fields, suggesting the underlying API may not always return 'id' or 'title'. This defensive coding indicates an implicit dependency on the notebooklm-py library's schema that is not documented. If fields are optional, the MCP server should explicitly document which fields may be missing and what the fallback values mean.