MCP Server for extracting and managing YouTube video knowledge
This server demonstrates solid fundamentals with 44 well-named tools, consistent descriptions averaging 180 - 220 chars (within baseline), and tool annotations for all. However, schema completeness cannot be verified from the provided source: while tool definitions reference schema objects (e.g., buildBrainSchema, askBrainSchema), the actual JSON Schema definitions are not visible in the code sample. Descriptions are good but generic in places, lacking dependency hints and error recovery guidance. The per-tool inspection reveals naming is verb-forward and clear, descriptions answer WHAT and WHEN, but many lack actionable constraints or examples of when to call instead of similar tools. No visible schema validation rules, parameter type declarations, or output field descriptions in the sample. Error handling patterns are not evident from the registry code.
Search everything a creator has said, across every video in their brain. Returns the passages that match, each with the timestamp and a link that opens the video at that moment, so every claim can be checked. Matches the words as spoken, so phrase the query the way the creator would say it. Requires build_brain first.
Read a channel's videos into a searchable corpus of timestamped passages, so you can later ask what its creator has said about anything. Long-running: several hundred videos is several hundred fetches. Safe to interrupt and call again — it continues where it stopped, and on a finished brain it picks up new uploads. The since and minDurationSeconds filters describe the brain, not just the call: narrowing one discards the passages of the videos it excludes, and widening it reads them again. Ask it questions with ask_brain.
Report whether yt-dlp and ffmpeg are installed, their versions, and whether yt-dlp is stale. Call this first when tools start failing unexpectedly — an outdated yt-dlp is the most common cause.
Permanently delete a channel brain: its passages, its manifest and its profile. The transcripts it was built from stay cached, so rebuilding is much faster than the first build was.
Permanently delete every thumbnail saved for a channel, with its manifest. Use this to start over, for example to fetch again at a different quality.
Schema definitions not visible in source code. Tool registry references schema objects (buildBrainSchema, askBrainSchema, etc.) but actual JSON Schema type definitions, parameter types, constraints, and output structures cannot be verified from provided sample. Applied conservative 45 across all tools.
Parameter descriptions missing from visible source. While tool-level descriptions are present, the individual parameter descriptions that LLMs rely on to disambiguate 'type', 'mode', 'limit', and other inputs are not shown in the code sample. The Zod library is listed as a dependency, suggesting validation exists, but the description text for each parameter (required per pattern:tool-description) cannot be verified.
No visible error handling or recovery guidance. Tool descriptions do not include 'what to do if X fails' or error classification (retryable vs. user-fixable vs. fatal). For example, get_transcript says 'Results are cached locally' but does not explain what happens if a caption fetch fails, or whether the LLM should retry or call a different tool.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 41 | 2025-06-18+ | v2 |
Permanently delete a saved summary or skill note, or the entire library entry for a video. This removes files from disk and cannot be undone.
Summarize a playlist or channel in one call: per-video metadata, chapter markers, and optionally transcript word counts. Use this to survey a body of content before deciding what to read in full.
Download a YouTube video to local disk. Use the quality parameter for automatic format selection with smart fallbacks, or formatId for a specific format from list_formats. Returns the downloaded file path, title, and format details.
Write a video transcript to disk as SRT, WebVTT or plain text, ready to import into a video editor such as Premiere, Resolve or CapCut.
Cut a time range out of a video as audio only, in an editor-friendly format (mp3, m4a, wav, flac, opus). Use for podcast pulls and voice-over sourcing. Requires ffmpeg.
Cut a time range out of a YouTube video without downloading the whole thing. Give start and end, or a chapter name. Pair with search_transcript to find the moment first. Requires ffmpeg.
Cut several time ranges out of one video in a single call. A range that fails is reported individually rather than losing the clips that succeeded. Requires ffmpeg.
Capture a single still image from a video at a given timestamp, without downloading the file. Use for thumbnails and reference frames. Requires ffmpeg.
Save every video thumbnail of a channel, plus its avatar and banner, under ~/.youtube-knowledge/thumbnails. One listing per tab, then the images from YouTube's image hosts directly, largest available first; each saved image is recorded with the size actually decoded from it. Safe to interrupt and call again: images already on disk are kept, missing ones fetched. Choose tabs to include shorts and streams.
List videos from a YouTube playlist or channel. Returns video IDs, titles, durations, upload dates, and URLs. Sorted by playlist or channel order.
What a brain actually covers: how many videos were read, how many had no captions, how many are still outstanding, the channel's upload rhythm and speaking rate, and the phrases it repeats across videos. Read this before trusting an answer built from ask_brain.
Get metadata for a YouTube channel. Returns channel name, handle, subscriber count, description, and channel URL.
Extract chapter markers and timestamps from a YouTube video. Returns chapter titles with start and end times. Not all videos have chapters. Returns empty list if none found.
Read a sample of a video's comments, with replies when includeReplies is set. Returns a completeness receipt saying how many of the video's comments this is — the real total comes from a separate metadata read, because a comment fetch reports only what it extracted. This is a sample, not a page: YouTube exposes no cursor for a live comment read.
Report what this server can actually prove it holds: one completeness receipt per harvested target, plus whether ANY of them is incomplete. Reads only the local store and never contacts YouTube, so it is cheap to call before making any claim about the data. Check anyIncomplete before describing anything as a full history.
Read back a summary or skill note previously saved with save_to_library. Returns the markdown content plus the saved metadata.
Get metadata for a YouTube playlist. Returns title, channel, video count, last updated date, and description.
Return a video's thumbnail, or a channel's avatar or banner, as an image you can look at, with its real pixel size. Tries the largest image YouTube serves and falls back to smaller ones. Locally, an image saved by fetch_channel_thumbnails is served from disk.
Extract the full transcript from a YouTube video. Supports auto-generated and manual captions. Returns plain text with word count and detected language. Results are cached locally.
Fetch transcripts for up to 25 videos in one call, each capped so the batch cannot flood the context. Videos that have no captions are reported individually rather than failing the whole call.
Get detailed metadata for a single YouTube video. Returns title, channel, duration, upload date, view count, like count, comment count, description, tags, and thumbnail URL.
Catalogue every video a channel lists, across all of its tabs (videos, shorts, streams, releases, podcasts), into the local store. Returns a completeness receipt saying whether the catalogue is provably whole. Call it again to continue: videos are keyed on their id, so a re-run costs requests and never data.
Extract a video's comments and replies into the local store, with a completeness receipt. There is no comment cursor — an interrupted run keeps nothing for that video, so to get more, run it again with a larger maxComments. Re-running upserts on the comment id, so it costs requests and never data.
List every channel brain built locally, with how much of each channel is indexed and whether a written profile exists.
List the thumbnails saved for a channel: each file path, its decoded size and which image variant it is, plus the avatar and banner. Works offline from the saved manifest.
List all available download formats for a YouTube video. Returns format IDs, extensions, resolutions, FPS, codecs, and file sizes. Grouped by video+audio, video-only, and audio-only.
List all saved items in the local YouTube knowledge library. Returns titles, channels, content types, tags, and save dates. Optionally filter by tag. Sorted by most recently saved.
Remove harvested comments, catalogues, or one author's comments everywhere, from the local store. Requires confirm: true — harvested comments cost hours of network that no local cache can replay. Use the author filter to honour an individual's erasure request.
Search comments already in the local store, with full-text matching, real pagination and thread lookup. Never contacts YouTube. The total counts rows in the store, not the number of comments the videos have — the coverage receipts alongside it say which videos are only partly harvested.
Search the videos already catalogued in the local store, by channel, title, date, duration or how many comments are held. Never contacts YouTube. The coverage receipts alongside the results say which channels are only partly catalogued.
Rebuild the full-text search index from the notes on disk. Use this if search_library results look stale or incomplete, for example after editing files by hand.
Check the harvested store and, if it is damaged, move it aside and start a fresh one. The damaged file is kept, never deleted. Call this when a harvest tool reports STORE_CORRUPT.
Store a written account of a creator alongside their brain — voice, recurring arguments, how their thinking has changed. Write it from passages returned by ask_brain and cite them, because this server cannot check a claim it was handed. Overwrites any existing profile for that channel.
Save a summary or skill note to the local YouTube knowledge library. Overwrites existing content of the same type for the same video. Returns the saved file path.
Search YouTube for channels by keyword or phrase. Returns channel names, handles, subscriber counts, descriptions, and URLs.
Full-text search across every saved summary and skill note, ranked by relevance. Returns matching excerpts with the video IDs needed to read the full note.
Find a phrase or pattern inside a video transcript and return each match with its timestamp and a link that opens the video at that moment. Use this instead of reading a whole transcript when you need to locate or cite a specific moment.
Search YouTube for videos by keyword or phrase. Returns video IDs, titles, durations, channels, view counts, and URLs. Results sorted by relevance.
Add, remove or replace the tags on a saved library item. Tags are how list_library filters, so this is the way to reorganize a growing library. The replace parameter discards all existing tags.
Destructive tools lack confirmation or dry-run patterns. repair_store, delete_library_item, delete_brain, delete_channel_thumbnails, and prune_harvest are annotated destructiveHint:true but descriptions do not mention a confirm parameter, dry-run mode, or undo capability. Agents may inadvertently delete data without recovery.
Missing dependency hints and tool selection guidance. Descriptions say WHAT each tool does but not WHEN to call it instead of a similar tool. For example, search_videos vs. fetch_videos vs. query_videos are semantically similar but the descriptions do not clarify: search_videos is for free-form queries, fetch_videos requires a playlist/channel ID, query_videos searches the local harvested store. LLMs will waste reasoning cycles deciding between them.
Output structure not documented. Tool descriptions do not state what fields are returned or how to chain to downstream tools. For example, search_videos returns 'video IDs, titles, durations, channels, view counts, and URLs' but does not specify the JSON field names (video_id? id? videoId?), nor which fields downstream tools like get_video_info or extract_clip expect. Forcing the LLM to infer field names risks broken chains.
Pagination/limits not documented for result-returning tools. list_library, list_brains, query_comments, query_videos, list_formats, and search_library likely accept limit and offset parameters, but the descriptions do not mention max results, pagination strategy, or how to handle large datasets. LLMs will not know whether to expect 10, 100, or 10,000 results.
Tool annotations present but descriptions do not reinforce them. Most tools have idempotentHint and openWorldHint annotations, but the descriptions do not explicitly state 'Safe to retry' or 'Reaches YouTube to fetch live data' or 'Uses only local cache'. LLMs read descriptions before annotations, annotations alone do not convey intent.