No-fuss YouTube transcripts for MCP - no API keys required! A Model Context Protocol server to get YouTube video transcripts.
YouTube MCP has solid tool naming and generally complete schemas, but suffers from overly verbose descriptions, missing output schema documentation, weak error handling guidance, and lack of per-tool structure hints. All 5 tools are explicitly registered with Zod schemas and descriptions present, but descriptions exceed production baselines (avg 284 chars vs baseline 194) and lack actionable error recovery paths. No tool annotations (readOnlyHint, destructiveHint) despite claiming READ_ONLY risk. Output responses are JSON-stringified text blobs without documented structure, forcing LLMs to parse unstructured data. No pagination or result-limiting guidance despite tools like search_videos and get_channel_videos returning potentially large lists.
Retrieves a list of videos from a specified YouTube channel. This tool is useful for getting all videos uploaded by a specific channel.
Retrieves the full transcript of a specified YouTube video. This tool is useful for understanding video content without watching it, or for extracting textual information from videos. FORMATTING GUIDANCE (optional - user instructions override): When creating summaries, consider using: **Key Points with Timestamps:** Use [MM:SS] or [HH:MM:SS] inline references. **Structure:** Break into logical sections. **Context:** Include video title and channel. Example: 'The speaker explains TypeScript generics [05:30] and shows practical examples [08:15].' This formatting is optional - always follow any specific user instructions instead.
Captures a screenshot from a YouTube video at a specified timestamp. Requires yt-dlp for full-quality captures; falls back to lower-resolution YouTube storyboards (320x180 max) if yt-dlp is not installed.
Searches YouTube for channels matching the specified query. You can specify a sort order for the results (default: rating).
Searches YouTube for videos matching the specified query. Returns a list of video results with title, video ID, and channel information. Results are sorted by rating by default for better quality content.
Tool descriptions are excessively verbose (avg 284 chars, baseline 194). The get_transcript description includes 255+ chars of formatting guidance that belongs in documentation, not the tool spec. LLMs waste tokens parsing narrative guidance instead of receiving concise, action-focused descriptions.
No output schemas documented. All tools return `{ type: 'text', text: JSON.stringify(...) }`, raw JSON strings. LLMs cannot plan downstream chains or extract specific fields without parsing unstructured text. Example: search_videos returns a JSON string; the agent has no way to know it includes 'videoId', 'title', 'channelId' fields without executing and inspecting.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
No error handling guidance or recovery paths. If get_transcript fails (video unavailable, no captions, network timeout), the tool returns only a raw error. No message telling the LLM to retry, call search_videos, or accept plaintext fallback. Silent failures and opaque error messages disable agent self-correction.
Missing pagination and result-limit guidance. search_videos, search_channels, and get_channel_videos can return many results (get_channel_videos caps at 200, others unbounded). No mention of pagination, limit defaults, or expected response size. An LLM naively calling search_videos with a broad query could flood context or timeout.
No tool annotations despite claiming READ_ONLY risk classification. Tools should declare `readOnlyHint: true` to signal idempotency and no side effects. This allows agents to safely retry without fear of duplicates or state changes. Current FastMCP version supports annotations; they are simply not used.
Parameter chunkSize, silenceThreshold, and quality lack numeric constraints (min/max). chunkSize could be 0, negative, or absurdly large (1GB), causing failures or memory exhaustion. silenceThreshold could be negative. quality is free-form string without guidance on valid yt-dlp format selectors. No validation constraints in schema or description.
Optional parameters mix boolean flags (chunkBySilence, skipSponsor, plainText) without clear interdependencies. The description says 'chunkBySilence is true, this specifies silenceThreshold' but does not warn that silenceThreshold is ignored if chunkBySilence is false. Undocumented dependencies cause LLMs to pass conflicting option combinations.
Output field naming may not match input field naming. If search_videos returns a JSON array with fields like 'videoId', 'title', 'channelId' but get_transcript expects a 'videoUrl' (human-friendly), there is a mismatch. LLMs must reason about field mapping (e.g., constructing a URL from a videoId). The response structure is not documented, so this cannot be verified, a strong signal of missing schema documentation.