An MCP server implementation that integrates with yt-dlp, providing video and audio content download capabilities (e.g. YouTube, Facebook, Tiktok, etc.) for LLMs.
A solid MCP server with 10 well-defined tools. All tools have clear descriptions (119-234 chars on average) and comprehensive input schemas with types, enums, patterns, and ranges. Naming is consistent (verb_noun convention: ytdlp_<action>_<noun>). However, several tools have undocumented output schemas, and error handling lacks recovery guidance. The server demonstrates production intent with proper schema validation (per test regression in tests/test-tools-schema.mjs) and tool annotations (readOnlyHint). Main gaps: no documented return schemas, limited per-parameter guidance on format/constraints in descriptions, and error responses likely lack actionable recovery paths. Strong foundation with room for output documentation and error UX.
Download audio from a video URL in the best available quality.
Download transcript or captions from a video.
Download a video from a URL in the specified resolution with optional trimming.
Download subtitles for a video in a specified language.
Retrieve comments from a video with support for sorting, pagination, and threading.
Get a concise summary of video comments in human-readable format.
Retrieve detailed metadata for a video including title, description, duration, uploader, and more.
Output schemas not documented. While input schemas are comprehensive and valid JSON Schema, tool responses lack documented return types. LLMs cannot plan downstream operations or extract specific fields without knowing what structure to expect.
Brief parameter descriptions lack actionable constraints. Parameters like 'resolution' in ytdlp_download_video describe WHAT but not FORMAT or EDGE CASES. Descriptions should state: 'Preferred video resolution; if unavailable, falls back to best available. Enum: 480p, 720p, 1080p, best.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Get a concise summary of video metadata in human-readable format.
List all available subtitle languages for a given video.
Search for videos on YouTube using keywords with pagination and date filtering support. This tool queries YouTube's search API and returns matching videos with titles, uploaders, durations, and URLs. Supports pagination for browsing through large result sets and filtering by upload date.
ytdlp_list_subtitle_languages description is only 59 chars ('List all available subtitle languages for a given video.'). Lacks guidance on when to call it, what the LLM should do next, or whether subtitles are guaranteed to exist.
ytdlp_download_audio description is 68 chars and omits key details: does this require the URL to be a video? Does it auto-select bitrate or allow configuration? Can the LLM specify format (mp3, aac, m4a)?
No documented error handling or recovery guidance. If a download fails (e.g., video region-locked, account suspended), the tool likely returns a raw yt-dlp error. LLMs need actionable messages like 'Video unavailable in your region. Try a VPN or request region-specific mirror.'
No dry-run or confirmation for WRITE operations. ytdlp_download_video and ytdlp_download_audio modify filesystem. Agents should be able to preview what will happen without committing. A confirmation-request pattern prevents accidental mass downloads.
ytdlp_search_videos accepts maxResults 1-50 (reasonable), but offset parameter lacks documented behavior at boundary. What happens if offset=1000 but total results=50? Does it return empty array or error? Should the description state 'Returns empty if offset exceeds result count'?
URL parameter validation relies on format='uri'. While JSON Schema format hints are good, descriptions should state: 'YouTube, Facebook, TikTok, and other supported yt-dlp platforms. For YouTube: use video ID, URL, or playlist URL.'