MCP proxy for SubDownload — YouTube knowledge base for AI agents. Exposes SubDownload's hosted MCP server as a local stdio MCP server suitable for Docker deployment.
SubDownload MCP demonstrates solid definition quality with well-structured tool schemas, clear naming conventions, and comprehensive descriptions. All 9 tools follow verb_noun patterns (search_, fetch_, get_, list_). Descriptions are detailed and contextual, ranging 100-250 chars, explaining WHAT each tool does, WHEN to use it, and key caveats. Input schemas use proper JSON Schema with type definitions, enums, and minLength/maximum constraints. Tool annotations are present and well-justified (readOnlyHint, destructiveHint, idempotentHint, openWorldHint). However, there are gaps in output schema documentation, response structures are not formally declared in the visible code. Error handling mentions exist (e.g., 'Errors with NO_CAPTIONS if the video has no captions, fall back to transcribe_video') but lack systematic recovery guidance. Parameter descriptions are strong but a few could be more specific about valid formats (e.g., 'lang' parameter says 'ISO 639-1' but does not show an enum or examples for common codes).
Fetch the existing official transcript (subtitles/captions) of a YouTube video, with per-segment timestamps and language detected. Errors with NO_CAPTIONS if the video has no captions — fall back to transcribe_video in that case to generate one with AI ASR. This call is free.
Fetch consolidated YouTube video metadata with numeric types — title, channel, duration, view count, publish date, thumbnail, description, captions availability. Does NOT include the transcript itself; call fetch_transcript or transcribe_video for that. Cheap, fast, free.
Poll the status of an ASR task created by transcribe_video. Returns one of `queued`, `downloading`, `transcribing`, `finalizing`, `done`, or `failed`. When status is `done`, includes the full transcript with timestamps. Recommended polling interval: 3-5 seconds. Free — does not consume credits.
Get the most recent videos from a YouTube channel — convenience wrapper over list_channel_videos with no pagination. Best for 'what did this creator publish recently?' style queries.
List all videos from a YouTube channel ordered by publish date (newest first), with pagination. Returns up to 30 per page plus a `continuation` token if more results exist. For just the most recent handful, prefer get_channel_latest_videos for simplicity.
Output schemas are not documented. While input schemas are comprehensive with types, descriptions, and constraints, the visible code does not declare what fields or structure each tool returns. LLMs cannot predict downstream tool chains without knowing what data to expect.
Error handling lacks systematic recovery guidance. fetch_transcript mentions 'Errors with NO_CAPTIONS if the video has no captions, fall back to transcribe_video' but this pattern is not consistently applied. Other tools do not document what errors are possible, whether they are retryable, or what alternatives exist.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 65 | 2026-07-28+ | v2 |
Resolve a YouTube @handle, channel URL, video URL, or raw channel ID into canonical channel info (channel ID, name, handle, subscriber count, video count, avatar). Call this first when you only have a handle or URL but need a channel ID for the other channel-scoped tools.
Search for specific videos within a single YouTube channel. Restricts results to the given channel. Use after resolve_channel if starting from a handle. Useful for 'find Karpathy's video about backpropagation' style queries.
Search YouTube globally for videos, channels, or playlists on any topic. Returns up to 50 results with metadata. Use this for topic-based discovery when the user has not specified a channel — for searching within a known channel use search_channel_videos instead.
Start an asynchronous AI ASR (Whisper) transcription of a YouTube video. Returns immediately with a task_id and estimated_wait_seconds; the actual transcription runs in the background. Poll status with get_asr_task. Use this when fetch_transcript returned NO_CAPTIONS or when the video has no captions. Costs 5 credits, debited only on successful completion.
Language parameter (lang) uses informal description 'ISO 639-1 language code to select among multilingual captions (e.g. 'en', 'zh', 'ja')' without an enum constraint or pattern. LLMs may invent language codes or pass incorrect values. Should be either an enum of supported codes or a formal pattern constraint.
pagination mechanism in list_channel_videos uses 'continuation' token but does not specify max result count per page, format of the continuation token, or how to detect end-of-results. Description says 'Returns up to 30 per page' but this is not enforced or documented in the schema.
Parameter 'limit' in search_youtube has max=50 but description says 'Default: 20' without declaring a default in the schema. LLMs will not apply the default, they will pass no value if limit is omitted, relying on server-side behavior that is not formally declared.