MCP server for getting transcript/subtitles from youtube videos using yt-dlp
The server implements 3 YouTube-focused tools with moderate definition quality. All tools have names starting with action verbs (get_*), descriptions present, and basic input schemas. However, several gaps reduce overall quality: (1) Two tools share nearly identical naming and responsibilities without clear distinction (get-youtube-video-transcript-and-title vs get-youtube-video-title-only), violates pattern:tool composition rule requiring singular responsibilities; (2) Parameter descriptions are minimal and lack detail about constraints, expected formats, or when to use each tool; (3) Output schemas are completely undocumented, LLMs cannot predict response structure; (4) Error handling is basic (error messages exist but lack recovery guidance); (5) Tool descriptions, while present, fall short of the 50-200 char LLM-optimized range and lack dependency hints (e.g., when to call title-only vs transcript-and-title).
Get comments from a youtube video
Get the title of a youtube video
Get transcript and title from a youtube video. In stdio mode, responses longer than 15 KB are split into chunks. If the response starts with '--- Part 1/', call this tool again with chunk=2, chunk=3, etc. to get the remaining parts.
Two tools operate on the same resource (YouTube transcript/title) with overlapping responsibilities without clear distinction. get-youtube-video-transcript-and-title returns both transcript+title; get-youtube-video-title-only returns only title. Tool descriptions never explain WHEN to call each, forcing LLMs to guess or test both.
No output schemas documented for any tool. LLMs cannot predict the structure of responses (fields, types, array vs object). This prevents planning of downstream operations and forces exploratory calls. For chunked transcript tool, the chunking protocol (--- Part 1/N syntax) is mentioned in description but not formalized in schema.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter descriptions are minimal or missing. Example: 'sortby' enum in get-youtube-video-comments has description 'Sort order for comments' but never explains what 'top' vs 'newest' means semantically. LLMs cannot infer intent from vague descriptions.
Tool descriptions lack dependency hints and multi-tool guidance. No description explains whether to use get-youtube-video-transcript-and-title or get-youtube-video-title-only first, or when comments are available vs transcripts. This prevents agents from planning efficiently.
Error handling is basic. Fetch errors return strings like 'Error fetching subtitle: <message>' or 'No transcript available.' These errors do not tell LLMs what to do next (retry, try alternative video, ask user, etc.) or categorize them as retryable vs fatal. Chunking logic returns 'Chunk X out of range' but does not guide the LLM to stop requesting.