A server built on the Model Context Protocol (MCP) that enables direct downloading of YouTube video transcripts, supporting AI and video analysis workflows.
The server provides 5 read-only tools for YouTube transcript extraction with generally clear names and descriptions. However, there are significant quality gaps: (1) tool descriptions are inconsistent in depth and some parameters lack detailed guidance; (2) schemas are present but parameter descriptions in the schema do not always match the natural-language descriptions in tool documentation; (3) no output schemas are documented, leaving LLMs to guess at response structure; (4) error handling is minimal, the code catches YouTubeTranscriptError but provides no recovery guidance; (5) tool composition is reasonable but there is semantic overlap between get_transcripts and get_transcript (both are aliases). The naming convention (verb_noun: get_transcripts, get_video_info) is solid. Descriptions average ~200 chars, which is within baseline (194 chars). However, parameter documentation is sparse in the schema itself, the descriptions 'Optional language code...' are too brief to guide LLM behavior on edge cases like missing language codes.
List available transcript languages for a YouTube video. Use this before retrying with a specific language code.
Extract a timestamped transcript from a YouTube video. Use this when the user needs quotes, chapter-like notes, or time-coded references. **Parameters:** - `url` (string, required): YouTube video URL or ID. - `lang` (string, optional): Language code for transcripts. If omitted, the best available caption track is used.
Alias of `get_transcripts` for compatibility with other YouTube transcript MCP servers. Extract and process transcripts from a YouTube video. **Parameters:** - `url` (string, required): YouTube video URL or ID. - `lang` (string, optional): Language code for transcripts. If omitted, the best available caption track is used. - `enableParagraphs` (boolean, optional, default false): Enable automatic paragraph breaks. **IMPORTANT:** If the user does *not* specify a language *code*, **DO NOT** include the `lang` parameter in the tool call.
Extract and process transcripts from a YouTube video. **Parameters:** - `url` (string, required): YouTube video URL or ID. - `lang` (string, optional): Language code for transcripts (e.g. 'en', 'uk', 'ja', 'ru', 'zh'). If omitted, the best available caption track is used. - `enableParagraphs` (boolean, optional, default false): Enable automatic paragraph breaks. **IMPORTANT:** If the user does *not* specify a language *code*, **DO NOT** include the `lang` parameter in the tool call. Do not guess the language or use parts of the user query as the language code.
Semantic overlap: get_transcripts and get_transcript are aliases with identical schemas and near-identical descriptions. This causes LLM confusion about which to call and violates the single-responsibility principle.
Output schemas are not documented. The tools return structured data (transcripts array, title, language, source) but LLMs have no schema to verify response shape. This forces LLMs to infer structure and risks parsing errors.
Parameter descriptions in schema are minimal ('Optional language code for transcripts (e.g. ...)') and do not explain LLM behavior edge cases. The natural-language description warns 'DO NOT include lang if not specified', but the schema description does not emphasize this.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Fetch basic YouTube video metadata and available transcript languages without returning the full transcript.
Error handling provides no recovery guidance. YouTubeTranscriptError is caught but responses lack actionable hints like 'Try calling get_available_languages(url) to see which languages are available.'
get_video_info and get_available_languages have short, vague descriptions (60-65 chars). When to call get_video_info vs get_available_languages is unclear, both retrieve language info. Description should explain the distinction and use cases.
get_timed_transcript schema includes enableParagraphs parameter (from copy-paste of getTranscriptsInputSchema) but the description never mentions it. This parameter may not be supported for timed transcripts.