A reliable Model Context Protocol server for extracting YouTube video transcripts
The YouTube Transcript MCP server provides two focused, read-only tools with clear naming (verb-first: get_*, list_*) and reasonable descriptions (156 - 186 chars, within 34 - 392 baseline). Input schemas are present with type definitions and parameter descriptions. However, output schemas are entirely undocumented, the critical gap for Definition Quality. The server returns structured JSON (video metadata, segments with timing), but LLMs have no formal specification of the return structure, violating pattern:tool. Error handling is present but generic; descriptions lack dependency hints and recovery guidance. Parameter validation is minimal. These gaps prevent the score from reaching 70+.
Extract transcript from a YouTube video. Handles various URL formats and provides detailed error messages. Full transcript is saved to a temp file to reduce context usage.
List all available transcript languages for a YouTube video. Useful for debugging transcript availability.
Output schemas are completely undocumented. LLMs do not know the structure of get_transcript's return value (video_id, language, language_code, is_generated, segment_count, total_characters, segments array). Without formal output schema documentation, LLMs cannot confidently plan downstream parsing or field extraction.
get_transcript description does not explain when to use list_transcripts first or how to discover available languages. The description mentions 'defaults to en' but omits the dependency hint: 'If the preferred language is unavailable, call list_transcripts() first to see available options.'
Error responses are generic TextContent strings. The code returns 'Error: URL is required' and 'Error: {str(e)}', which do not guide recovery. Per pattern:recovery-guide, errors must suggest next steps: 'Invalid URL. Try a YouTube video link like https://youtube.com/watch?v=abc123 or call list_transcripts() to verify the video exists.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 63 | - | v1 |
The 'lang' parameter description includes example values ('e.g., en, de, fr') instead of relying on machine-readable constraints. Per pattern:constrained-input, if the tool supports only specific languages, enumerate them as an enum. Example values invite LLMs to hallucinate unsupported codes like 'english' or 'pt_BR'.
list_transcripts is a discovery tool but does not explicitly state its return structure or when an LLM should call it. The description says 'List all available transcript languages... Useful for debugging transcript availability' but omits: 'Returns a list of language codes and names. Call this first if you need a specific language other than English.'
No input validation metadata for error classes. The code catches TranscriptsDisabled, VideoUnavailable, and NoTranscriptFound but returns plain-text error strings. Per pattern:error-classification, errors should be categorized as retryable, user-fixable, or fatal. E.g., VideoUnavailable is fatal (do not retry); NoTranscriptFound is user-fixable (try list_transcripts or a different video).