MCP server for extracting YouTube video transcripts
This server demonstrates solid definition quality with all four tools having proper schemas, clear descriptions, and good parameter documentation. Naming follows verb_noun convention well. However, several opportunities for improvement exist: output schemas are not explicitly documented, error handling lacks recovery guidance, and parameter relationships could be clearer. The tools are read-only and relatively straightforward, which limits complexity but also means some patterns (error classification, destructive operation safeguards) do not apply. Average description length is 85 characters (within baseline), and all tools have type-defined parameters with descriptions.
Extract frames from a YouTube video at specified timestamps
Extract the transcript from a YouTube video
Extract the transcript from a YouTube video as structured JSON
Search a YouTube transcript by keywords or regex
Output schemas are not explicitly documented. Tools return structured data, but LLMs cannot determine the exact field names, types, and nesting without seeing the response schema. This forces LLMs to guess at downstream field mappings and risks errors when chaining tools.
Error handling lacks recovery guidance. The code returns RuntimeError and generic exceptions without telling the LLM what to do next (e.g., 'Try with different language codes', 'Verify video is public'). Error responses should be actionable.
extract_frames has vague output description. The tool returns frames 'at specified timestamps' but does not document the frame format (Base64 image data, file paths, Image objects, etc.), resolution, or how they are indexed in the response array.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
get_transcript_json description differs minimally from get_transcript. The description says 'Extract the transcript from a YouTube video as structured JSON', but does not explain why an LLM should choose this over get_transcript (what does 'structured JSON' mean when both are called via JSON-RPC?). Descriptions should clarify the distinction.
Parameter dependencies not documented. search_transcript has 'regex' parameter, the description says 'If true, interpret query as a regex pattern' but does not explain what happens if the regex is invalid, or what error the LLM should expect. Dependencies between parameters (query format depends on regex flag) should be explicit.