MCP server that exposes YouTube video metadata extraction, transcript retrieval, and Whisper-based transcription capabilities as tools
The server provides three well-named YouTube tools with clear descriptions that guide LLM selection. All tools have input schemas with type definitions and parameter descriptions. Output schemas are partially documented via struct tags. Tool annotations are present (toolAnnotations=true). However, there are gaps in error handling guidance, output schema documentation is incomplete, and some parameter descriptions lack sufficient detail about formats and constraints. The descriptions are generally good (120-170 chars) but could be more prescriptive about recovery paths.
Extract video metadata including caption availability. Check 'Has Captions' field to determine which transcript tool to use: if true, use get_youtube_transcript (free); if false, consider transcribe_youtube_whisper (paid).
Get existing YouTube captions/transcript (FREE). Only works if the video has captions - check metadata first. Fails if no captions available.
Create transcript using OpenAI Whisper API (PAID). Requires OPENAI_API_KEY environment variable to be set. Use only when videos have no captions and user explicitly agrees to incur costs. Always ask user for confirmation before calling this tool.
Output schemas are declared via struct tags (jsonschema) but not explicitly documented in tool registration or descriptions. LLMs cannot see the output structure (mcpMetadataOutput, mcpTranscriptOutput) to plan downstream usage.
Error handling is implicit. Descriptions mention failure conditions (e.g., 'Fails if no captions available' for get_youtube_transcript, 'Requires OPENAI_API_KEY' for Whisper) but provide no recovery guidance. No recovery_guide pattern implemented.
Parameter 'include_timestamps' in get_youtube_transcript lacks format/constraint details. Parameter 'include_timestamps' in transcribe_youtube_whisper states 'Reserved for future use' but is still present and optional, unclear if LLMs should use it.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2025-06-18+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Tool composition assumes sequential discovery: metadata first (to check has_captions), then transcript or Whisper. This is documented in metadata description but forces multi-step planning. No batch/composite variant exists.
The Whisper tool description warns 'Always ask user for confirmation before calling this tool' but MCP provides no confirmation/dry-run mechanism. This places the burden on the agent/client to implement confirmation logic.