MCP server for downloading YouTube videos, transcribing audio using OpenAI Whisper API, and managing transcription caching
The YouTube MCP server has three tools with mixed quality. Tool naming follows verb patterns (download_and_transcribe, transcribe_youtube, get_transcription_cache), but descriptions lack depth and parameter documentation is minimal. Input schemas are present but lack type information beyond basic strings. Output schemas are completely undocumented. Critical security issue: the server handles YouTube URLs but does not validate or sanitize them, and depends on an external cookies.txt file for authentication without documenting this dependency. Error handling is generic (HTTPException, generic try-catch blocks) with no recovery guidance for LLMs. The caching mechanism is present but not well-exposed in tool descriptions.
Download a YouTube video and transcribe its audio using OpenAI Whisper API. Returns cached transcription if available.
Retrieve a cached transcription by URL if it exists
Transcribe a YouTube video URL using OpenAI Whisper API with audio chunking and caching support
Tool descriptions are generic and lack actionable context. 'Download a YouTube video and transcribe its audio using OpenAI Whisper API' does not explain when to use download_and_transcribe vs transcribe_youtube, what the output contains, or prerequisites.
Input schemas present but incomplete. Parameters are typed as strings with descriptions, but no constraints (min/max length, format patterns, valid value enums) are declared. For 'language', no enum of valid language codes provided.
Output schemas completely undocumented. Tools return transcription text and cache UUIDs, but LLMs have no schema to understand the response structure, field names, or types. This violates the documented output schema requirement.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 19 | - | v1 |
No error handling guidance for LLMs. Code contains generic try-catch blocks with logger.error() but no structured error responses telling agents what to do next (retry, ask user, try alternative tool). Exception types are not categorized as retryable or fatal.
Security: Cookies file required but not validated in tool parameters. Server checks for cookies.txt existence at download time, but this dependency is not documented in tool descriptions. No input sanitization for YouTube URLs, could be vulnerable to command injection or SSRF if URL is passed unsanitized to yt_dlp.
Duplicate tool functionality. download_and_transcribe and transcribe_youtube perform nearly identical operations (download + transcribe), but descriptions do not clarify when to use one vs the other. This violates the single responsibility principle and confuses LLM tool selection.
Cache behavior not exposed in tool interface. get_transcription_cache exists but download_and_transcribe and transcribe_youtube mention caching in descriptions without exposing cache parameters or return fields. LLMs cannot reason about when cached results are returned vs new transcriptions generated.
No pagination or result limits documented. Transcription output could be arbitrarily long (entire video transcripts). No guidance on maximum token limits or pagination for responses that exceed context windows.
Missing parameter descriptions for 'language' parameter. Description states 'Language code for transcription (optional, defaults to auto-detect)' but does not specify format (ISO 639-1, ISO 639-3, BCP 47?) or provide examples. LLM will not know which format to use.