A TypeScript MCP server implementation for extracting transcripts from YouTube videos
The server defines 2 tools with explicit schemas via Zod, descriptions, and some output structure. However, there are significant gaps: (1) Output schemas are partially defined but incomplete, ToolTranscribeYoutubeOutputSchema is declared but the GET_TRANSCRIPT output is inferred; (2) parameter descriptions are minimal (single short phrases, not actionable); (3) error handling is generic (handleError delegates to a centralized function with no visible recovery guidance); (4) no tool annotations (readOnlyHint, destructiveHint); (5) no enumerated constraints despite opportunities (e.g., status values, language codes). The transcribe_youtube tool is destructive (WRITE risk) but has no dry-run or confirmation step. get_transcript is read-only but lacks clarity on pagination or result limits. Naming is adequate (verb_noun convention) but descriptions are too terse to guide LLM selection effectively.
Get existing transcript by video ID
Extract transcript from YouTube video with progress reporting and save to configurable folder
Destructive tool (transcribe_youtube, WRITE risk) lacks dry-run or confirmation mechanism. Agents may inadvertently overwrite transcripts or create duplicates without user approval.
Output schemas incomplete. ToolTranscribeYoutubeOutputSchema includes 'next_action' field (suggesting multi-step composition) but does NOT fully document the return type. ToolGetTranscriptOutput is inferred from usage (readTranscriptFile return), not explicitly visible in the provided code. Missing documentation of return fields (transcript array structure, metadata fields) forces LLMs to guess.
Parameter descriptions are terse single-line phrases ('YouTube video URL', 'YouTube video ID'). These fall below the 50-char minimum.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Tool descriptions lack actionable context. Current descriptions (45-65 chars) state WHAT the tool does but NOT WHEN to use it or what happens next. For example, transcribe_youtube says it 'extracts transcript' but does not clarify: (a) it modifies server state (saves files), (b) it may fail if video has no subtitles, (c) the next logical action is to call get_transcript. Minimal descriptions reduce LLM confidence in tool selection.
No tool annotations. Neither transcribe_youtube (destructive, should have destructiveHint=true) nor get_transcript (read-only, should have readOnlyHint=true) declares MCP tool annotations. These hints are critical for agents to reason about side effects and retry safety.
Error handling is opaque. The CallToolRequestSchema handler calls handleError(error) on failures, but error responses are not visible in the provided code. No evidence of 'what to do next' guidance, categorization (retryable vs. user-fixable vs. fatal), or self-correcting hints. LLMs receive raw exceptions with no recovery path.
No pagination or result limits documented for get_transcript. If a transcript contains hundreds of segments, the full JSON will be returned in text and structuredContent. This can bloat the context window and degrade LLM reasoning. Baseline pattern requires pagination and limits (cap 20-50 items) with a 'next_cursor' or similar.
Parameter 'url' in transcribe_youtube lacks format constraints. No regex pattern, length limits, or clarification of accepted formats (e.g., youtube.com/watch?v=..., youtu.be/..., or playlist URLs). LLMs may pass invalid or ambiguous values.
No natural-identifier fallback. get_transcript requires a 'videoId' (system ID), not a human-readable title or channel name. If an LLM only knows the video title, it must call transcribe_youtube first to extract the ID, forcing an unnecessary extra call. Per pattern:tool-chain, lookup tools should accept human-friendly identifiers.