The server provides 4 well-named tools with complete input schemas and descriptions. Tool names follow verb_noun convention (fetch_*, list_*) and are action-oriented. All parameters have type definitions and descriptions. However, output schemas are not documented in the source code, parameter descriptions lack depth and constraint details, and error handling guidance is minimal. The implementation includes defensive practices (input sanitization, thread-safe client creation, custom error handlers), but the tool definitions themselves don't expose enough structure for optimal LLM reasoning.
Tools (4)
fetch_metadataread only50/100
Fetch metadata about a YouTube video
fetch_playlistread only50/100
Fetch a playlist of videos from YouTube
fetch_transcriptread only50/100
Fetch the transcript of a YouTube video
list_transcriptsread onlysource verified70/100
List all available transcripts for a YouTube video
Output schemas not documented. Tool descriptions state what is returned (e.g. 'Fetch the transcript') but do not specify the structure of returned data (fields, types, formats). LLMs cannot plan downstream tool calls or extract chaining IDs without knowing what fields to expect.
Parameter descriptions lack actionable constraints. The 'format' enum for fetch_transcript is documented, but 'languages' array lacks guidance on ISO 639-1 code format, what happens if an invalid language is requested, or how the default is determined. The 'limit' integer in fetch_playlist has no min/max bounds specified.
Error handling does not guide LLM recovery. The code implements _handle_transcript_error() with human-friendly messages (e.g. 'No transcript found... Use list_transcripts to see available languages'), but these are returned as plain strings without structure. LLMs cannot reliably parse suggestions or understand whether an error is retryable, user-fixable, or fatal without explicit classification.
Recommendations
Document output schemas for all four tools. For fetch_transcript, explicitly describe the structure: 'Returns an array of transcript objects, each with: {start: number (seconds), duration: number (seconds), text: string}. Formatted according to the requested format parameter (json, text, webvtt, srt, or pretty).' For list_transcripts: 'Returns an array of available transcript objects with language, language_code (ISO 639-1), and is_generated (boolean) fields.' For fetch_metadata: 'Returns a metadata object with fields: title, description, channel, channel_id, channel_url, upload_date (YYYY-MM-DD), duration_seconds, view_count, like_count, tags (array), categories (array), thumbnail_url, chapters (array), age_limit, is_live, webpage_url.' For fetch_playlist: 'Returns an object with: playlist_id, playlist_title, videos (array with video_id, title, channel, duration), total_count (integer), requested_limit.'
Add constraint details to parameter descriptions. For fetch_transcript 'languages': 'ISO 639-1 language codes (e.g., en, es, fr). If not provided, defaults to the video\'s default language. If specified languages are unavailable, an error is returned with available options.' For fetch_playlist 'limit': 'Maximum number of videos to fetch (1-50, default 20). Larger limits may incur higher latency.'
Implement structured error responses. Instead of returning error strings, return a JSON object: {"error": "TranscriptNotFound", "message": "No transcript found...", "recovery": "Call list_transcripts() to see available languages", "retryable": false}. This enables LLMs to parse error type and follow guidance programmatically.
Tool descriptions are generic and lack discovery context. 'Fetch the transcript of a YouTube video' does not explain when to call this vs list_transcripts, what happens if a transcript is unavailable, or what the expected latency is. This forces the LLM to reason through tool selection rather than reading clear intent signals.
No pagination guidance for list_transcripts or fetch_playlist. fetch_playlist accepts a 'limit' parameter but does not document: what is the maximum limit? Is there a default? Does it return a next_cursor or page token for continuation? Without pagination structure, the LLM cannot reliably fetch large results.
list_transcriptsfetch_playlist
Expand tool descriptions with discovery context. For fetch_transcript: 'Retrieves the full transcript for a YouTube video. Call list_transcripts() first to check available languages and whether the transcript is auto-generated. Transcripts may be slow to fetch (5-15s) depending on video length and API availability.' For list_transcripts: 'Lists all available transcripts for a video, including language, language code, and whether the transcript is auto-generated. Use this to verify a language is available before calling fetch_transcript.'
Add pagination structure to fetch_playlist. Extend the 'limit' parameter description: '(default 20, max 50)'. Extend the output description to include 'has_more: boolean' and 'next_page_token: string (if has_more is true, pass this to a follow-up call to get the next batch).' This enables continuation of large playlists.
Document timeout behavior and retry guidance. Add a note to tool descriptions: 'This tool calls external YouTube APIs and may timeout after 30 seconds. If a timeout occurs, retry with a smaller limit or fewer languages.' This helps agents decide whether to retry or escalate.
Add tool annotations (readOnlyHint) to indicate these are safe, read-only operations. This signals to the agent that retries are safe and that no data loss is possible.