MCP server for YouTube data extraction - transcripts, comments, and search
This server has solid naming conventions and schemas across all 5 tools. Tool names use action verbs (get, search) appropriately. However, descriptions lack the depth and specificity needed for production-grade LLM guidance. Many descriptions are under 100 characters and fail to explain WHEN to use each tool or dependencies between them. Parameter descriptions are present but often generic ('Language code (en, ko, ja, es, etc.)' instead of 'ISO 639-1 language code. Defaults to English (en). Supports: en, ko, ja, es, fr, de, pt, ru, zh, hi, ar'). Output schemas are not explicitly documented in code, only partially inferable from implementation. Error handling returns raw exceptions without recovery guidance. No tool annotations (readOnlyHint/destructiveHint/idempotentHint) despite all being read-only, which is a missed opportunity for agent reasoning.
Fetch replies to a specific YouTube comment. Use repliesToken from getComments response.
Fetch YouTube video comments with pagination. Supports sorting by relevance or time. Use nextPageToken for pagination.
Extract transcript/subtitles from a YouTube video. Returns full text with video metadata. Supports multiple languages and optional timestamps.
Get YouTube video metadata: title, views, publish date, channel, comment count, and pagination tokens for comments.
Search YouTube for videos, channels, and playlists. Returns results with thumbnails, views, and channel info.
Output schemas not documented. Tool descriptions mention return values (e.g. 'Returns full text with video metadata') but do not specify structured output format, field names, or data types. LLMs cannot plan downstream calls or extract data reliably without knowing response structure.
Descriptions lack LLM optimization guidance. Descriptions are 50-100 characters and do not explain WHEN to use each tool, dependencies, or prerequisites. E.g. getCommentReplies requires a repliesToken from getComments, which is stated once but not emphasized in either description. getVideoInfo does not explain it returns metadata needed for other calls.
Tool annotations missing. All 5 tools are read-only and safe to retry, but none carry readOnlyHint=true. This forces agents to reason conservatively about idempotency and retry safety, increasing latency and context waste.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | 2025-06-18+ | v2 |
| 2026-03-09 | D | 55 | - | v1 |
Parameter descriptions are generic and lack constraints. E.g. 'Language code (en, ko, ja, es, etc.)', does it accept only these 4? All ISO 639-1 codes? The ellipsis suggests more but does not enumerate them. Similarly, 'Max results (1-50)' and 'Max comments to return (1-100)' state bounds in prose rather than schema constraints, which LLMs may not reliably parse.
Error handling provides no recovery guidance. createErrorResponse() returns raw exception messages without actionable next steps. E.g. 'Invalid YouTube URL or video ID' gives the LLM nothing, should it retry? Search for the video? Ask the user? Compare to production pattern: 'Invalid video URL. Expected format: https://youtube.com/watch?v=VIDEO_ID (11 characters) or just the video ID. Try search_youtube() if you only have a title.'
Pagination not fully documented. getComments, getCommentReplies, and searchYoutube accept pageToken parameters, but tool descriptions do not explain the pagination protocol (does nextPageToken come in response? is it returned as a field or in metadata?) or how to detect end-of-results.
URL/ID handling inconsistency. Multiple tools accept 'url' or 'video ID' but do not specify what formats are valid beyond '11-character video ID' in getTranscript. Does getVideoInfo accept shortened URLs? Does searchYoutube accept youtube.com/watch?v=ID or only the ID itself? Ambiguity forces trial-and-error.