MCP server for extracting YouTube video transcripts, metadata, and comprehensive video information using headless-youtube-captions
Server provides 7 tools with explicit schemas and descriptions. Tool naming follows verb_noun conventions consistently (get_*, search_*). Descriptions are present and moderately detailed (80-200 chars), meeting baseline minimums. However, several critical gaps emerge: (1) output schemas are NOT documented anywhere, the code shows ListToolsRequestSchema returns tool definitions but actual response structures are invisible; (2) parameter descriptions lack implementation details (e.g., what happens on invalid language code? What format does 'segment' use?); (3) error handling is not evident in the visible code, no recovery guidance, no error categorization; (4) the server uses an external library (headless-youtube-captions) with @ts-ignore, obscuring actual return types; (5) no pagination guidance despite tools like get_channel_videos accepting maxVideos (suggests potential unbounded results). Tool composition is sound, each tool has a single responsibility, and naming is clear. Parameter constraints are present (enums for sortBy, limits for maxResults). However, the absence of documented output schemas and error recovery patterns significantly limits agent planning capability.
Extract videos from a YouTube channel with pagination support
Retrieve comments from a YouTube video
Extract comprehensive video metadata including description, upload date, like count
Extract transcript/captions from a YouTube video
Navigate to a video or channel page from search results
Search for specific videos within a YouTube channel
Search across all of YouTube and return structured results
Output schemas completely undocumented. Code shows tool definitions but actual response structures are invisible. LLMs cannot plan downstream actions without knowing what fields are returned.
Error handling strategy not visible. No error categorization (retryable vs user-fixable vs fatal), no recovery guidance. Tools calling external services (headless-youtube-captions) can fail silently or return cryptic errors. Agents have no way to diagnose or self-correct.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 19 | - | v1 |
Parameter descriptions lack implementation detail. Example: 'lang' says 'Language code for captions (e.g., "en", "es", "ko")' but does not state what happens if the language is unavailable, what the fallback is, or whether partial matches (e.g., 'en-US' vs 'en') work. LLMs must guess.
Pagination pattern unclear for list tools. get_channel_videos and search_channel_videos accept maxVideos/maxResults but do not document: (a) whether results are truncated if the limit is exceeded, (b) how to retrieve the next page, (c) total count or next_cursor. Without pagination metadata, agents cannot handle large result sets.
External library (@ts-ignore headless-youtube-captions) obscures actual return types. Code imports functions but their signatures and error modes are hidden. Cannot verify whether tool responses match schema definitions or if error paths are handled.
Parameter 'segment' in get_youtube_transcript lacks context. Description says 'Segment number to retrieve (1-based). Each segment is ~98k characters.' But does not explain: (a) how many segments exist for a given video, (b) what happens if segment exceeds the total, (c) how to discover segment count beforehand. Agent has no way to iterate.
No tool-annotation hints (readOnlyHint, destructiveHint, idempotentHint). All tools are correctly marked READ_ONLY in the rubric, but the MCP server does not declare this in the tool definition. Modern agents benefit from explicit safety metadata.
Caching implementation (transcriptCache, searchCache) is in-memory and per-server. No TTL management visible in tool definitions, so agents are unaware that results may be stale or cached. Leads to stale-data bugs if videos/channels change between calls.