Citation-ready YouTube transcripts, comments, and channel intelligence for AI agents over MCP
YouTube Research MCP demonstrates solid definition quality with 14 well-named tools, comprehensive parameter schemas, and detailed descriptions. All tools follow verb-noun naming conventions (search-, get-, research-, compare-). Schemas are properly typed with JSON Schema format including enums, descriptions, and defaults. However, output schemas are entirely undocumented, the server provides no structured definition of what each tool returns, forcing LLMs to infer response structure. Error handling guidance is minimal. Tool descriptions average ~150 chars and are adequate but could be more LLM-optimized per the 50-200 char baseline. Parameter descriptions are present on all params (100% coverage), a strong point. Security is handled correctly: API keys are server-side injected via environment variables, not exposed as parameters.
Compare transcripts from multiple YouTube videos to find common themes, overlapping content, and differences.
List all available caption tracks for a YouTube video, including language codes and auto-generated status.
[Requires YOUTUBE_API_KEY] Get information about a specific YouTube channel including subscriber count, video count, and view count.
Get enhanced transcript for one or multiple videos with formatting options, metadata, and optional video details.
Extract key moments/highlights from a YouTube video transcript using intelligent segmentation.
Get transcript segments with intelligent segmentation by time duration or equal length segments.
Get a summary of a YouTube video transcript with customizable summary length, keywords, and segment-level summaries.
Output schemas are completely undocumented. No tool defines what fields it returns, field types, or structure. LLMs cannot plan downstream calls or validate responses.
Error handling lacks recovery guidance. Tool descriptions do not explain what to do if the tool fails (e.g., 'if transcript unavailable, try get-captions-available first'). No error categorization (retryable vs user-fixable).
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
[Requires YOUTUBE_API_KEY] Get comments from a YouTube video with optional filtering by relevance or time, and include reply threads.
[Requires YOUTUBE_API_KEY] Get detailed information about a specific YouTube video including title, description, publication date, channel, view count, like count, and comment count.
Get the full transcript/captions for a specific YouTube video with optional language parameter. Returns complete transcript with timestamps.
Research a YouTube video transcript with search and filtering capabilities. Supports query-based research, time range filtering, context lines, and pagination with citation-ready output.
Research and compare multiple YouTube video transcripts. Applies the same search and filtering logic across multiple videos and returns comparative results.
Search within a single video's transcript for specific terms with context and pagination support.
[Requires YOUTUBE_API_KEY] Search for YouTube videos with advanced filtering options. Supports parameters: query (required), maxResults (1-50), channelId, order (date, rating, viewCount, relevance, title), type, videoDuration, publishedAfter, publishedBefore, videoCaption, videoDefinition, regionCode
Missing pagination limit enforcement on list-like results. search-videos, get-video-comments, search-transcript accept maxResults or similar but no documented cap. API could return very large payloads, bloating LLM context.
Descriptions for some tools are generic. E.g., 'Extract key moments/highlights from a YouTube video transcript using intelligent segmentation' does not explain WHEN to use this vs research-video or get-transcript-summary, or what constitutes a 'moment'.
Parameter dependency underdocumented. E.g., research-video accepts both startSeconds/endSeconds AND a query, unclear if query is applied before or after time filtering, or if both are required together.