Extract and search YouTube video transcripts
The server has well-structured tools with clear naming conventions (all verb_noun pattern) and comprehensive parameter schemas. All four tools follow consistent patterns with good descriptions and type definitions. However, there are significant gaps in output schema documentation, error handling guidance, and response structure optimization. The tools accept both URLs and video IDs (good UX), but responses are returned as plain strings rather than structured JSON, limiting LLM composability. Tool descriptions are adequate (100-150 chars each) but could be more prescriptive about when to use each variant.
Get transcripts for multiple YouTube videos in a single request (max 10 videos).
Get the full transcript of a YouTube video in the specified language and format.
Get a transcript organized into time-based chunks for easier analysis of long videos.
Search for keywords or phrases in a YouTube video transcript and return matching segments with timestamps.
All tools return unstructured plain-text strings instead of structured JSON objects. LLMs cannot reliably parse markdown-formatted responses for downstream chaining.
No documented output schema. Tool descriptions state what they do, but LLMs have no specification of returned fields, types, or structure. This violates the critical pattern that tools returning lists must document schema.
Error handling returns plain-text error strings (e.g. 'Error: Invalid YouTube URL...') with no guidance on recovery. LLMs cannot determine if error is retryable, user-fixable, or fatal.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 60 | - | v1 |
batch_transcripts lacks input validation and upper bound enforcement. Description states 'max 10 videos' but code does not enforce or error if exceeded. LLMs may pass lists longer than 10.
search_transcript does not document what context_segments value should be for best results. Parameter range is constrained (0-10) but guidance on typical use is missing.
get_transcript_summary does not specify expected chunk count, memory footprint, or maximum response size. Long videos chunked into 1-minute segments could return thousands of entries, no pagination or limit.
Rate limiting is global (30 calls/min) with no per-tool or per-resource guidance. No rate-limit headers returned to client. LLM has no way to know when it's approaching limit.