Returns YouTube video data for use by the host AI: (1) caption transcript text, (2) title / channel / description from the watch page. Does not call OpenAI or any LLM — you write the article or summary yourself from this content.
Server defines 2 tools with reasonable naming (verb_noun pattern) and adequate descriptions. Both tools have input schemas with type definitions and parameter descriptions. However, output schemas are not documented, error handling lacks recovery guidance, and descriptions could be more detailed per LLM-optimized standards. The server follows a clean, single-responsibility pattern but falls short of production-grade polish.
Fetch caption transcript text for a video (plus title, channel, and description from the watch page).
Fetch title, channel, and description text from the video watch page (no captions API). Use when you only need the uploader description, or when transcripts are unavailable.
Output schemas not documented. Tool descriptions do not specify the structure of returned JSON objects (e.g., keys like 'video_id', 'video_title', 'transcript', 'error'). LLMs cannot plan downstream operations or extract specific fields without knowing the response structure.
Error messages lack recovery guidance. Errors return JSON objects with only an 'error' string (e.g., 'YouTube rate limit or IP block; try again later'). However, there is no structured guidance on whether the error is retryable, user-fixable, or fatal. Per pattern:recovery-guide, error responses should tell the LLM what to do next.
Parameter descriptions are minimal. 'target_lang' is described as 'Optional ISO language code to prefer for captions (e.g. en, de)' but does not state format constraints (2-character code, case sensitivity, or list of supported languages). Per pattern:constrained-input, format expectations should be explicit.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | <=2025-11-25 | v2 |
Flexible video_id parameter. Both tools accept 'Full watch URL or 11-character video ID' but do not document the URL format (full watch URL like 'https://www.youtube.com/watch?v=...' or short form?). This ambiguity forces LLMs to guess or retry on format errors.
No pagination or limits. Both tools return full results without documented size limits or pagination tokens. If a transcript is extremely long (120,000+ chars per MAX_TRANSCRIPT_CHARS in app.py), the response could exhaust the agent's context window. Per pattern:paginated-result, large results should support pagination.