Extracts available subtitles or generates them if they don't exist
Single tool with reasonable naming and schema structure, but significant gaps in descriptions and error handling. The tool name 'get_youtube_transcript' follows verb_noun convention (positive), and the input schema is present with proper JSON Schema structure. However, the tool description, while present, is generic and doesn't adequately explain when/why to use it relative to alternatives or prerequisites. Parameter descriptions exist but lack detail about format/constraints. Most critically, error handling is absent from the visible code, no guidance for LLMs when transcript lookup fails, when Whisper transcription errors occur, or how to retry. Output schema is not documented in the visible tool definition.
Search for a YouTube video and get its transcript. If no official transcript exists, generates one using Whisper.
Tool description is generic and lacks context for LLM selection. Current: 'Search for a YouTube video and get its transcript. If no official transcript exists, generates one using Whisper.' Does not explain prerequisites (internet connectivity, Whisper availability), failure modes, or when to call this vs other tools.
Output schema not documented. The tool returns a dict but the structure is not visible in the tool definition. LLMs cannot plan downstream actions or extract required fields (transcript text, video metadata, duration, etc.) without knowing the response shape.
Parameter 'query' description lacks format/constraint guidance. Current: 'Search query for the YouTube video (e.g., 'What is an API? by MuleSoft')'. The example is problematic per the rubric (LLMs reuse examples literally). Should specify: minimum/maximum length, whether creator name is required, whether to include timestamps, handling of ambiguous/non-existent videos.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 56 | - | v1 |
No error handling or recovery guidance. Code shows multiple failure paths (video not found, transcript disabled, Whisper failures, download failures, model file missing) but tool definition contains no error descriptions or actionable next steps for LLMs. When failures occur, agents have no guidance.
'force_whisper' parameter description is minimal ('Force the use of Whisper for transcription, even if an official transcript is available'). Does not explain why an LLM would want to force Whisper (e.g., faster, different language support, accuracy preference), performance implications, or fallback behavior if Whisper fails.