mcp-subtitle has 2 tools with reasonable naming but critical gaps in schema documentation, parameter descriptions, and error handling. Both tools have input parameters but the schema is not explicitly visible in the source code, only inferred from parameter names. Descriptions are present but brief (70-80 chars). The 'cancel_subtitles' tool has questionable semantics (marked READ_ONLY but suggests state mutation). No output schema documented. Error handling is entirely absent from the tool definitions. The implementation includes complex logic (Whisper transcription, yt-dlp integration, caching) but none of this is surfaced in tool descriptions to guide LLM selection or understand prerequisites.
Cancel an ongoing subtitle extraction task for a given URL and language.
Extract subtitles from a YouTube video, using existing subtitles if available, otherwise transcribe the audio using Whisper.
Input schemas not explicitly visible in source code. Only parameter names and descriptions are provided; no JSON Schema type definitions, minLength, maxLength, enums, or patterns are declared.
Output schema completely undocumented. get_subtitles returns subtitle content, but structure is undefined, is it a string, JSON object with metadata, array of segments? Agents cannot plan downstream operations without knowing the response shape.
Tool descriptions are too brief (45 - 70 chars). 'Extract subtitles from a YouTube video...' does not explain WHEN to use this tool, what makes it different from other subtitle extractors, or that it triggers heavy ML workload (Whisper transcription can take minutes). Missing prerequisites and operational expectations.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
'cancel_subtitles' marked as READ_ONLY risk, but its semantics involve stopping/cancelling a running task, this is a state mutation, not a read. Risk annotation is incorrect and misleading.
No error handling guidance. Code includes complex failure modes (audio download failed, Whisper transcription error, subtitle parsing failed) but tool definitions provide no recovery hints. Agents will not know what to do if get_subtitles fails.
Language parameter accepts free-form strings ('auto', 'en', 'zh', etc.) but no enum constraint is declared. LLMs cannot infer valid language codes and may hallucinate invalid values like 'english' or 'mandarin' instead of 'en' or 'zh'.
Parameter description for 'language' mentions 'Defaults to auto' but does not specify what languages are supported, how they map to Whisper language codes, or that invalid languages will silently fall back to 'auto'.
No documentation of operational constraints: get_subtitles can take minutes to transcribe long videos (Whisper is CPU-bound). No timeout declared, no indication of cost/latency expectations, no guidance on when cancellation is appropriate.
Tool definitions do not explain the caching behavior. If a subtitle is already cached, does get_subtitles return instantly or re-extract? This determines whether the LLM can safely call it multiple times without side effects.