MCP server for video transcripts, downloading videos, and auto-generating subtitles from multiple platforms
transcript-mcp has 11 tools with mostly complete schemas and descriptions. Naming is action-verb-based and clear (get-, list-, download-, generate-, transcribe-). Descriptions are well-written and exceed 100 chars on average, explaining what the tool does and when to use it. However, there are significant gaps in error handling, output schema documentation, parameter validation guidance, and security considerations. Parameter descriptions are generally good but lack detail on ranges, formats, and constraints. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite multiple tools having side effects. Output schemas are not documented for any tool, making it difficult for LLMs to plan downstream calls. Error handling lacks recovery guidance and categorization.
Download a video from any supported platform (YouTube, Vimeo, etc.) to local storage. Returns the file path of the downloaded video.
Generate subtitles for a local video file using AI speech-to-text (OpenAI Whisper or local whisper). Creates an SRT or VTT file alongside the video.
Retrieve the transcript of a video from supported platforms (YouTube, Bilibili, Vimeo, etc.). Accepts various URL formats and returns the full transcript with timestamps.
List all downloaded video files in the storage directory or a specified directory.
List all available transcript languages for a video from any supported platform.
Transcribes audio via Whisper. Preferred: audio_url (most token-efficient; server fetches bytes). audio_base64 is for small clips only (<= ~60KB raw per call). audio_path only works when the MCP host shares a filesystem with the caller (often false on Claude.ai / Claude Code). For larger payloads in sandboxed environments, use transcribe_upload_start / transcribe_upload_append / transcribe_upload_finalize. Server re-encodes to Opus 16kHz mono 16kbps before Whisper unless skip_compression=true. Long audio (>5min) or async=true returns a job_id; poll transcribe_get_job.
No output schema documented for any tool. LLMs cannot plan downstream calls or extract chaining IDs without explicit documentation of response structure.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite multiple state-modifying tools. download-video, generate-subtitles, transcribe-* tools all have side effects but lack explicit annotations. This makes it difficult for agents to reason about safety and retry logic.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
Cancel a pending or processing async transcription job.
Poll the status of an async transcription job. Returns status, result when completed, or error if failed.
Append a chunk of base64-encoded audio data to an active upload session.
Finalize a chunked audio upload and trigger transcription. Returns job_id for async processing or the transcription result.
Start a chunked audio upload session for large files. Returns a session token to use with subsequent append/finalize calls.
Error handling lacks recovery guidance. No error response examples or categorization (retryable, user-fixable, fatal). Error messages would benefit from actionable next steps for LLMs.
Parameter validation constraints not fully specified in descriptions. 'format', 'quality', and 'engine' enums are declared in schema but lack constraint explanation in descriptions. Constraints like file path format, URL allowlist (TRANSCRIPT_MCP_URL_ALLOWLIST), and size limits are undocumented.
Security consideration: transcribe-audio accepts audio_url parameter, which requires TRANSCRIPT_MCP_URL_ALLOWLIST environment variable, but this constraint is NOT mentioned in the parameter description. LLMs won't know which URLs are permitted.
Multi-parameter dependencies undocumented. transcribe-audio has mutually exclusive audio input modes (audio_path, audio_base64, audio_resource_uri, audio_url) but descriptions don't state this. LLMs may pass multiple conflicting parameters.
No pagination or result limits documented. list-downloads and list-transcript-languages could return large unbounded results. No mention of pagination support or max result count guidance.
Destructive operations lack confirmation or dry-run support. download-video and generate-subtitles write to disk, transcribe-cancel-job is irreversible, but no confirmation pattern is documented.