High-performance video transcription MCP server using whisper.cpp for faster transcription
The server defines 4 tools with explicit schemas and descriptions. Tool naming follows verb-first conventions (transcribe_, check_, list_). Descriptions are substantive (194 - 200 chars on average, within the 10-1024 baseline). However, several parameters lack descriptions, error handling guidance is minimal, and output schemas are undocumented. The transcribe_video tool is well-parameterized with enums (model) and constraints; check_dependencies and list_supported_sites have empty schemas (no parameters documented); list_transcripts has optional output_dir but lacks output schema details. No error recovery guidance ('if transcription fails, try...') and no discussion of what fields the LLM can expect in results.
Check if all required dependencies (yt-dlp, ffmpeg, whisper models) are installed
List all video platforms supported by yt-dlp (1000+ sites including YouTube, Vimeo, TikTok, Twitter, Facebook, Instagram, educational platforms, and more)
List all available transcripts in the output directory, sorted by modification time (newest first)
Transcribe videos from 1000+ platforms (YouTube, Vimeo, TikTok, Twitter, etc.) or local video files using whisper.cpp (4-10x faster than Python whisper!). Downloads/extracts audio and generates transcript in TXT, JSON, and Markdown formats.
Output schemas are not documented for any tool. LLMs cannot plan downstream operations or extract fields without knowing return structure. E.g., does transcribe_video return {transcript: string, file_path: string, duration: number}? Does list_transcripts return [{name, path, modified_time, size}]?
check_dependencies and list_supported_sites have empty input schemas ('properties': {}). While technically valid for no-parameter tools, they lack descriptions of what the tool returns or when to call it. The description says 'Check if all required dependencies are installed' but does not document output format (list of status objects? boolean? per-component status?).
Error handling is absent. No guidance on what to do if transcription fails (out of disk? unsupported format? network error?). No error categorization (retryable vs. fatal). The tool description does not mention prerequisites ('requires yt-dlp, ffmpeg, and whisper models installed') or failure modes. Pattern recovery-guide demands 'User not found. Try search_users()', equivalent guidance is missing here.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 40 | - | v1 |
transcribe_video's 'output_dir' parameter description references get_default_output_dir() via format! macro, which produces runtime text like '~/.cache/video-transcriber-mcp'. This works but is slightly brittle, if the default changes, the description needs manual sync. A clearer description would be: 'Optional absolute or relative path. If omitted, transcripts are saved to ~/.cache/video-transcriber-mcp (configurable via $HOME).'
list_transcripts has an 'output_dir' parameter with description 'Optional output directory on the server. Leave unset.' This is confusing, if the agent should leave it unset, why is it a parameter at all? The note 'on the server' suggests security/sandboxing constraints but does not explain what happens if the agent passes a value. Either remove the parameter or document its valid range and behavior clearly.