MCP server that gives Claude Code the ability to watch and understand videos — extracts frames via ffmpeg and processes audio via multiple backends
The server defines 6 tools with mostly complete schemas and descriptions, but falls short of production quality. Tool naming follows verb_noun convention well (video_watch, video_info, video_analyze, video_detail). Descriptions are present and adequate (ranging 80-200 chars), meeting the minimum baseline of 10-1024 chars, but lack depth about error recovery and use-case guidance. Input schemas are visible and typed for all tools, a significant strength. However, several critical gaps emerge: (1) parameter descriptions lack constraint information (e.g., fps='auto' vs numeric not explained; resolution bounds missing), (2) output schemas are not documented anywhere in the visible source, (3) error handling is absent from descriptions, no guidance on what happens if video_path doesn't exist or ffmpeg fails, (4) the video_setup and video_configure tools expose API keys as parameters, violating secret-injection pattern. Overall, this is solid B-tier infrastructure work but misses the polish required for production agent use.
Run ffmpeg-based analysis filters on video to detect scene changes, black intervals, freeze frames, motion, blur, exposure, silence, and loudness
Update MCP server configuration settings for frame extraction, audio analysis, and indexing behavior
Extract frames and audio from a specific time segment of a video with optional detailed frame descriptions using an LLM
Get metadata information about a video file without extracting frames
Configure the MCP server with API keys and backend preferences for audio transcription
Extract and analyze frames from a video file with configurable fps, resolution, and optional audio transcription
API keys and credentials exposed as tool parameters (video_setup: openai_api_key, gemini_api_key). Credentials must never appear in tool params; they are logged and enter prompt history. Use server-side secret injection via environment variables or config files instead.
No output schemas documented. The visible source shows parameter schemas clearly, but tool descriptions and code do not specify what fields video_watch, video_info, video_analyze, and video_detail return. LLMs cannot plan downstream calls or extract data without knowing output structure.
Parameter constraints not documented in descriptions. Example: fps accepts 'auto' or a number, but the description does not explain the difference or when to use each. resolution parameter lacks min/max bounds (what if an agent passes 10000?). include_audio, describe_frames, and other booleans lack use-case context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2026-07-28+ | v2 |
Error handling and recovery guidance completely absent. No description tells the agent what to do if video_path doesn't exist, ffmpeg fails, audio transcription times out, or API rate limits are hit. Agents will be blocked with no actionable recovery path.
video_setup tool mixes configuration concerns (backend selection, API key injection, model selection) in one tool. This violates single-responsibility principle and forces the agent to call a setup tool before using the server, which is awkward for multi-agent scenarios.
video_configure tool has 11 parameters, all optional. The description does not clarify which combinations make sense (e.g., does audio_chunk_trigger_seconds apply only when using Gemini audio model?). Parameter dependencies are undocumented.