MCP server for extracting key frames from videos and returning them as base64 images for Claude analysis
The server defines 2 tools with explicit Zod schemas, good descriptions, and proper tool annotations. Both tools have verb-noun naming (analyze_video, get_video_info), structured input schemas with type constraints and descriptions, and tool annotations (readOnlyHint, destructiveHint, idempotentHint). Descriptions are detailed (194-250 chars) and explain WHAT each tool does, WHEN to use it, and what data it returns. Error handling returns actionable messages. Main gaps: output schemas are not formally documented (only examples in description text); parameter dependencies (mode vs sceneThreshold) could be more explicitly documented; no pagination/limits guidance for batch operations (though tool set is small enough this is minor).
Extract key frames from a short video file and return them as images for visual analysis. Supports .mp4, .webm, and .gif files up to 60 seconds. Two extraction modes: - "smart" (default): Detects scene changes to capture meaningful transitions. Falls back to interval mode if no scene changes are detected. - "interval": Extracts frames at equal time intervals. Returns a metadata text block followed by JPEG image blocks with timestamps. Use this to review UI flows, state transitions, animations, and visual changes in short screen recordings. Args: - filePath (string): Absolute path to the video file - mode ("smart" | "interval"): Extraction mode (default: "smart") - maxFrames (number 1-20): Maximum frames to extract (default: 10) - sceneThreshold (number 0-1): Scene change sensitivity for smart mode (default: 0.3)
Get metadata about a video file without extracting frames. Returns JSON with duration, resolution, fps, codec, format, and file size. Args: - filePath (string): Absolute path to the video file
Output schema not formally documented. Both tools return structured JSON + image blocks, but the return type structure is described only in natural language, not as a formal JSON Schema. LLMs cannot reliably extract or validate output fields without explicit schema.
Parameter dependency (sceneThreshold only applies when mode='smart') is mentioned in description but not formally declared. LLMs may pass sceneThreshold with mode='interval', wasting context on a no-op parameter.
No guidance on maxFrames default interaction with mode. If mode='smart' and no scene changes detected, does maxFrames still limit the fallback interval extraction? Behavior is implemented but not documented in description.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 79 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |