MCP server for video analysis — extracts transcripts, key frames, OCR text, and metadata from video URLs and local files. Supports Loom, YouTube and other yt-dlp platforms, direct video URLs, and local video files.
This is a well-structured video analysis MCP server with clear naming conventions, comprehensive tool coverage, and detailed parameter schemas. All 8 tools follow verb_noun naming (analyze_, get_) and have non-trivial descriptions (100 - 500 chars). Input schemas are complete with proper type definitions and parameter descriptions for all tools. However, there are notable gaps: (1) output schemas are not documented for any tool, LLMs cannot predict what fields to expect from results; (2) error handling guidance is absent, tools lack descriptions of failure modes and recovery steps; (3) some parameters lack granular constraint documentation (e.g., maxWidth has no min/max bounds stated in descriptions); (4) no tool annotations (readOnlyHint, idempotentHint) are present despite all tools being READ_ONLY. The caching behavior (10-min default + forceRefresh flag) is documented in analyze_video but not explained as an idempotent pattern that LLMs should understand. Overall, naming and basic schema structure are solid, but output documentation and error guidance are the primary weaknesses.
Deep-dive on a time range. Combines burst frames + filtered transcript + OCR + mini-timeline. Use when the user asks about a specific part of the video.
Full analysis. Use by default when a video URL appears. Returns transcript + frames + metadata + OCR + timeline. detail="brief" → fast, metadata + truncated transcript, no video download; detail="standard" → default, scene-change frames + full transcript + OCR; detail="detailed" → dense 1fps sampling, more frames, thorough OCR. fields parameter allows returning only specific fields. Cached for 10min — use forceRefresh=true to re-analyze.
Batch version of analyze_video. Use when given a list of sources (e.g. a folder of local files). Runs them with a concurrency limit and returns one structured result per source (counts + warnings, or a per-item error). Pair with MCP_WRITE_SIDECARS=1 for resumable bulk processing.
Single frame at a timestamp. Use when the transcript reveals an interesting moment and you want to see it.
N frames in a narrow time range. Use for motion, animations, fast UI changes.
Frames only (scene-change or dense=true for 1fps). Use when you need visuals without transcript.
No output schemas documented for any tool. LLMs cannot predict result structure, fields, or types returned. This breaks downstream tool chaining and forces agents to guess what fields are available.
Missing error handling guidance. Descriptions do not explain failure modes (e.g. 'URL not supported', 'video too long', 'OCR failed') or recovery steps. Agents will not know whether to retry, ask user, or abandon.
No tool annotations present. All tools are READ_ONLY and idempotent (caching + forceRefresh), but lack readOnlyHint and idempotentHint to signal this to LLMs. This forces conservative agent planning.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 57 | 2026-07-28+ | v2 |
Metadata + comments + chapters. No video download needed.
Transcript only. Faster than analyze_video when you only need what was said. Whisper fallback for videos without native transcripts.
Numeric parameters (maxWidth, maxFrames, frameCount, concurrency) lack min/max bounds in descriptions. E.g. maxWidth: 'Maximum frame width in pixels' omits valid range. LLMs may pass invalid values like 10000 or negative numbers.
analyze_videos batch tool lacks per-item error reporting. Description says 'counts + warnings, or a per-item error' but does not clarify whether partial success is possible or how to distinguish which of N sources succeeded/failed. LLMs cannot recover from selective failures.
get_transcript and other tools reference 'Whisper fallback' and 'native transcripts' but do not document which video platforms provide native vs. fallback transcripts. LLMs cannot predict result quality or understand why some URLs may take longer.