MCP server for video analysis and processing, exposing tools for video processing, multimodal search (text, image, caption, visual), clip generation, and video metadata retrieval
MOSAIC MCP server provides 9 video analysis tools with reasonable structure but significant quality gaps. Tool naming follows verb_noun convention well (process_video, search_text, generate_clips). Descriptions exist for all tools but are inconsistent in quality and length, some are detailed (process_video: 163 chars) while others are shorter but adequate. Input schemas are visible and include type definitions, but lack comprehensive validation constraints. Output schemas are not documented in the provided source, which is a critical omission per the rubric. Error handling is not evident from the source code. Security considerations for file operations (WRITE risk on process_video and generate_clips) are not addressed with validation or permission gates. The server is HTTP-based (good for protocol readiness), but the definition quality is held back by missing output schema documentation and lack of error recovery guidance.
Extract video clips from search results. Creates clip files for each hit with timing information.
Get information about a processed video. Returns frame count, index paths, and collection names.
List all available processed videos in the system.
Process and index a video file. Extracts keyframes, generates captions, transcribes audio, and stores embeddings in vector databases (FAISS + ChromaDB).
Search for relevant frames using AI-generated captions. Good for finding specific objects, actions, or scenes described in text.
Search video frames using image similarity (CLIP embeddings). Returns similar frames with clip timing parameters.
Output schemas are not documented. The rubric requires documented return types for 100% of A-grade tools. Without knowing what fields search_text, search_image, generate_clips, and others return, LLMs cannot plan downstream tool calls or extract chaining IDs (video_id, timestamps, frame counts).
Missing error handling and recovery guidance. Tools like process_video (WRITE) and generate_clips (WRITE) have destructive consequences but provide no error classification (retryable vs fatal), no recovery hints (e.g., 'Check disk space if write fails'), and no confirmation/dry-run mechanism per pattern:confirmation-request.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Search video transcript using text query. Returns matching transcript segments with timestamps.
Search video frames using CLIP visual-semantic embeddings. This uses the same CLIP model to encode text and find visually similar frames. More powerful than caption-based search for visual content.
Summarize the video content using its transcript.
Destructive operations (process_video, generate_clips) lack security gates and permission declarations. No evidence of input sanitization against path traversal (video_path, output_dir parameters accept raw user strings). No permission model (read:video, write:storage).
Parameter descriptions lack actionable constraints. fps defaults to 30.0 with no bounds documented (should state range, e.g., '1.0 - 120.0'). k defaults to 5 with no max documented; searching 1000 results could blow context.
list_videos has an empty input schema ({}), making it unclear what optional params are accepted (pagination, filter, sort). No limit, offset, or filter parameters documented.
No pagination support documented for search tools (search_text, search_image, search_caption, search_visual). They accept k parameter but no offset/limit or cursor for retrieving beyond top-k results. Large result sets will blow context window.
get_video_info and summarize_video descriptions are generic and lack actionable details. 'Get information about a processed video' (62 chars) does not explain what 'index paths' and 'collection names' mean to the LLM or when to use this vs list_videos.
generate_clips 'hits' parameter description is vague: 'List of search hits with timing information (from search tools)'. Does not specify the exact structure (array of objects with fields?), required fields, or data types. LLM must guess the format.