A multimodal video processing and search MCP server that processes videos, extracts clips based on queries, searches by speech/captions/images, and answers questions about video content.
The server provides 4 tools with basic schema definitions and descriptions, but several critical gaps prevent higher scoring. All tools have input schemas with type information and parameter descriptions (good baseline), but descriptions lack depth and context about when to use each tool vs alternatives. Parameter descriptions are present but minimal (averaging ~40-50 chars, below the 72-char baseline). No output schemas are documented anywhere in the codebase, LLMs cannot predict what these tools return. Tool naming follows verb_noun convention (positive), but there is no distinction between the two query-based tools (get_video_clip_from_user_query vs get_video_clip_from_image) in terms of when to use one vs the other. Error handling, idempotency guidance, and permission scopes are entirely absent. The tools appear functional for their domain (video processing) but fall short of production-grade quality that would confidently guide an LLM's selection and use.
Use this tool to get an answer to a question about the video.
Use this tool to get a video clip from a video file based on a user image.
Use this tool to get a video clip from a video file based on a user query or question.
Process a video file and prepare it for searching.
No output schemas documented. LLMs cannot predict what these tools return (field names, types, structure), forcing them to reason blindly about downstream chaining. Required for proper tool composition.
Ambiguous tool selection between get_video_clip_from_user_query and get_video_clip_from_image. Descriptions do not explain the distinction or when to use one vs the other. LLMs will struggle to choose correctly.
process_video description does not state whether it modifies state or what it returns. Agents need to know: is this idempotent? What does it prepare? Can it be retried safely?
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Descriptions are brief (30-50 chars) and lack context about prerequisites, expected formats, and when to call the tool. Below the 194-char baseline for A+ tools.
No error handling guidance. If a video file is not found or processing fails, LLMs have no recovery path. Missing recovery guides for retryable vs user-fixable vs fatal errors.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. Clients and LLMs cannot distinguish safe read-only tools from write operations.