Veo 3 Video Generation MCP server built with FastMCP. Provides Google Veo video generation capabilities through MCP tools, supporting text-to-video and image-to-video generation with multiple models and aspect ratios.
This server has 5 tools with schemas visible in the source code. Naming follows verb_noun conventions (veo_generate_video, veo_check_generation, etc.) which is appropriate. However, there are significant quality gaps: (1) Output schemas are not documented anywhere in the visible code, tool descriptions mention what is returned but no structured schema is provided for LLM planning; (2) Parameter descriptions are present but inconsistent in quality and detail; (3) No error handling guidance or recovery paths documented; (4) Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint) which would clarify state mutation; (5) Tool descriptions are verbose (200-350+ chars) when best practice is 50-200 chars. The tools do have input schemas with types and descriptions visible, which prevents a lower score, but the lack of output schema documentation and error handling guidance significantly limits production readiness.
Check the status of a video generation subprocess. This monitors the subprocess that is handling the video generation, not the Gemini API directly. The subprocess handles all API polling. If the generation has an operation_id, you can also use veo_check_operation to directly query Google's API for the operation status.
Check the status of a video generation operation directly via Google's API. Query the operation status directly from Google's Gemini API servers for real-time status information.
Download completed videos from a generation session. Downloads all generated videos from the specified generation session to the local filesystem. Videos are stored on Google's servers for 2 days, so download within this period to save permanently.
Generate videos using Google's Veo models. Supports text-to-video, image-to-video, and both combined. This tool starts a background video generation process that initiates video generation with specified parameters and monitors progress in the background. Returns session_id for monitoring progress with veo_check_generation.
List all video generation operations with their current status. Shows all active and completed generation operations managed by the local generation manager.
Output schemas not documented. Tool descriptions describe what is returned (e.g., 'Returns session_id') but no structured JSON Schema output is provided. LLMs need documented return types to plan downstream tool calls and extract field names correctly.
No error handling or recovery guidance. Tools describe what they do but do not explain what errors may occur, how to distinguish retryable from fatal failures, or what the LLM should do next on failure. Missing 'Try search_users() first' style dependency hints.
Tool descriptions are 200-350+ chars, exceeding the best-practice 50-200 char range. Verbose descriptions waste tokens and bury key selection criteria. E.g., veo_generate_video description is ~280 chars; trim to essential 'Generate videos from text/image using Google Veo models. Returns session_id for monitoring progress.'
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Missing tool annotations (destructiveHint, readOnlyHint, idempotentHint). veo_download_video and veo_generate_video are destructive or have side effects (file I/O, API calls with billing impact), but the schema does not declare this. Annotations help agents reason about safety.
veo_list_operations returns all active/completed operations with no documented limit or pagination. If hundreds of operations exist, returning all risks context window exhaustion. Add 'Returns at most 100 operations; use cursor-based pagination for older results.' and implement limit/offset parameters.
Inconsistent parameter descriptions. 'number_of_videos' has detailed range (1-4) but 'duration_seconds' says '2-15' without noting the veo-3.0-fast exclusion clearly in parameter description. 'resolution' says 'if supported by model' vaguely; which models support what?