Veo MCP demonstrates good definition quality overall. All 8 tools are explicitly registered with complete descriptions (194 - 400 chars, well above the 10 - 1024 baseline). Tool naming follows verb_noun convention with clear, unambiguous patterns (veo_list_*, veo_get_*, veo_text_to_video, veo_image_to_video). Parameters are well-typed with JSON Schema enums for model selection and aspect ratio. However, output schemas are not documented, responses are described in narrative form rather than formal schema declarations, which forces LLMs to infer structure. Error handling is absent: no recovery guidance, no categorization of errors, no validation error messages. Tools like veo_get_task return prose descriptions of task states ('processing', 'succeeded', 'failed') but the actual response format (fields, types, nested objects) is not specified. Most tools are lightweight wrappers around a Veo API, so they inherit the API's data model, but that is not documented in the MCP layer.
Get the 1080p high-resolution version of a generated video. By default, Veo generates videos at a lower resolution for faster processing. Use this tool to get the full 1080p version of a completed video. Use this when: - You need a higher resolution version for production use - The initial video generation is complete and you want to upscale - You need a clearer, more detailed video output Note: The video must be in 'succeeded' state before requesting 1080p version. Returns: Task ID and the 1080p video information including the new video URL.
Get guidance on writing effective prompts for Veo video generation. Shows how to structure prompts for best video generation results. Following these tips helps Veo understand your vision and generate more accurate and higher quality videos. Returns: Complete guide with prompt structure, examples, and tips.
Query the status and result of a video generation task. Use this to check if a generation is complete and retrieve the resulting video URLs and metadata. Use this when: - You want to check if a generation has completed - You need to retrieve video URLs from a previous generation - You want to get the full details of a generated video Task states: - 'processing': Generation is still in progress - 'succeeded': Generation finished successfully - 'failed': Generation failed (check error message) Returns: Task status and generated video information including URLs and state.
Query multiple video generation tasks at once. Efficiently check the status of multiple tasks in a single request. More efficient than calling veo_get_task multiple times. Use this when: - You have multiple pending generations to check - You want to get status of several videos at once - You're tracking a batch of generations Returns: Status and video information for all queried tasks.
Output schemas not documented. Tools return prose descriptions of results (e.g. 'Task status and generated video information including URLs and state') but formal JSON Schema for response payloads is missing. This forces LLMs to infer the structure of returned objects, increasing hallucination risk and preventing reliable chaining.
Error handling is absent. No recovery guidance, error categorization, or actionable error messages. Tools do not specify what errors are possible, whether they are retryable, or what the LLM should do next (e.g. if veo_text_to_video fails with 'invalid_prompt', should the agent retry with a different prompt or ask the user?).
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 75 | 2026-07-28+ | v2 |
Generate AI video from one or more reference images using Veo. This creates a video using your image(s) as reference frames. The video will animate from/between your provided images according to the prompt. Image modes: - 1 image: First-frame mode - the video starts from your image - 2-3 images: First-last frame mode - video interpolates between images - veo31-fast-ingredients model: Multi-image fusion - blends elements from all images Use this when: - You have a specific image you want to animate - You want consistent visual style from a reference - You need to create a video transition between two images For video generation from text only, use veo_text_to_video instead. Returns: Task ID and generated video information including URLs and state.
List all available Veo API actions and corresponding tools. Reference guide for what each action does and which tool to use. Helpful for understanding the full capabilities of the Veo MCP. Returns: Categorized list of all actions and their corresponding tools.
List all available Veo models and their capabilities. Shows all available model versions with their features, supported actions, and image input rules. Use this to understand which model to choose for your video generation. Model comparison: - veo3/veo3-fast: Improved quality, 1-3 images supported - veo31/veo31-fast: Latest models, 1-3 images supported - veo31-fast-ingredients: Multi-image fusion mode (ingredients2video action) Returns: Table of all models with their capabilities and image rules.
Generate AI video from a text prompt using Veo. This creates a video from scratch based on your text description. Veo will interpret your prompt and generate a matching video clip. Use this when: - You want to create a video from a text description - You don't have a reference image to use - You want maximum creative freedom for Veo For video generation starting from an image, use veo_image_to_video instead. Returns: Task ID and generated video information including URLs and state.
veo_list_models and veo_list_actions return unstructured text tables rather than structured JSON. This requires LLMs to parse plain text output, which is error-prone and wastes tokens. Example: veo_list_models returns a markdown table string, not a JSON array of model objects with typed fields.
veo_get_tasks_batch accepts both task_ids and trace_ids arrays, but the relationship between them and filtering by type, created_at_min, created_at_max is underdocumented. Are these filters AND or OR? If both task_ids and trace_ids are provided, which takes precedence?
veo_text_to_video and veo_image_to_video accept optional callback_url, but there is no documentation of what the callback payload looks like, when it is sent, or what the server should respond with. This makes it impossible for an LLM or user to set up correct callback handling.
No idempotency or dry-run guidance for write operations. veo_text_to_video and veo_image_to_video are write operations that incur costs. No mention of whether these are idempotent, what happens if an LLM retries with the same parameters, or whether a dry-run/estimate mode exists.
veo_image_to_video parameter image_urls description says 'Maximum 3 images' but does not explain what error occurs if more than 3 are provided, or whether the tool silently drops excess images.