The server defines 8 tools with reasonable naming conventions (verb_noun pattern), mostly good descriptions, and complete input schemas. However, there are gaps in output schema documentation, some parameter descriptions lack specificity, and error handling guidance is minimal. Tool names follow the pattern well (veo_list_*, veo_get_*, veo_text_to_*, veo_image_to_*), and descriptions are generally adequate (50-250 chars). The main weakness is the absence of documented output schemas and limited error recovery guidance for non-idempotent operations like video generation.
Get the 1080p high-resolution version of a generated video. By default, Veo generates videos at a lower resolution for faster processing. Use this tool to get the full 1080p version of a completed video. Use this when: - You need a higher resolution version for production use - The initial video generation is complete and you want to upscale - You need a clearer, more detailed video output Note: The video must be in 'succeeded' state before requesting 1080p version.
Get guidance on writing effective prompts for Veo video generation. Shows how to structure prompts for best video generation results. Following these tips helps Veo understand your vision and generate more accurate and higher quality videos.
Query the status and result of a video generation task. Use this to check if a generation is complete and retrieve the resulting video URLs and metadata. Use this when: - You want to check if a generation has completed - You need to retrieve video URLs from a previous generation - You want to get the full details of a generated video Task states: - 'processing': Generation is still in progress - 'succeeded': Generation finished successfully - 'failed': Generation failed (check error message)
Query multiple video generation tasks at once. Efficiently check the status of multiple tasks in a single request. More efficient than calling veo_get_task multiple times. Use this when: - You have multiple pending generations to check - You want to get status of several videos at once - You're tracking a batch of generations
Output schemas not documented. Tools like veo_text_to_video, veo_image_to_video, and veo_get_task return complex structured data (video URLs, task status, metadata) but the response schema is not declared. LLMs cannot plan downstream tool calls or extract specific fields without knowing what the response contains.
Error handling and recovery guidance missing. The tools do not document what errors can occur (e.g., invalid task_id, generation failure, API rate limit) or what the LLM should do next (retry, ask user, call a different tool). Non-idempotent write operations (veo_text_to_video, veo_image_to_video) especially need recovery paths.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 63 | - | v1 |
Generate AI video from one or more reference images using Veo. This creates a video using your image(s) as reference frames. The video will animate from/between your provided images according to the prompt. Image modes: - 1 image: First-frame mode - the video starts from your image - 2-3 images: First-last frame mode - video interpolates between images - veo31-fast-ingredients model: Multi-image fusion - blends elements from all images Use this when: - You have a specific image you want to animate - You want consistent visual style from a reference - You need to create a video transition between two images For video generation from text only, use veo_text_to_video instead.
List all available Veo API actions and corresponding tools. Reference guide for what each action does and which tool to use. Helpful for understanding the full capabilities of the Veo MCP.
List all available Veo models and their capabilities. Shows all available model versions with their features, supported actions, and image input rules. Use this to understand which model to choose for your video generation. Model comparison: - veo3/veo3-fast: Improved quality, 1-3 images supported - veo31/veo31-fast: Latest models, 1-3 images supported - veo31-fast-ingredients: Multi-image fusion mode (ingredients2video action)
Generate AI video from a text prompt using Veo. This creates a video from scratch based on your text description. Veo will interpret your prompt and generate a matching video clip. Use this when: - You want to create a video from a text description - You don't have a reference image to use - You want maximum creative freedom for Veo For video generation starting from an image, use veo_image_to_video instead.
Parameter descriptions lack specificity on constraints and expected formats. For example, veo_text_to_video's 'prompt' parameter says 'Be descriptive about scene, subject, action, camera movement, lighting, and style' but does not specify: minimum length, maximum length, unsupported characters, or what happens if a prompt is too vague. veo_get_task's 'trace_id' is marked optional but its purpose and format are unclear.
veo_get_tasks_batch parameter relationships not documented. The tool accepts optional task_ids, trace_ids, offset, limit, type, and timestamp filters, but does not explain: which filters are mutually exclusive, what is the default behavior if none are provided, or how pagination interacts with filters. This forces LLMs to guess.
No pagination metadata in tool descriptions. veo_get_tasks_batch accepts 'limit' and 'offset' parameters but does not document: What is the maximum recommended batch size? What is the default limit? Does the response include a total count or next_cursor for proper pagination? This can cause LLMs to request too many items or fail to paginate correctly.
Tool descriptions include example values (e.g., 'Examples: A white ceramic coffee mug...') which LLMs may reuse literally in real calls. Best practice is to rely on enums and format constraints instead of examples in prose.