Model Context Protocol server for Google AI services (VEO 3, Imagen 4, Gemini, Lyria 2)
The server provides 5 tools with explicit schemas and descriptions. However, there are systematic issues: (1) Tool naming lacks action verbs, 'veo_generate_video', 'imagen_generate_image', 'gemini_generate_text', 'lyria_generate_music' all use 'generate' but are module-prefixed in a way that obscures the action. 'check_operation_status' is the only clearly verb-led tool. (2) Descriptions are present but generic (34 - 72 chars), below the baseline average of 194 chars. They state WHAT but not WHEN to use or key prerequisites. (3) Schemas are complete and well-structured with proper JSON Schema types, enums, and defaults, this is a strength. (4) Parameter descriptions exist but are minimal (8 - 30 chars) vs. baseline of 72 chars. (5) No output schemas documented, the code shows responses are wrapped in {content: [{type: 'text', text: JSON.stringify(...)}]} but the actual return structure is not specified. (6) Error handling is minimal, the code shows try/catch blocks that re-throw errors without actionable guidance. (7) Security: outputStorageUri is a parameter on multiple tools, which could allow exfiltration if not properly validated; no evidence of permission checks. Overall, this is a competent but minimal implementation, tools work, but descriptions and error handling fall short of production quality.
Check the status of a long-running operation
Generate text using Google Gemini models
Generate photorealistic images using Google Imagen 4
Generate music using Google Lyria 2 (up to 60 seconds)
Generate videos using Google VEO 3 (5-8 seconds with audio)
Tool descriptions are below baseline (28 - 61 chars vs. average 194 chars). Descriptions state WHAT but not WHEN to use or how these tools relate to each other.
No output schemas documented. Tools return JSON wrapped in {content: [{type: 'text', text: '...'}]}, but the actual payload structure is not specified. LLMs cannot plan downstream operations without knowing return fields.
Tool naming uses product/module prefixes (veo_, imagen_, gemini_, lyria_) instead of clear resource types. E.g. 'veo_generate_video' obscures whether this generates videos, animations, or clips. 'video_generate' or 'generate_video' would be clearer.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
Parameter descriptions are minimal (8 - 30 chars). E.g., 'Text prompt for Gemini' does not explain constraints, format, or typical use cases. Baseline is 72 chars.
Error handling is minimal. Code shows try/catch blocks that re-throw errors without actionable guidance. E.g., if 'prompt' is missing, the error message is 'Prompt is required', but LLMs cannot recover from this without a suggestion (e.g., 'Did you omit the prompt? Try again with a text description.').
Multiple tools accept 'outputStorageUri' as a parameter, allowing clients to specify GCS bucket URIs. No validation or permission checks are visible. If a malicious agent passes an adversary's GCS bucket, the tool may write generated media there, causing data exfiltration.
No documented status values for 'check_operation_status'. LLMs cannot know what status strings to expect or how to branch on them. Should document: 'Returns {status: pending|succeeded|failed|cancelled, result?: {...}, error?: {...}}'.
Descriptions do not clarify which tools have side effects. E.g., 'Generate videos' reads as instantaneous, but the code likely returns an operation ID for polling. This invites LLMs to expect immediate results and miss the polling pattern.