MCP server for media generation and analysis using Google Gemini APIs, supporting image generation, video creation, text-to-speech, and image analysis
The server defines 4 tools with explicit schemas and descriptions in src/main.py using FastMCP decorators. Tool names follow verb_noun convention (create_visualization, text_to_speech, analyze_image, create_video). However, descriptions are generic and lack actionable context for LLM selection. Parameter descriptions exist but are minimal (e.g., 'Aspect ratio of the generated image' without constraints or examples). No output schemas are documented, callers cannot see what fields to expect. Error handling is absent from the code; no indication of how failures are communicated or what recovery paths exist. Security: API keys are configured via environment variables (good), but the middleware enforces a single x-api-key header without granular permission scoping. Tool definitions are directly visible in source (not inferred), but the implementation delegates to submodules (src.tools.image, audio, video) whose actual logic is not provided for review.
Analyze an image using Gemini vision model and return detailed description or analysis
Generate a video from a text prompt using Google Veo video generation models with optional image conditioning
Generate an image based on a text prompt using Gemini image generation model
Convert text to speech audio using Gemini TTS model with configurable voice
Output schemas completely undocumented. No definition of what fields are returned, their types, or structure. LLMs cannot plan downstream calls or extract relevant data.
Tool descriptions lack actionable context for LLM selection. 'Generate an image based on a text prompt' is generic, does not explain when to prefer this over other generative tools, what quality to expect, or prerequisites. Baseline: 50-200 chars with WHEN/WHY/WHAT structure.
Parameter descriptions lack constraints and formats. 'aspect_ratio' accepts values like '1:1', '16:9', '9:16' but the description does not enumerate valid options or format requirements. Should use enum constraint or explicit pattern statement.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
No error handling guidance in tool code or documentation. How do tools report failures? Are errors retryable, user-fixable, or fatal? Does the tool return structured error objects or exceptions? LLMs have no recovery path.
analyze_image accepts BOTH image_url and image_base64 as optional parameters with unclear semantics. Description does not state which takes precedence, whether both are mutually exclusive, or what happens if both are provided. This invites ambiguous LLM calls.
create_video similarly accepts image_url and image_base64 without documenting mutual exclusivity or precedence.
No tool annotations visible (readOnlyHint, destructiveHint, idempotentHint). create_visualization and create_video are write operations and should carry destructiveHint:true. analyze_image is read-only and should carry readOnlyHint:true. Current MCP spec (2026-07-28) supports these for LLM planning.