MCP server for Google GenMedia APIs (Gemini Image, Veo, Chirp, Lyria)
This server has significant gaps in definition quality. While tool names follow verb_noun conventions and descriptions are present, descriptions are often too long (violating the 10-1024 char guideline), lack clear LLM-friendly structure (WHAT/WHEN/prerequisites), and parameter descriptions are minimal or absent. Input schemas are visible but lack proper type constraints and enums where appropriate. Error handling is present but generic (returns dict with 'error' key without recovery guidance). Output schemas are undocumented, the server returns generic dicts without declaring structure to the agent. No tool annotations (readOnlyHint/destructiveHint). The combine_audio_video tool is the only one marked as WRITE risk, but this metadata is not surfaced in the MCP tool definition itself. Japanese-language descriptions add ambiguity for English-speaking LLMs. The server_info tool is self-referential and relatively useless for planning. Overall, the tools are functional but poorly optimized for LLM decision-making.
動画と音声を ffmpeg で合成する。 前提: ffmpeg がシステムにインストールされている必要があります。 インストール: macOS: brew install ffmpeg、Ubuntu: apt install ffmpeg
テキストから画像を生成する(Gemini)。 reference_image を指定すると、その画像を入力にした編集(参照画像編集)を行う。
Lyria モデルでテキストから音楽を生成する。 Lyria 3 Pro(デフォルト): - 最大約 184 秒の楽曲を生成(プロンプトで秒数を指定可能) - ボーカル・歌詞対応(プロンプトに歌詞を含めると歌付き楽曲を生成) - [Verse], [Chorus], [Bridge] などのセクションタグで構成を制御可能 - BPM やテンポはプロンプト内で自然言語で指定(例: "120 BPM") - MP3 形式で出力 - インストのみにしたい場合は "Instrumental only, no vocals" と指定 Lyria 3 Clip: - 30 秒のクリップを生成 - その他は Lyria 3 Pro と同じ機能 Lyria 2: - 30 秒固定のインストゥルメンタル音楽を生成 - negative_prompt と seed パラメータに対応 - WAV 形式で出力 注意: - Lyria 2 は Vertex AI または OAuth 認証方式でのみ利用可能です - Lyria 3 は API Key 方式でも利用可能です
Chirp 3 HD でテキストを音声に変換する。 注意: この機能は Vertex AI または OAuth 認証方式でのみ利用可能です。 API Key 方式では利用できません。
Veo モデルでテキストから動画を生成する。 注意: 動画生成には数分かかる場合があります(ポーリングで完了を待機)。
Descriptions written in Japanese and overly technical. LLMs are primarily English-trained; non-English descriptions create friction. Additionally, descriptions exceed 200 chars for most tools (generate_music is 500+ chars), violating the 10-1024 guideline and burying key selection criteria. Example: 'generate_music' description is a wall of text about model variants rather than a crisp 'Generate music from text prompts. Supports multiple Lyria models with varying duration and feature support.'
Parameter descriptions are minimal or missing critical constraints. For example, 'model' parameters accept string but lack enum values or valid options. LLMs will hallucinate model names. 'aspect_ratio' has no enum (should enumerate valid values like '16:9', '9:16', '1:1'). 'duration_seconds' lacks min/max bounds, agents may pass absurd values like 999999 or 0. 'number_of_videos' unbounded.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Veo モデルで画像から動画を生成する(Image-to-Video)。 注意: 動画生成には数分かかる場合があります(ポーリングで完了を待機)。
MCP サーバーの情報と利用可能なツール・モデルの一覧を返す。
No input schemas visible for parameters. The tool definitions show parameter dictionaries with type and description, but it is unclear whether these are enforced as JSON Schema constraints in the actual MCP tool registration. Source code shows @mcp.tool() decorator without explicit schema argument. This means LLMs receive no machine-readable parameter validation, they must infer constraints from text descriptions alone, which is error-prone.
Output schemas not documented. Tools return generic dicts (e.g., combine_audio_video returns {'output_path': str, 'video_path': str, 'audio_path': str, 'error': str, 'code': str}). LLMs have no way to know what fields are present in success vs error cases. generate_music and others return custom structures but with no formal schema declaration, forcing LLMs to guess field names and types.
Error handling is generic and non-actionable. combine_audio_video returns {'error': str, 'code': str, 'hint': str}, but the LLM receives no guidance on whether to retry, ask the user, or give up. Pattern:recovery-guide requires 'what to do next' in error responses. Current errors like 'FFMPEG_NOT_FOUND' with hint 'brew install ffmpeg' are actionable, but generate_* tools likely throw opaque API errors with no recovery path.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Current MCP spec supports tool annotations to mark read-only operations, destructive operations, and idempotent ones. This server has WRITE risk for combine_audio_video but does not surface this in the MCP tool definition. Agents cannot distinguish safe exploratory calls from destructive ones without annotations.
server_info tool is self-referential and provides little value for agent planning. Its description 'MCP サーバーの情報と利用可能なツール・モデルの一覧を返す' (MCP server info and tool/model list) is generic. LLMs already have tool descriptions in the protocol; calling server_info to learn about tools is redundant. This tool should either be removed or redesigned to return dynamic, actionable metadata (e.g., available model names, current quotas, pricing tiers).
reference_image parameter in generate_image accepts both 'GCS URI' and 'local path' but lacks an enum or clear validation. LLMs may pass invalid paths or malformed URIs. Description should clarify: 'A path to an image file: either a local filesystem path (e.g. /tmp/image.png) or a GCS URI (e.g. gs://bucket/image.png). Local paths must be within the configured output directory.'
image_gcs_uri parameter in generate_video_from_image requires a GCS URI but the description says '参照画像の GCS URI (例: gs://bucket/image.jpg)'. This is too minimal, it does not state whether the bucket must be publicly readable, whether the service account has permissions, or what happens on permission errors. Add: 'GCS URI of an image file. The service account running this server must have read access to the bucket.'