Generate images from text with Claude! Easy-to-install MCP server for Google Gemini 2.5 Flash Image generation, editing, and composition.
This server has 4 image manipulation tools with reasonable naming (verb-noun format) and documented schemas. However, there are significant gaps: descriptions are generic and lack LLM-optimized guidance, parameters lack validation constraints and relationship documentation, output schemas are not documented, and error handling does not provide recovery guidance. The tools are composable but would benefit from clearer parameter constraints and more specific descriptions. Tool definitions are directly visible in src/index.ts, but output schema documentation is absent from the code review.
Compose multiple images together using a text prompt with Google Gemini 2.5 Flash
Edit an existing image using a text prompt with Google Gemini 2.5 Flash
Generate an image from a text prompt using Google Gemini 2.5 Flash
Apply a style to an image using a text description with Google Gemini 2.5 Flash
Descriptions lack LLM-optimized guidance. All four tools use generic descriptions (e.g., 'Generate an image from a text prompt using Google Gemini 2.5 Flash') without explaining WHEN to use each tool vs alternatives, what constraints exist, or what the typical workflow is. Descriptions should be 50 - 200 characters and answer: what does it do, when should the LLM call it, what does it return?
Output schemas are not documented. The code shows successful responses may contain 'imageBase64' and 'mimeType' fields, but this is not visible in the tool definitions provided. LLMs cannot plan downstream calls (e.g., save the image, send it in an email) without knowing the response structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 25 | - | v1 |
Parameters lack validation constraints and relationship documentation. The mimeType parameter states it is optional but does not document allowed values (the code shows ALLOWED_MIME_TYPES = ['image/png', 'image/jpeg', 'image/jpg', 'image/webp', 'image/gif']). Imagepaths in compose_images are array items without min/max length constraints. saveToFilePath is undocumented regarding path traversal rules.
Error handling does not provide recovery guidance. The code includes retry logic and timeouts (MAX_RETRIES=3, REQUEST_TIMEOUT=60s) but tool definitions do not advertise error conditions, retry eligibility, or what an LLM should do if a call fails (e.g., 'Try again with a shorter prompt' or 'Check image file format is JPEG/PNG').
Parameter descriptions are minimal. Across all tools, parameter descriptions are 1-2 sentences and lack actionable detail. For example, 'Preferred output MIME type (e.g., image/png, image/jpeg)' should specify that only png/jpeg/webp/gif are supported, what happens if an unsupported type is requested, and what the default is.
No distinction between tool purposes. generate_image, edit_image, and compose_images all manipulate images but the descriptions do not clarify when to use edit_image vs compose_images (both modify existing images). An LLM may conflate them or pick the wrong tool.
No documentation of file format or encoding. The code handles base64 encoding internally but the tool interface does not explain whether imagePaths expects relative or absolute paths, what encodings are supported, or whether paths are resolved relative to the working directory or a fixed output directory.