An MCP server that generates and transforms images using Google's Gemini model with support for text-to-image generation, image transformation, and multi-turn image editing conversations.
The server defines 3 tools with visible schemas and descriptions. However, there are significant gaps in quality: (1) descriptions are functional but generic, lacking LLM-optimized guidance on when to use each tool and what distinguishes them; (2) parameter descriptions exist but are brief and lack constraints (format examples, ranges, enum values); (3) output schemas are not documented, the response structure is opaque to downstream callers; (4) error handling is minimal, no actionable recovery guidance; (5) no input validation messaging or constraints visible in parameter definitions. Tool naming is clear and verb-driven (generate_, transform_, chat_), which is good, but the definitions lack the richness needed for confident LLM decision-making and multi-tool composition.
Have a multi-turn conversation about an image. Send an image and chat prompt, get a text response from Gemini about the image. Supports continuous conversation by passing previous messages.
Generate an image based on the given text prompt using Google's Gemini model.
Transform an existing image based on the given text prompt using Google's Gemini model.
No output schema documented. Tools return image bytes or text, but the response structure is not formally defined. LLMs cannot plan downstream calls or extract required fields (e.g., image URL, dimensions, format) without guessing.
Parameter 'model' lacks constraints. Description says 'default: gemini-3-pro-image-preview' but does not list other valid options, format restrictions, or how to discover available models. LLMs will guess invalid model names.
Tool descriptions lack context on when to use each vs. the others. 'Generate' vs 'transform' distinction is not explained. 'chat_about_image' mentions multi-turn but does not explain how conversation state is managed or when to use it instead of generating a new image.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 35 | - | v1 |
Parameter 'encoded_image' format constraint is documented in description ('data:image/[format];base64,[data]') but is informal. No regex pattern or format constraint in schema. Parameter validation happens server-side with no LLM-facing guidance for malformed input.
No error handling or recovery guidance. If Gemini API fails (invalid model, rate limit, authentication), the tool will raise an exception with no actionable message for the LLM. LLM cannot self-correct or retry intelligently.
'previous_messages' parameter in chat_about_image is documented as 'Optional list of previous message objects' but lacks structure definition. How should messages be formatted? What fields are required? LLMs will guess and pass malformed structures.
No idempotency hints. Image generation and transformation are stateful operations that may have side effects (file I/O, API charges). Tools do not declare idempotentHint=false or destructiveHint=true. Agents may not know these are risky to retry.