MCP server for AI image generation with OpenAI DALL-E, Google Imagen, Gemini, and Flux via Replicate
This MCP server defines 4 image generation tools with explicit Zod schemas and basic descriptions. All tools are explicitly registered with inputSchema objects and handler functions visible in src/server.ts. Schemas use proper Zod validation with type constraints (enums, optional fields, positive integers). However, descriptions are terse (50-100 chars), parameter descriptions are minimal or absent, output schemas are not documented, and error handling lacks recovery guidance. The naming convention 'image.generate.{provider}' is clear and follows verb-noun patterns. All parameters are properly typed via Zod but lack detailed documentation about constraints, formats, and relationships. Output is returned as mixed text+image content parts without formal schema documentation. No toolAnnotations (readOnlyHint/destructiveHint/idempotentHint) are present despite all tools being write operations (WRITE risk classification).
Generate an image using Google Gemini via @google/genai (default gemini-2.5-flash-image-preview). Requires GOOGLE_API_KEY.
Generate an image using Google (e.g., Imagen 3). Requires GOOGLE_API_KEY and GOOGLE_IMAGEN_ENDPOINT. Returns a saved file path and optional base64.
Generate an image using OpenAI (default model gpt-image-1). Returns a saved file path and optional base64.
Generate an image using Replicate models: Flux 1.1 Pro (default), Qwen Image, or SeedDream-4. Requires REPLICATE_API_TOKEN.
Parameter descriptions are missing or incomplete. Zod schema shows type constraints but Zod objects lack descriptions for most parameters (width, height, size, format, seed, quality, style, background, returnBase64, filenameHint). LLMs cannot infer the purpose of 'size' vs 'width'/'height', whether 'format' applies to output or input, or what 'quality' means per provider.
Output schema is not documented. Tools return a mixed array of {type:'text', text:'...'} and {type:'image', data:'...', mimeType:'...'} but no formal schema is provided. LLMs cannot predict what fields to expect (e.g., does 'path' appear in the text? what format is the base64?). Clients must infer structure from code inspection.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 55 | - | v1 |
Tool descriptions lack WHEN/HOW guidance. 'Generate an image using OpenAI (default model gpt-image-1)' states WHAT but not WHEN to use OpenAI vs Google vs Replicate, what dependencies exist (OPENAI_API_KEY required), or what the output represents. LLMs cannot decide between providers.
Missing tool annotations for write operations. All 4 tools modify state (generate and save files) and carry financial cost (API calls to OpenAI, Google, Replicate). Zod schemas lack .describe() calls and tools lack destructiveHint or cost warnings. LLMs treat them as read-only safe operations.
Parameter constraints not documented in descriptions. 'width' and 'height' are marked as positive integers via Zod but no description states minimum (e.g., 256px), maximum (e.g., 2048px), or typical values. 'size' parameter is documented as 'Size string (e.g. 1024x1024)' but does not state format exactly, whether it overrides width/height, or valid values per provider.
Error handling lacks recovery guidance. Tool implementations (visible in providers/*.ts imports) are not shown, but handler functions in server.ts have no try-catch blocks visible; errors bubble up as raw exceptions. LLMs receive no guidance on why a generation failed (rate limit? invalid prompt? API key missing?) or what to do next.
Environment variable dependencies are mentioned in descriptions but not enforced at tool level. 'image.generate.openai' description states it uses 'default model gpt-image-1' but code shows 'OPENAI_API_KEY' is required via process.env check. LLMs do not know that calling this tool without OPENAI_API_KEY will fail.
Overlapping and ambiguous parameter semantics across tools. All tools accept both 'width'/'height' and 'size' parameters but no description clarifies: does 'size' override width/height? Are they mutually exclusive? What happens if both are provided? LLMs cannot reason about priority or expected behavior.