Image Generation MCP Server - Provides tools for generating images using Google Gemini, OpenAI, and xAI Grok APIs.
Two image generation tools with well-documented parameters and comprehensive descriptions. Both tools have complete JSON Schema definitions with type constraints (Literal enums) and default values. Descriptions are detailed and informative, exceeding the 50-200 char guideline slightly but remaining useful for LLM selection. Parameter descriptions clearly explain constraints, valid values, and behaviors. Output schemas are documented inline. However, both tools exhibit identical concerning patterns: (1) they write files to disk without safeguards against path traversal or overwrite conflicts, (2) error handling lacks recovery guidance for LLMs, (3) no permission gates despite destructive file operations. Tool naming follows verb_noun convention ('generate_image_*'). Composition is clear, each tool wraps a single API provider without conflating concerns.
Generate an image using Google Gemini API (Nano Banana 2 by default). Default model is gemini-3.1-flash-image-preview (Nano Banana 2), released Feb 2026. It combines the quality of Nano Banana Pro with the speed of Gemini Flash. OUTPUT LOCATION: Images save to current working directory by default with auto-generated names like "gemini_20240126_143052_a1b2c3d4.png". Use output_path for custom location. MODELS: - gemini-3.1-flash-image-preview (default): Nano Banana 2 -- fastest + highest quality, 4K resolution, character consistency for up to 5 characters, 14-object fidelity, expanded aspect ratios, advanced text rendering, optional thinking mode - gemini-2.5-flash-image: Previous generation (Nano Banana 1), max 2K. WARNING: may be deprecated June 17, 2026 with the broader 2.5 Flash family. - gemini-3-pro-image-preview: Pro-tier quality for production assets ASPECT RATIOS (Nano Banana 2 adds ultra-wide/tall): - Square: 1:1 - Landscape: 3:2, 4:3, 5:4, 16:9, 21:9 - Portrait: 2:3, 3:4, 4:5, 9:16 - Ultra-wide (NB2 only): 4:1, 8:1 - Ultra-tall (NB2 only): 1:4, 1:8 IMAGE SIZE (Nano Banana 2): - 512px: Minimum, fastest (~3-8s) - 1K: Default if not specified (~5-10s) - 2K: Balanced quality (~10-15s) - 4K: Maximum resolution (~15-25s) THINKING MODE (Nano Banana 2 only): - minimal (default when enabled): Light reasoning pass, small quality boost - high: Deep reasoning, best for complex scenes or precise compositions
Generate an image using OpenAI's GPT Image API. OUTPUT LOCATION: Images save to current working directory by default with auto-generated names like "openai_20240126_143052_a1b2c3d4.png". Use output_path for custom location. MODELS: - gpt-image-2 (default): Flagship model (April 2026). ~2x faster than gpt-image-1, ~99% text accuracy across Latin/CJK/Hindi/Bengali, supports arbitrary resolutions (divisible by 16, aspect ratio 1:3 to 3:1, up to 3840x2160 experimental). - gpt-image-1.5: ~4x faster than gpt-image-1, 20% cheaper, good text rendering - gpt-image-1: Original model, deprecated October 23, 2026 - gpt-image-1-mini: Cost-efficient variant of gpt-image-1 SIZES: 1024x1024 (square), 1536x1024 (landscape), 1024x1536 (portrait), auto (model decides) QUALITY: low (fastest), medium, high (best detail), auto (model decides based on prompt) BACKGROUND: opaque (solid background), transparent (PNG only, for logos/icons), auto
Unsafe file write without path validation or overwrite prevention. Both tools accept output_path parameter and write to it directly using open(output_path, 'wb'). No check for path traversal (e.g., '../../etc/passwd'), symlink attacks, or confirmation before overwriting existing files. An agent could be tricked into writing to arbitrary locations or destroying user data.
Error responses lack recovery guidance. When API calls fail (e.g., missing API key, HTTP 400, no candidates in response), the tool raises ValueError with minimal context. LLMs receive raw error messages like 'Gemini API error (400): ...' with no suggestion for next steps. Should include recovery actions: 'Check GOOGLE_API_KEY is set', 'Verify prompt is not too long', or 'Try a simpler prompt'.
No permission gates or audit logging. These tools perform I/O operations (file writes) that could be abused by a compromised or untrusted agent. No user ID is captured, no audit trail of who called what, no permission checks before execution. Should log: caller identity, timestamp, output_path, model used, and result status.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 71 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
API keys exposed via environment variables without validation or rotation guidance. Tools read GOOGLE_API_KEY and rely on OpenAI client library (which reads OPENAI_API_KEY). No mechanism to rotate keys, no rate limiting, no validation that credentials are valid before accepting a request. If an agent runs in an untrusted environment, keys could be exfiltrated.
Missing idempotency guards and duplicate detection. If an agent retries a failed call with the same parameters, both tools will generate a new image and write a new file (with a unique UUID). This violates idempotent-operation pattern. Should either deduplicate by prompt hash or document that repeated calls are NOT safe and may incur extra API costs.
No input validation on output_path. Both tools use the output_path parameter directly in open(output_path, 'wb') with no checks. An agent could pass paths like '/dev/zero', '/etc/passwd', or '../../secrets.txt'. Should validate that output_path is within an allowed directory (e.g., current working directory or a temp folder) and reject paths with '/..' or leading '/'.
generate_image_openai description mentions arbitrary resolutions 'divisible by 16, aspect ratio 1:3 to 3:1, up to 3840x2160 experimental' but the size parameter is constrained to a fixed enum [1024x1024, 1536x1024, 1024x1536, auto]. This discrepancy is confusing, the description implies custom sizes are supported, but the schema does not. Either add a 'custom_size' parameter accepting arbitrary dimensions, or clarify that arbitrary sizes are NOT exposed via MCP.
generate_image_gemini's thinking_level parameter is marked optional but the description states 'Only supported by gemini-3.1-flash-image-preview; ignored for other models.' This creates a silent failure mode, if an agent selects gemini-2.5-flash-image and sets thinking_level='high', the parameter is silently ignored. Should either validate and reject invalid combinations early with a clear error, or document the constraints more explicitly in the parameter description.