MCP server exposing OpenAI's gpt-image-2 for image generation and iterative editing.
This server exhibits strong definition quality overall. All 6 tools have clear verb-noun naming patterns (generate_image, edit_image, start_edit_session, continue_edit_session, end_edit_session, list_edit_sessions). Tool descriptions are comprehensive (194 - 300+ chars), well above the 10 - 1024 baseline. All parameters have type definitions and descriptions. Schemas are explicitly defined using Zod with proper enum constraints for categorical parameters (quality, background, output_format). Output schemas are documented. Error handling is present with actionable recovery guidance via describeOpenAIError(). The server demonstrates LLM-aware prompt engineering in descriptions (e.g., 'use ALL CAPS or quote literal text you want rendered verbatim'). Key strengths: extensive parameter documentation (size constraints explained, file format specs detailed), composition-friendly tool chain (session tools properly ordered: start→continue→end), and security consciousness (user parameter for abuse monitoring, no API keys exposed as params). Weaknesses: moderate issues around parameter relationships (size/quality/background interdependencies not fully cross-documented), limited per-parameter constraint details in some schemas (e.g., output_compression range 0 - 100 is mentioned but not formally constrained in schema), and tool descriptions could be more concise for some parameters (several descriptions exceed 200 chars, risking token waste). Risk annotation is present (WRITE/READ_ONLY) but security implications of stateful sessions and in-memory storage not fully documented.
Apply another edit turn to an existing session. The previous turn's output image is used as the input. Use short, focused prompts like "make the sky more orange" or "add a small boat on the horizon"; include "keep everything else the same" to limit drift. Returns the new image and the updated session.
Edit or compose images with gpt-image-2. Give 1–8 input images plus a text prompt; optionally include a PNG mask whose transparent regions mark what to change (mask applies to the first image). Great for: swap backgrounds, retouch products, combine multiple reference images into one composition, maintain a character across scenes. gpt-image-2 always processes inputs at high fidelity (no input_fidelity knob needed). Pass background: "transparent" (with png/webp output) to cut the subject out onto alpha. The edited image is saved to disk and returned inline.
Free an iterative-edit session. Safe to skip — sessions are in-memory only and are discarded on server restart — but calling this frees the memory immediately.
Generate an image from a text prompt using OpenAI's gpt-image-2 model. The image is written to disk and also returned inline so you can see it. gpt-image-2 handles photoreal, illustrations, infographics, multilingual text (incl. CJK), and complex structured visuals. Transparent backgrounds are supported via background: "transparent" (png/webp output only). Sizes accept presets or any custom "WxH" where edges are multiples of 16, max edge 3840px, aspect ratio within 1:3–3:1, total pixels 655K–8.29M.
Parameter constraint documentation incomplete for numeric parameters. output_compression is described as 0 - 100 range in text but not formally constrained in JSON Schema (missing minimum/maximum). This invites LLMs to pass invalid values.
Parameter interdependencies not fully cross-documented. background='transparent' requires output_format='png'|'webp', but the description of background does not explicitly state this constraint, nor does the parameter description for output_format explain its dependency. This forces LLMs to infer relationships.
Session state lifecycle not fully documented. start_edit_session creates in-memory state; continue_edit_session and end_edit_session depend on this state. The risk of session loss on server restart is mentioned once in end_edit_session but not prominently in start_edit_session or continue_edit_session descriptions, risking agent confusion.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 74 | 2025-06-18+ | v2 |
List all active iterative-edit sessions and their metadata (creation time, turn count, last prompt, last image path).
Begin a stateful multi-turn edit session. Returns a session_id you then pass to continue_edit_session to iteratively refine the image (each turn uses the previous turn's output as the input). Use end_edit_session when done.
Tool description length exceeds 200 chars for several parameters (prompt, size, background, output_format). While detailed, descriptions this long risk token waste and may be harder for LLMs to parse quickly. Consider summarizing and moving advanced constraints to separate documentation.
list_edit_sessions has empty input schema (no parameters), but output schema is not visible in the provided code. Cannot verify that session metadata return fields match the data stored by start_edit_session (lastImageBase64, lastImageMime, lastImagePath, lastPrompt, outputDir, outputFormat). This breaks the tool-chain pattern.