MCP server for OpenAI's GPT image generation and editing capabilities
Server defines 3 image generation/manipulation tools with generally good schemas and descriptions. Tool names are action-oriented (create-, edit-, varies-), which is appropriate. All three tools have detailed parameter schemas with proper JSON Schema typing. However, there are several quality gaps: (1) Output schemas are not documented, the return type structure is unclear from the tool definitions, making it hard for LLMs to plan downstream actions; (2) Error handling is minimal, no guidance on recovery paths, retryability, or user-fixable vs fatal errors; (3) Descriptions, while present, focus on API parameters rather than when/why to use each tool; (4) No tool annotations (destructiveHint, readOnlyHint, idempotentHint) despite all three tools being write operations (WRITE risk declared); (5) The schema uses Zod refinements that are not exposed to the server.tool() registration, potentially causing schema visibility issues for clients; (6) File output handling is complex (absolute path validation, multi-image indexing) but underdocumented in parameter descriptions. The three tools are well-structured with proper enums, min/max constraints, and optional fields, but the overall tool system lacks maturity in error guidance and composition clarity.
Generate images using OpenAI's gpt-image-1 model with comprehensive customization options for backgrounds, formats, quality, and size.
Edit existing images using OpenAI's gpt-image-1 model with optional mask-based editing, supporting quality and size customization.
Generate image variations of an existing image using OpenAI's gpt-image-1 model with customizable quality, size, and output format.
Output schema not documented. Tools return base64 data or file paths, but the response structure is not declared. LLMs cannot infer what fields to expect or how to chain subsequent calls.
No error handling guidance. Tools provide no recovery paths (e.g., 'Invalid prompt format, try rephrasing' or 'File too large, compress to <25MB'). LLMs receive raw errors with no actionable next steps.
No tool annotations despite write semantics. All three tools have WRITE risk but lack destructiveHint or idempotentHint annotations. LLMs cannot determine if tools are safe to retry or if they produce side effects.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Descriptions emphasize API parameters (background, moderation, output_compression) over user intent. Why would an LLM call 'create-image' vs 'edit-image'? The description should clarify use cases, not parameter lists.
Zod schema refinements not exposed to server.tool(). Code uses (createImageSchema as any)._def.schema.shape to work around refinements, but the actual validation logic (e.g., transparent background requires png/webp output_format) is enforced only in handler code, not visible in the schema. This breaks client-side validation and introspection.
Complex file_output semantics underdocumented. Parameter descriptions mention automatic indexing ('image_1.png', 'image_2.png' for n > 1) and auto-switching to file output if base64 exceeds 1MB, but these behaviors are not clearly stated in the schema descriptions, forcing LLMs to infer or guess.