A Model Context Protocol server for OpenAI's gpt-image-1 model
Single tool 'create_image' with well-structured schema and clear naming. The tool starts with 'create_' verb reflecting the action (pattern:tool compliant). Input schema is fully visible in src/index.ts with comprehensive parameter validation using Zod. Description is clear and action-oriented. However, output schema is not documented in the source code, the function saves images to disk and returns file paths, but the exact response structure is not explicitly defined or commented. Error handling in the source is basic (try-catch blocks with some error messages), but recovery guidance is minimal. The tool avoids exposing secrets (OPENAI_API_KEY is injected via environment variable, not as a parameter). No composition concerns, this is a single-purpose tool.
Generate new images using OpenAI's gpt-image-1 model
Output schema not documented. The tool saves images and returns file paths, but the exact response structure (fields, types, examples) is not visible in the code or comments. LLMs cannot plan downstream calls without knowing what fields are returned.
Error responses lack recovery guidance. The code contains generic error handlers (try-catch with console.error) but does not return structured error objects with actionable next steps. E.g., if the OpenAI API call fails, the agent receives a raw error message without guidance on whether to retry, adjust the prompt, or call a different tool.
No confirmation/dry-run pattern for write operations. create_image is destructive (writes files to disk, consumes API quota). The tool should support a 'dry_run' parameter or return a preview before committing, preventing accidental resource waste.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Parameter descriptions lack actionable constraints. E.g., 'prompt' is described as 'The text prompt for image generation (max 32000 characters)', good. But 'size' enum values lack rationale. Should explain: '1024x1024 = square, 1536x1024 = landscape, 1024x1536 = portrait. Use auto for gpt-image-1 to choose automatically.' This helps LLMs pick the right option.