MCP server for generating and editing images using OpenAI's image APIs (DALL-E)
Server provides 2 tools with reasonable parameter schemas and descriptions, but has significant gaps in clarity, error handling guidance, and output documentation. Tool names follow verb_noun conventions (generate_image, edit_image) which is good. However, descriptions lack LLM-optimization: they are moderately detailed but do not clearly distinguish WHEN to use each tool vs alternatives, or what the response structure looks like. Parameters are well-typed with enums for size/quality, but descriptions are generic ('The text prompt...') without guidance on prompt engineering or length limits. Output schemas are not documented in the provided code, no indication of what fields are returned or how pagination/multiple images are handled. Error handling mentions logging but provides no recovery guidance (e.g., 'if API rate-limited, retry after X seconds'). The code shows helper functions for image validation and file saving, but these safeguards are not reflected in tool descriptions, so LLMs won't know input constraints exist.
Edit an existing image by providing a text prompt describing the edits. Optionally provide a mask to specify which areas to edit
Generate one or more images from a text prompt using OpenAI's image generation API
Output schema not documented. Code saves images to disk and returns file paths, but tool descriptions do not specify return format, fields, or how multiple images (n > 1) are handled. LLMs cannot plan downstream operations without knowing the response structure.
Descriptions lack LLM-optimization guidance. Both tool descriptions are under 100 characters and do not explain WHEN to use generate_image vs edit_image, what prompt quality matters, or prerequisites (e.g., 'provide absolute file paths for edit_image'). Descriptions should be 50 - 200 chars and answer: what does it do, when should I call it, and what should I expect back?
Input parameter descriptions are generic and lack actionable constraints. The 'prompt' parameter says 'The text prompt for image generation' but does not explain: Is there a length limit? Should the prompt be detailed or brief? What happens if the prompt is empty? The 'size' and 'quality' enums are clear, but 'prompt' description should explain best practices and constraints.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
File path parameters require absolute paths but descriptions do not warn or guide. 'image_path' is marked required with description 'Absolute path to the input image (PNG, <= 25MB)' but LLMs often pass relative paths or names without full paths. Description should say: 'Provide the full absolute file path (e.g., /home/user/images/photo.png). Relative paths will fail.'
No error recovery guidance. Code logs errors but tool descriptions provide no guidance on what to do if API fails, rate limits, or validation rejects input. Error responses should include actionable next steps: 'Image validation failed: PNG format required. Try converting your input with PIL or ImageMagick, then retry.'
No documentation of default behavior when parameters are omitted. Code shows defaults (n=1, size='auto', quality='auto') but tool descriptions do not explain: what happens when quality='auto'? Does the API choose based on prompt, or is there a fallback? Does size='auto' scale based on aspect ratio? LLMs need explicit defaults to reason about behavior.