MCP server for generating images using DALL-E 3
The server implements two image generation tools with explicit Zod schema definitions and reasonable parameter coverage. However, descriptions lack depth, output schemas are not formally documented, error handling is minimal, and the composition does not follow single-responsibility patterns. Tools use clear verb_noun naming (generate_image, generate_image_batch) and accept relevant parameters with type constraints via enums. Parameter descriptions are present but brief (average ~50 chars, below the rubric baseline of 72). The batch tool conflates multiple concerns (sequential processing, error accumulation) that could be split. No documentation of return structure or pagination strategy. Error messages are generic ('Error generating image') without recovery guidance. Security: credentials are injected via environment variable (correct), but output sanitization and input validation are minimal.
Generate an image using DALL-E 3. Returns the saved file path and revised prompt.
Generate multiple images using DALL-E 3. Processes sequentially and returns all file paths.
Output schemas not formally documented. Tool descriptions state what is returned ('file path and revised prompt') but do not specify the structured response schema (filePath: string, revisedPrompt: string, etc.). LLMs cannot reliably extract or chain these outputs without explicit schema documentation.
Error handling lacks recovery guidance. Errors return generic text like 'Error generating image: {err.message}' without actionable recovery steps. Rate limit retries are implemented internally (10s delay on 429) but not exposed to the LLM, so it cannot understand retry semantics or backoff strategy.
generate_image_batch conflates two concerns: batch processing and error accumulation. The tool returns partial results with per-item error strings mixed into a single text response. This violates single-responsibility and makes error classification (retryable vs fatal) ambiguous for the LLM.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Parameter descriptions are present but lack depth. Example: 'Image dimensions' (17 chars) does not explain constraints, trade-offs, or guidance on which size to choose. Descriptions should include expected use cases, e.g., 'Image dimensions: 1024x1024 (square, default), 1024x1792 (portrait), 1792x1024 (landscape). Portrait/landscape sizes may be rejected for certain prompts; use 1024x1024 if generation fails.'
No input validation or constraint documentation for the 'prompt' parameter. The description 'The image generation prompt' is generic. DALL-E 3 has constraints (max length ~4000 chars, no restrictions on style but some content policies). These constraints should be documented so LLMs avoid invalid prompts.
Tool descriptions do not state whether operations are idempotent or have side effects. generate_image writes to the filesystem, the LLM should understand this is a state-modifying operation. This affects retry safety and error recovery strategies.