MCP server for AI image generation with DALL-E, GPT-Image-1, and extensible model support
Single tool 'generate_image' has a solid verb_noun name and comprehensive parameter documentation with type information and defaults. However, the tool description lacks actionable context about when/why to use it, and critically, no output schema is documented despite the response type being mentioned. Parameter descriptions are present but inconsistent in clarity. The tool is registered via @mcp.tool() decorator with explicit schema in code, so inferred scoring does not apply. Overall: above-average parameter definition, but missing output documentation and recovery guidance keeps this in fair territory.
Generate images from text descriptions using AI models. Args: prompt: Text description of the desired image style: Style preset (default, photorealistic, illustration) size: Image dimensions (1024x1024, 1792x1024, 1024x1792) n: Number of images to generate (currently only 1 supported) model: Specific model to use (dalle-3, dalle-2, gpt-image-1) Returns: ImageGenerationResponse with image URLs and metadata
Output schema not documented. Tool description mentions 'ImageGenerationResponse with image URLs and metadata' but the actual response fields, types, and structure are not specified. LLMs cannot infer what fields to extract from the response or plan downstream operations.
Tool description is generic and under-specifies context. It reads like API documentation ('Generate images from...') rather than LLM-optimized guidance. No statement of WHEN to use this tool vs alternatives, what prerequisites exist (OpenAI API key, account), or typical use cases. Expected 10 - 1024 chars with actionable guidance; this is 173 chars but lacks 'when' and 'why' framing.
No error handling or recovery guidance in tool definition. Code logs errors and raises RuntimeError but does not document what can go wrong (API rate limits, invalid model, bad API key) or guide the LLM on recovery. Per pattern:recovery-guide, errors must tell the agent what to do next.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 54 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 11 | - | v1 |
Parameter 'n' description claims 'currently only 1 supported' but parameter type allows any integer and default is 1. This constraint is buried in the description, not enforced as an enum or min/max validation rule. LLMs cannot parse prose constraints, use formal JSON Schema validation (enum: [1]) instead.
Parameter 'model' defaults to null and description says 'Specific model to use (dalle-3, dalle-2, gpt-image-1)' with no enum constraint. Free-form string invites hallucinated model names. Should be an enum: ['dalle-3', 'dalle-2', 'gpt-image-1', 'default'] or similar to let LLMs pick valid options.
Parameter descriptions lack format/range guidance. 'size' description lists options (1024x1024, 1792x1024, 1024x1792) in prose instead of as an enum constraint. 'style' description lists presets but no enum. Without formal constraints, LLMs guess at valid values.
No output pagination or limiting strategy documented. If image generation returns multiple results or metadata, tool description should state max results and whether pagination is supported. Currently opaque to the LLM.