An MCP (Model Context Protocol) server that integrates with OpenAI's gpt-image-1 model and Google Gemini for text-to-image generation and image editing services
The server provides 12 tools with basic structure, but exhibits significant gaps in schema completeness, parameter descriptions, and error handling guidance. Most tools have generic, short descriptions (under 50 chars) that fail to guide LLM tool selection. Parameter schemas are present for most tools but lack constraints (enums, ranges, patterns) and many parameter descriptions are minimal or missing. No output schemas are documented. Error handling is absent, no recovery guidance, retryability hints, or actionable error messages visible. The architecture shows effort (multi-provider abstraction, template system, proper validation utilities exist) but the MCP interface itself is underspecified. Tool composition is reasonable (separate concerns for generation, editing, listing) but lacks idempotency guarantees and confirmation patterns for destructive ops.
Delete a generated image from storage
Edit or manipulate existing images with text prompts and optional masks using the provider API
Estimate the cost of generating images with specified parameters
Generate images from text prompts using OpenAI DALL-E, Google Gemini, or other configured providers
Get detailed metadata and generation information for a specific image
Get detailed capability information for a specific image generation model
Check the health and status of configured image generation providers
No output schemas documented for any tool. LLMs cannot plan subsequent calls or know which fields are available without executing. This breaks tool chaining patterns.
Parameter descriptions are generic and lack actionable constraints. E.g., 'size' param description 'Image dimensions as WxH (e.g., 1024x1024, 1536x1024, 1024x1536, 3840x2160) or auto' uses examples instead of enums. 'quality' has no explanation of quality differences. Agents cannot disambiguate when to use high vs medium.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Get details for a specific prompt template including parameters and examples
List previously generated images with filtering and pagination options
List available image generation models and their capabilities across all configured providers
List available prompt templates for image generation organized by category
Render a prompt template with provided parameter values to generate the final prompt text
Tool descriptions are under 50 chars and fail to answer 'when to use this instead of a similar tool'. E.g., 'Generate images from text prompts using OpenAI DALL-E, Google Gemini, or other configured providers' (105 chars, but generic). 'Edit or manipulate existing images' lacks detail on what edits are supported. LLMs struggle to select the right tool.
delete_image has no confirmation or dry-run pattern. Destructive operations must support user confirmation before execution. LLMs make mistakes, a bare delete with no safety check risks data loss.
No error handling or recovery guidance visible in tool definitions. If an API call fails (rate limit, invalid model, authentication), LLMs have no hint on what to retry or how to self-correct. Error messages must be actionable.
Enum constraints missing. 'output_format' accepts 'png, jpeg, webp' but is defined as free-text string, not enum. 'quality' (auto, high, medium, low) and 'style' (vivid, natural) and 'moderation' (auto, low) should be enums. Free-form strings invite hallucinated values.
No idempotency guarantees. generate_image with n=3 called twice will produce 6 images; no de-duplication or idempotency key mechanism. Agents retry on ambiguous failures, non-idempotent tools risk duplicate charges and duplicate records.
render_template expects a 'parameters' object (type: object) but has no schema for what fields that object should contain. LLMs cannot know which template parameters are valid without calling get_template first, adds discovery overhead.
list_generated_images and similar pagination tools lack documentation on what happens when offset > total or when no results match filters. Return structures (total_count, has_more, next_cursor) are not specified. LLMs cannot implement continuation loops.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in tool definitions. delete_image should be marked destructiveHint=true. All read-only tools (list_*, get_*) should be marked readOnlyHint=true. These annotations guide agent safety.