CLI for AI image generation using OpenAI, Google Gemini, and xAI Grok
The server defines 3 image generation tools with complete input schemas and adequate descriptions. All tools follow verb_noun naming (generateOpenAIImage, generateGeminiImage, generateGrokImage) and include comprehensive parameter constraints via enums. However, tool descriptions are generic and lack LLM-optimization guidance on when to select each tool vs alternatives. Parameter descriptions are present but often minimal (e.g., 'Image size' without format guidance). Output schema is not documented. Error handling in the tool definitions is absent, no guidance on what errors occur or recovery steps. The server is a working CLI tool repurposed for MCP, not designed ground-up for agent composition.
Generate or edit images using Google Gemini's image generation models
Generate or edit images using xAI Grok's image generation models
Generate or edit images using OpenAI's image generation models
Tool descriptions lack LLM-optimization and selection guidance. All three tools have identical generic descriptions ('Generate or edit images using [Provider]'s image generation models'). LLMs cannot distinguish when to choose OpenAI vs Gemini vs Grok without explicit guidance on cost, speed, quality, or feature differences.
No documented output schema. Tool definitions do not specify what fields are returned on success or what structure is used. Line 'console.log(JSON.stringify(result.data, null, 2))' in cli.ts suggests a result object with .data and .ok fields, but this is not formally documented for the MCP schema.
Error handling guidance missing. No error responses documented, no recovery instructions, no categorization of retryable vs fatal errors. The server returns {'ok': false, 'error': <message>} but does not guide LLMs on next steps (e.g., 'Try reducing image resolution' or 'API rate limit, retry in 60 seconds').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 56 | - | v1 |
Parameter descriptions are minimal and lack LLM-actionable details. Example: 'Image size' (Gemini tool) does not explain the consequence of each enum value or when to use 512 vs 1K vs 2K. 'Aspect ratio' lacks guidance on portrait vs landscape intent. Descriptions should match pattern baseline of 50-100 chars per parameter.
No tool composition or chaining guidance. All three tools write a local file (output_path). The server does not document whether results can be chained (e.g., edit output from one model with another) or how to reference generated images in subsequent calls.
Overlapping and confusing enum values across models. OpenAI uses 'gpt-image-2', 'gpt-image-2-2026-04-21', 'gpt-image-1.5'. Gemini uses 'gemini-3.1-flash-image', 'gemini-3.1-flash-lite-image'. Grok uses 'grok-imagine-image-2.0', 'grok-2-image'. No descriptions explain model capabilities, cost, or when each is appropriate. LLMs may pick randomly.
Missing idempotency and retry semantics. Tools write files but do not document whether repeated calls with same params and same output_path are safe or will fail. The CLI enforces --force flag, but MCP tool does not expose this safety mechanism clearly.