Multi-provider MCP server for AI image generation (Gemini, Jimeng) with intelligent provider selection
Mixed quality across 8 tools. Strengths: all tools have descriptions, most have input schemas with type constraints, tool annotations present (destructiveHint/idempotentHint), comprehensive parameter documentation in generate_image and img_generate. Weaknesses: generate_image and img_generate share near-identical responsibility (image generation) creating ambiguity; some tools (output_stats, list_models) have minimal/generic descriptions; schemas visible for Python tools but TypeScript tool (img_generate) inferred from description only; maintenance tool lacks guidance on consequences of destructive operations; no clear error handling examples in source; output schemas not explicitly documented; missing pagination details in list_image_jobs.
Generate images using Gemini or other AI providers with support for multiple providers, aspect ratios, and advanced options
Retrieve information about a previously executed image generation job
Generate an image (Jimeng 4.5 / Seedream 4.5). Use ratio for aspect ratio (1:1/3:4/4:3/16:9/9:16/21:9) or size (WxH/2K/4K). Output includes MEDIA and Markdown image link when URL is available.
List recent image generation jobs with optional filtering
List available AI models and their capabilities
Perform maintenance operations including cleanup of expired files, local file cleanup, quota checks, and database hygiene
Get statistics about generated images and output directory usage
Two image generation tools with overlapping responsibility: 'generate_image' and 'img_generate' both produce images but with different provider/parameter models. LLMs will struggle to choose between them. Split by provider (generate_image_gemini, generate_image_jimeng) or merge into a single canonical tool.
output_stats has a 21-character description ('Get statistics about generated images and output directory usage') that is generic and does not clearly state WHY an LLM should call this tool or when it's useful. Missing context: e.g., 'Call before generating to check available storage quota and directory size. Returns disk usage and file counts.'
list_models description ('List available AI models and their capabilities') does not explain what the response structure is or how to use it. LLMs need: 'Returns available models per provider. Use this to understand which models support what features (e.g., grounding, extended thinking) before calling generate_image.'
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
Upload image files to Files API for use in subsequent operations
img_generate tool is defined in TypeScript (openclaw-plugin/index.ts) and tool definition is not visible in provided source code. Tool registration and schema are INFERRED from description only.
maintenance tool accepts 'dry_run' but does not document what the tool returns when dry_run=true (does it return what WOULD be deleted? A count? A list?). Destructive tools require explicit confirmation/dry-run output documentation to guide agent decision-making.
list_image_jobs accepts 'limit' and 'offset' for pagination but does not document: what is the default limit? Is there a maximum? What field(s) identify the next page? Missing pagination guidance forces LLMs to guess reasonable values.
No explicit output schemas documented in source code for any tool. generate_image description states 'returns images as real MCP image content blocks, and also provides structured JSON with metadata' but the JSON structure is not formally documented. LLMs cannot plan downstream operations without knowing what fields to extract.
generate_image has 11 parameters with only 3 enum-constrained (aspect_ratio, model_tier, thinking_level). Others like 'provider' (free string) and 'resolution' (free string) invite hallucinated values. Add enums for provider (gemini|jimeng|openai|...) and resolution presets (1024x768|1536x1024|...) or document exact format.
generate_image parameter 'system_instruction' has no description. Is this a prompt-like directive? Does it affect all generated images or only certain models? Ambiguous parameter names without context cause misuse.
No error recovery guidance in any tool descriptions. If generate_image fails due to invalid prompt or quota exceeded, how should the LLM recover? Missing error classification (retryable vs user-fixable) forces agents to retry blindly.