MCP server for AI vision generation (images, video) via Claude Code
The server defines 3 tools with reasonable naming and visible input schemas. Descriptions are present and moderately detailed (averaging ~250 chars), but lack actionable guidance on when to use each tool versus alternatives. Parameters are typed (string, array) but lack enum constraints where beneficial (style, size). Output formats are JSON strings rather than structured objects, making post-processing harder for agents. Error handling is minimal, no guidance on recovery or invalid input self-correction. The tools follow single-responsibility principle (generate, refine, list), but composition guidance is weak.
Generate an image from a text prompt, optionally using reference images or feedback. Args: prompt: Detailed description of the image to generate (be specific). style: Optional visual style. One of: modern, minimal, professional, playful, dark, realistic, artistic, flat. size: Optional image size/aspect ratio. One of: square, portrait, landscape, wide, 4k, hd. model: Optional litellm model string to override the default (e.g. "gemini/gemini-2.5-flash-image", "gpt-image-1"). filename: Optional custom filename (without extension). working_directory: Directory where the image should be saved. Pass your project's working directory. feedback: Optional feedback on a previous generation (e.g. "make it darker", "remove the text"). When provided, prompt is treated as the original prompt and a new prompt is built from it + feedback. reference_images: Optional list of absolute paths to image files to use as visual references for generation.
List all previously generated images. Args: working_directory: Directory to look for generated images.
Take a rough idea and return an optimized image generation prompt. Use this before generate_image to craft a better prompt. Returns the refined prompt without generating an image. Args: idea: A rough description of what you want (e.g. "banana logo for my app"). style: Optional visual style. One of: modern, minimal, professional, playful, dark, realistic, artistic, flat. size: Optional image size/aspect ratio. One of: square, portrait, landscape, wide, 4k, hd.
Missing enum constraints on style and size parameters. Both accept free-form strings instead of declaring valid values (modern|minimal|professional|playful|dark|realistic|artistic|flat and square|portrait|landscape|wide|4k|hd). LLMs will hallucinate invalid values instead of selecting from the documented set.
Output schemas not documented. All three tools return JSON strings (via json.dumps()) but the rubric requires documentation of what fields LLMs should expect. generate_image returns {status, file, absolute_path, prompt_used, model, usage}; refine_image_prompt returns {refined_prompt, available_styles, available_sizes}; list_generated_images returns {count, images}. Without schema docs, agents cannot reliably extract or chain results.
Insufficient error guidance. If working_directory is invalid, if a reference image file does not exist (caught in _read_image_file but error is generic), or if the LiteLLM API fails, no recovery guidance is provided. Errors should indicate: is this retryable? Should I ask the user? What should I try next?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
generate_image description includes example values ('e.g. "gemini/gemini-2.5-flash-image", "gpt-image-1"') which LLMs tend to reuse literally. Replace with formal constraint: 'A LiteLLM-compatible model identifier following the pattern provider/model-name.'
list_generated_images has a weak description (50 chars: 'List all previously generated images.'). It does not explain when to call it (e.g., to discover available images before passing to feedback), whether it paginates, or what 'images' field structure is.
generate_image accepts reference_images as a list of file paths but does not document the required format, accepted image types (jpeg, png, webp, etc.), or the hard limit of 10 references. The code enforces max 10, but agents are not told this constraint.