MCP server for AI image & video generation, editing, and region repair. Powered by Gemini and OpenAI.
Pixel Surgeon MCP has well-structured tool definitions with good naming conventions and parameter schemas, but falls short of production quality in critical areas. All 9 tools follow verb_noun naming patterns (generate_image, edit_image, list_generated_images, etc.), and each has a description. However, parameter descriptions are generic and inconsistent, output schemas are not documented, and error handling lacks recovery guidance. The schema score is dragged down by missing documentation of return types and incomplete per-tool validation rules.
Edit an existing image with a text prompt describing the edits to apply
Generate an image from a text prompt using specified model and parameters
Generate a video from a text prompt using Google's Veo model
List all available image generation and editing models
Interactively select a region of an image in the viewer and apply edits to that region
List all previously generated images stored in the image gallery
List all previously generated videos stored in the video gallery
Output schemas are completely undocumented. No tool explicitly declares what it returns. LLMs cannot plan downstream calls or extract the correct fields without this information.
Parameter descriptions lack constraint details. E.g., 'image_size' accepts '512, 1K, 2K, 4K' but the description does not specify these are the ONLY valid values. LLMs cannot disambiguate between example values and constraints without explicit enums or constraint language.
No error handling or recovery guidance in tool descriptions. If an image generation fails, the LLM has no way to know whether to retry, ask the user for a revised prompt, or fall back to a different model. Descriptions should state 'On failure, try adjusting the prompt or selecting a different model.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
Repair or inpaint a region of an image by selecting an area and providing a repair prompt
Open the interactive image and video gallery viewer in a web browser
No documentation of pagination for list_* tools. If the image or video gallery contains hundreds of items, how does the LLM iterate through results? No limit/offset parameters are visible, and no total count is documented in the schema.
Some tool names are ambiguous about action. 'interactive_crop' does not clearly convey that edits are applied to the cropped region. A name like 'edit_cropped_region' or 'apply_edits_to_selection' would be clearer about the action taken.
'view_gallery' description states it opens the gallery 'in a web browser' but does not document what happens if no browser is available or fails to open. On headless systems, this tool will silently fail with no user feedback.
Parameter 'mask_url' in repair_image is described as 'optional' but the schema does not reflect this (no 'required' field limiting its enforceability). LLMs cannot distinguish optional from required parameters.