AI-powered image generation MCP server with multi-model support for Google Gemini and Imagen
pixelforge-mcp has 14 tools with generally clear verb-noun naming and comprehensive parameter schemas. However, descriptions lack depth for production use, parameter documentation is inconsistent, and critical output schema documentation is missing. The server defines tools with proper JSON Schema types and reasonable defaults, but lacks actionable error guidance, recovery hints, and field-level descriptions that would enable LLMs to compose tools effectively. Tool definitions are directly visible in server.py (not inferred), so no artificial capping applies.
Advanced image editing with inpainting, outpainting, and object replacement.
Analyze an image and return a detailed description including objects, colors, composition, and mood.
Apply a prompt template with variable substitution.
Compare two images and return visual differences and similarities.
Detect objects of a specific type in an image and return their locations and confidence scores.
Edit an existing image with text-guided modifications using Google Gemini.
Output schemas completely missing. All 14 tools describe inputs thoroughly but nowhere does the code or docs define what fields, types, or structure are returned. LLMs cannot plan downstream tool calls or extract required fields (e.g., image paths for chaining tools) without knowing response structure.
transform_image bundles 8+ distinct operations (crop, resize, rotate, flip, blur, sharpen, grayscale, watermark) into one tool with 18 parameters. Per pattern:tool, each operation should be a separate tool so agents can compose them independently. This violates single responsibility and makes the tool hard for LLMs to reason about.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Estimate the cost of a generation or editing operation based on model and operation type.
Extract all visible text from an image using OCR.
Generate an image from a text prompt using Google Gemini.
List available prompt templates, optionally filtered by category.
Enhance an image generation prompt to produce better, more detailed images.
Remove the background from an image and return a transparent PNG.
Apply local image transformations (crop, resize, rotate, flip, blur, sharpen, grayscale, watermark).
Upscale an image to higher resolution (up to 4x) using AI-powered enhancement.
detect_objects parameter 'target_objects' is a free-form comma-separated string with no enum or schema. This invites LLM hallucination of invalid object types. Should either enumerate valid objects or reference a canonical detection ontology.
apply_template 'variables' parameter is an untyped object. LLM has no way to discover what variables a template requires or their types. Should document expected structure or require variables to match template schema.
No error recovery guidance anywhere. Tools lack descriptions of failure modes, retryable vs. fatal errors, or what to do next (e.g., 'If image file not found, check path format or call list_templates to see available templates'). Agents are left guessing.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present in code or visible config. Tools that write files (generate_image, edit_image, etc.) should be marked destructive; read-only tools should be marked. This enables safe agent planning.
Descriptions are inconsistently detailed. Some tools (generate_image) have rich docstrings; others (extract_text, detect_objects) have minimal descriptions (60-65 chars). Most descriptions fall in the 'too brief' band (50-150 chars), lacking context on when to use the tool, what it returns, and failure modes.
No guidance on tool composition or dependencies. For example, list_templates and apply_template are related but the descriptions don't say so. Agents don't know that they should call list_templates first to discover valid template names, leading to wasted list_templates calls or apply_template failures.
No rate limiting or cost warnings in descriptions. Tools that call expensive Gemini APIs (generate_image with 4K, thinking_budget) lack guidance on cost, tokens consumed, or rate limit behavior. Agents might invoke these tools in loops without understanding the cost impact.