MCP server for Pixazo AI media generation API - images, videos, music, 3D models, and more
Pixazo MCP server provides 8 well-structured tools with explicit Zod schemas, descriptions, and tool annotations (readOnlyHint, destructiveHint). However, several definition quality gaps prevent a higher score: (1) Output schemas are NOT documented, the rubric requires 'Document the output schema' but the tools return generic 'content' objects with no typed schema specification; (2) Parameter descriptions lack critical constraint details (ranges, formats, patterns) that would prevent LLM hallucination; (3) Error handling is minimal, no recovery guidance, no actionable error messages documented; (4) Some tools return generic text responses instead of structured data; (5) Tools accept image/video URLs but lack validation guidance. The tools follow basic naming (verb_noun), have descriptions (100-200 chars, within baseline range), and explicit input schemas with enums. However, the lack of output schema documentation and minimal error handling guidance are significant gaps for production use. The model_enum list documents all available models clearly, which is a strength. Tool composition is clean (each tool has one responsibility), though some parameter interdependencies are not fully documented (e.g., when to use image_url vs video_url vs first_frame in generate_video).
Edit, upscale, remove background, relight, or transform images. 7 models: BRIA background removal, Crystal Upscaler, SeedVR, FireRed edit, relighting, face-to-sticker, inpainting.
Generate 3D models or NeRF from text or images. 5 models: Meshy Voxel, Meshy 3D, Rodin Gen 2 (text-based), Tripo 3D, Point-E.
Generate an image from a text prompt. 52 models available including Flux, Recraft, Ideogram, Stable Diffusion, GPT Image, Grok, Nano Banana, SDXL, and more.
Generate music or songs. 4 models: ACE Step (songs with lyrics), Lyria 2 (instrumental), ElevenLabs Music, Tracks.
Generate speech from text. 3 models: Google TTS, Azure TTS, ElevenLabs Premium.
Generate videos from text, images, or other videos. 42 models including Hailuo, Seedance, Kling (1.6/3.0/O1), Sora, Veo 3.1, Runway gen-4.5, WAN (2.2/2.5/2.6), LTX, Luma Ray 2, Mochi, VEED, Vidu, Higgsfield, and more.
Output schema not documented. Tools return generic { content: [{ type, text }] } responses with no typed schema specification for what the LLM should expect (URL format, metadata fields, error conditions).
Parameter constraint descriptions lack specificity. 'duration' in generate_music accepts a number but no min/max documented (e.g., '30-180 seconds'). 'resolution' in generate_video is described as 'e.g. 720P, 1080P' but no enum constraint or validation rules provided.
Error handling and recovery guidance absent. No documentation of what errors tools return, when they are retryable, or what the LLM should do if a generation fails (e.g., 'Rate limit exceeded. Retry after 60 seconds' or 'Invalid image URL. Check URL is accessible and in JPG/PNG format').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
List all available models by category (images, videos, music, 3D, TTS, try-on).
Virtual try-on for clothing and accessories. 3 models: Snapxai IDM-VTON, Xtreme IDM-VTON, ID-VTON.
Parameter interdependencies not documented. generate_video accepts image_url, video_url, first_frame, last_frame, and audio_url, but it is unclear which combinations are valid, which are mutually exclusive, and which model types require which parameters.
No input validation guidance documented. Tools accept URLs for images/videos but do not specify: format requirements (JPG, PNG, MP4?), size limits, encoding requirements, or what happens if a malformed URL is passed.
edit_image tool logic in code checks for specific model IDs to change 'image_url' to 'image' parameter, this is an API-specific quirk that should be encapsulated and not require LLM awareness. Suggests API contract is brittle.
No pagination or result limit guidance. While generate_image and similar tools are inherently single-result, list_models returns all available models with no documented limit. For future scale, pagination should be designed in.