MCP server for Z.AI image and video generation models (GLM-Image, CogView-4, CogVideoX-3, Vidu Q1, Vidu 2) via the Model Context Protocol
The server provides 10 well-intentioned tools with mostly complete schemas and descriptions. All tools have names starting with action verbs (list_, generate_, get_, download_) and descriptions present. However, several high-impact issues reduce the score: (1) Three tools have compound names with 'and' (generate_and_download, generate_and_download_video), violating single-responsibility principle. (2) Output schemas are not formally documented, responses are formatted as text strings via helper functions but the actual structure is not visible in parameter definitions. (3) Error handling exists but lacks recovery guidance and categorization. (4) Parameter descriptions vary in quality; some are thorough (e.g., generate_image) while others are terse (e.g., list_models has empty schema). (5) No pagination support for list tools despite potential for large result sets. (6) Tool composition issues: generate_and_download combines two concerns (generation + download) that should be separate, forcing the agent to do both even when only one is needed.
Download an image from a URL and return it as base64 or save to a file. Use this after generating an image to get the actual image data. Note: Z.AI image URLs expire after 30 days.
Generate an image and immediately download it in one operation. Combines generate_image and download_image for convenience.
Generate a video and immediately download it in one operation. Combines generate_video and download_video for convenience.
Generate an image synchronously from a text prompt. Returns the image URL directly. Use this for most image generation tasks.
Start an asynchronous image generation task. Returns a task ID to poll for results. Use this for long-running generations or when you need to process multiple images.
Generate a video asynchronously. Returns a task ID to poll for results. Supports text-to-video, image-to-video, and start-end frame generation across multiple models.
Tool names contain 'and', signaling multiple responsibilities. 'generate_and_download' and 'generate_and_download_video' should be split into separate tools (e.g., generate_image + download_image already exist independently). LLMs cannot compose partial operations when the tools force both to run together.
List tools (list_models, list_video_models) have empty input schemas (no parameters) but do not declare pagination, limits, or sorting options. No documentation of output structure (are they returning plain text, JSON, markdown?). Formatters return text strings but LLMs cannot reliably parse unstructured output.
Output schemas are not documented. Helper functions (formatImageResponse, formatAsyncStartResponse, formatVideoResultResponse, etc.) return text strings, but the actual JSON structure, field names, and types are invisible in the tool definition. LLMs cannot plan downstream operations without knowing what fields to extract.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Retrieve the result of an asynchronous image generation task. Use the task ID from generate_image_async.
Retrieve the result of an asynchronous video generation task. Use the task ID from generate_video.
List available Z.AI image generation models and their capabilities
List available Z.AI video generation models and their capabilities
Error handling returns text via formatError() but does not categorize errors as retryable, user-fixable, or fatal. No recovery guidance (e.g., 'invalid user_id, must be 6-128 characters'). Stack traces or raw API errors likely leave LLM unable to self-correct.
No pagination support on list tools. If list_models or list_video_models return large result sets, they will bloat the response and waste tokens. The rubric baseline calls for limit/offset/cursor parameters and documented result limits.
Polling tools (get_async_result, get_video_result) lack clear guidance on polling strategy: recommended interval, max retries, timeout. LLMs may poll indefinitely or with too-long intervals, wasting latency.