MCP server for Grok image generation and editing via xAI API
The server defines two tools with complete input schemas using Zod validation and clear descriptions. Both tools follow verb_noun naming conventions (generate_image, edit_image). Descriptions are substantive (150-200 chars) and explain what each tool does, prerequisites, and output format. All parameters have type definitions and descriptions. However, there are notable gaps: no output schema documentation, no error handling guidance for LLMs, no tool annotations (destructiveHint/readOnlyHint), and no structured response documentation. The tools are well-composed (single responsibility each), but the server lacks patterns for guiding recovery from failures.
Edit existing images using a text prompt. Provide 1-3 source images (as URLs, base64 data URIs, or local file paths) along with editing instructions. Returns edited image URLs in markdown format.
Generate images from text prompts using the Grok image model. Returns image URLs in markdown format. Note: URLs are temporary, download or process promptly.
No output schema documentation. Tools return markdown-formatted image URLs and revised prompts, but LLMs have no formal schema for the response structure. This forces LLMs to infer the response shape and risks misinterpretation of multi-image results.
Error responses lack recovery guidance. When API calls fail (timeout, rate limit, API error), responses return raw error messages without suggesting next steps (e.g., 'retry with smaller n value' or 'check API key'). Per pattern:recovery-guide, errors should tell LLMs what to do next.
No tool annotations. Neither tool declares readOnlyHint, destructiveHint, or idempotentHint. 'generate_image' and 'edit_image' modify external state (create images at xAI), so destructiveHint would clarify that these are non-idempotent write operations to LLMs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Parameter 'image_urls' in edit_image accepts local file paths, base64 URIs, and HTTP(S) URLs but provides no guidance on which formats are preferred or when to use each. LLMs may unnecessarily convert URLs to base64, wasting tokens and inviting encoding errors.
The 'n' parameter description ('Number of images to generate (1-10, default 1)') does not mention cost implications or rate limits. Agents may not realize that n=10 is expensive or could trigger API throttling, leading to wasteful or failed calls.
No documented behavior for temporary URL expiration. The description notes 'URLs are temporary, download or process promptly' but does not specify TTL or what happens when a URL expires. LLMs cannot plan for stale URLs without this information.