MCP server for generating images using Stable Diffusion/OpenAI API
Single tool 'generate-image' has a reasonable schema and some description, but critical gaps in parameter descriptions, missing output schema documentation, and no error guidance. The tool name lacks the verb-first convention (should be 'generate_image' or better 'create_image_from_prompt'). Input schema is present and validates well, but the description is sparse and doesn't explain when to use this tool vs alternatives, prerequisites (API key availability), or downstream behavior. No output schema is documented, LLMs cannot predict what fields the response contains or how to chain subsequent calls. Error handling is basic (generic 500-level responses without recovery hints). Security: API key is presumably injected via environment (good), but no explicit documentation of this pattern. The tool is a single-purpose write operation, which is good composition, but lacks idempotency guarantees or dry-run/confirmation pattern for destructive file I/O.
Generate an image using Stable Diffusion based on a text prompt
Tool name does not follow verb_noun convention; 'generate-image' uses hyphen and lacks clear verb prefix. Should be 'generate_image_from_prompt' or 'create_image' to signal the action and disambiguate from similar tools.
Tool description (52 chars) is too sparse. Does not explain WHEN to use this tool, what prerequisites exist (Stable Diffusion API key, internet access), or what downstream fields the response contains. Baseline for good descriptions is 50-200 chars with context for LLM selection.
Parameter 'destination' description says 'optional, saves to default folder if not provided' but does not specify what 'default folder' is, where it resolves to (user home? /tmp? $DEFAULT_DOWNLOAD_PATH env var?), or what happens if the path is invalid or already exists. Code shows fallback to Downloads but description does not document this.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | 2024-11-05+ | v1 |
No documented output schema. LLM cannot predict response structure. Code returns a text message ('Image generated and saved to: <path>'), but no ToolContent schema is documented. Pattern baseline: 100% of A+ tools document return types.
Error responses are generic ('Error generating image: <err>', 'Error saving image: <err>') without recovery guidance. Baseline pattern: error responses must tell the LLM what to do next. E.g., 'API rate limit exceeded; retry in 60s' vs bare 'Error generating image'.
Parameters 'width' and 'height' have type 'number' but no min/max constraints documented. Code defaults to 1920x1080 but schema says default is 1792x1024, inconsistency. Baseline: numeric parameters must specify min/max and constraints must match implementation.
No idempotency guarantee or dry-run pattern. Tool writes to filesystem with unique filename generation, but if filename generation fails mid-retry, behavior is undefined. Baseline pattern: non-idempotent tools should document retry semantics or offer confirmation steps.