MCP server for GPT Image 2 with selectable OpenAI API and ChatGPT web backends.
This server has three explicitly registered tools with schemas and descriptions. Tool naming follows verb_noun convention (generate_image, backend_status, browser_visibility), which is correct. However, parameter descriptions vary in quality and completeness. The generate_image tool has comprehensive schema documentation with proper constraints (min/max, enums), but lacks clear guidance on behavior (e.g., what happens on fallback from api to chatgpt-web). Error handling returns basic error text but lacks recovery guidance. Output schema is documented via structuredContent in code, but this is implicit rather than explicitly declared in the tool definition. The server does not implement tool annotations (readOnlyHint, destructiveHint) despite having tools with clear READ-ONLY (backend_status) and WRITE (generate_image, browser_visibility) semantics.
Show configuration and readiness for the API and ChatGPT web backends.
Show, hide, toggle, or inspect the ChatGPT web backend browser window.
Generate images using either the OpenAI gpt-image-2 API backend or the ChatGPT web automation backend.
Missing tool annotations (readOnlyHint, destructiveHint). backend_status is clearly READ-ONLY; generate_image and browser_visibility are destructive/write operations that should be annotated for safety.
Error responses lack recovery guidance. When image generation fails, the error message is raw (error.message or String(error)). Should return categorized errors with actionable next steps, e.g., 'API key missing. Check GPT_IMAGE_API_KEY env var and retry.'
Output schema not formally documented in tool registration. The code uses structuredContent to return JSON, but the tool definition does not declare the response schema. Agents cannot predict what fields will be present (status, backend, fallback_from, etc.).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2025-06-18+ | v2 |
generate_image fallback behavior undocumented. When backend='auto' or a selected backend fails, the tool may silently fall back to another backend. The response includes fallback_from, but the tool description does not explain this behavior or how to prevent it.
Parameter 'size' accepts arbitrary string values with no enum constraint. Description says 'API mode accepts OpenAI-supported sizes or auto' but does not list valid options. This invites LLM to hallucinate invalid sizes.
No idempotency guarantee documented. generate_image with identical prompt may produce different images. If an agent retries a failed call, it gets a different result. Should document whether calls are idempotent or require deduplication.
browser_visibility start_browser parameter interaction not documented. When action='show' and start_browser=false but no browser is running, what happens? Does it fail, return a status, or silently do nothing?