MCP server for image and video generation via OpenAI GPT Image, OpenAI Sora, and Google Vertex AI video models
This MCP server provides 13 tools for media generation (images, videos) via OpenAI and Google APIs. While all tools have descriptions and schemas are present, there are significant gaps in parameter descriptions, inconsistent schema rigor, and missing output documentation. The naming is action-oriented (openai-images-generate, google-videos-generate) which is good, but many parameter descriptions are minimal (under 20 chars in some cases) or lack detail about constraints, defaults, and dependencies. No error handling guidance is provided to help LLMs recover from failures. The server shows mid-tier quality typical of community-contributed tools, not production-grade.
Generate a video from a text prompt using Google Vertex AI video generation API
Download the video content of a completed Google video generation operation
Retrieve the status and metadata of a Google video generation operation
Edit or extend an existing image using OpenAI's image editing capabilities
Generate images using OpenAI's DALL-E 3 model (gpt-image-1.5)
Generate variations of an existing image using OpenAI's image variation API
Parameter descriptions lack specificity and constraints. 'The text prompt describing the image to generate' provides no guidance on length, content restrictions, or format. Parameters like 'n' lack min/max bounds. Descriptions should state: range (1-10), default values, format requirements, and examples of valid values.
No output schema documentation. Tools return results but do not document what fields are in the response (e.g., does openai-images-generate return { url, revised_prompt, data }?). LLMs cannot plan downstream calls without knowing response structure.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Generate a video from a text prompt using OpenAI's Sora model
Delete a generated video from OpenAI's Sora API
List all generated videos from OpenAI's Sora API
Generate a video remix from an image and prompt using OpenAI's Sora model
Retrieve details about a specific generated video from OpenAI's Sora API
Download the video content of a generated video from OpenAI's Sora API
Generate test images using sample images from the test sample directory
No error handling guidance. Tools do not document recovery paths. When openai-images-generate fails (quota exceeded, invalid prompt, API error), the LLM has no guidance on whether to retry, call a different tool, or ask the user. Error responses should state: is this retryable? What should I do next?
Ambiguous 'response_format' parameter semantics. All image/video generation tools accept response_format enum ['url', 'b64_json'] and tool_result enum ['resource_link', 'image']. The relationship between these two parameters is undocumented. Does response_format=url + tool_result=image work? Does b64_json require image and resource_link require url? Dependencies force LLMs to guess.
test-images tool lacks a clear purpose. Description is minimal ('Generate test images using sample images from the test sample directory') with no guidance on when to call it vs. other generation tools. Required parameters (tool_result, response_format) suggest it's a utility but the naming does not start with a clear action verb.
Parameter 'file' (output path) is optional on all tools but lacks guidance on: default location if omitted, allowed directories (security boundary), file naming conventions, overwrite behavior. The env vars MEDIA_GEN_DIRS exist in code but are not documented in tool descriptions, leaving agents blind to filesystem constraints.
openai-videos-list and openai-videos-retrieve do not document pagination or filtering. openai-videos-list accepts 'limit' but no offset/cursor. What is the default limit? How many videos exist total? How does an agent iterate through all videos?
Destructive tool openai-videos-delete lacks confirmation or dry-run mode. Agents can irreversibly delete videos. No pattern for 'confirm_before_execute' or way to preview what will be deleted. Risky for multi-step workflows.
Composition issue: openai-videos-create and openai-videos-remix both generate videos but serve different workflows (text→video vs image→video). No guidance on when to call which. Descriptions do not clarify the distinction for agent planning.