Unified MCP server for OpenAI multimodal APIs (Sora, Whisper, GPT Vision)
Sanzaru provides 22 tools across video, audio, and image domains with explicit schema definitions visible in server.py. Tool names follow verb-noun conventions (create_video, list_audio_files, transcribe_audio) and most descriptions are present and non-trivial (avg ~120 chars). However, significant gaps emerge: (1) Many parameter descriptions are sparse or missing detail about constraints (e.g., 'seconds' in create_video lacks range info; 'offset' and 'chunk_size' in _get_media_data have minimal guidance). (2) Output schemas are NOT documented anywhere in the visible source, no specification of what fields are returned or their structure. (3) Several tools mix concerns or have unclear responsibilities (e.g., create_image vs generate_image both do image generation; _get_media_data is marked 'do not call directly' but remains exposed). (4) Error handling is minimal, no evidence of recovery guidance or actionable error messages. (5) Tool definitions reference external services (OpenAI Sora, Whisper, DALL-E, ElevenLabs) but do not document failure modes, timeouts, or prerequisites. The server demonstrates competent naming and basic schema presence, placing it solidly in the C+ range rather than B-.
Internal tool used by the MCP App media viewer to fetch base64-encoded chunks of media data. Do not call directly — use view_media instead.
Chat with GPT-4o using audio input
Compress audio file to reduce file size
Convert audio between formats
Generate images using OpenAI DALL-E or image generation API
Generate video using OpenAI Sora API
Delete a video from OpenAI storage
Output schemas not documented. None of the 22 tools specify what fields they return or the structure of responses. This forces LLMs to guess at response format and prevents proper chaining between tools.
_get_media_data is marked 'do not call directly' but is still exposed as a callable tool. This violates single-responsibility principle and creates confusion. Internal utility tools should not be registered.
Duplicate/overlapping image generation tools: both create_image and generate_image perform image generation with different parameters. LLMs cannot disambiguate. Consolidate to one canonical tool or clearly document the distinction.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 63 | - | v1 |
Download a generated image to local storage
Download a generated video to local storage
Edit an existing image using OpenAI API
Generate images using GPT-4o vision with optional text-to-image generation
Check status of an image generation job
Check the status of a video generation job
List audio files in storage
List locally stored videos
List reference images for image operations
List generated videos from OpenAI
Prepare a reference image for use in image generation or editing
Generate a new video based on an existing one with a new prompt
Convert text to speech using OpenAI or ElevenLabs API
Transcribe audio using OpenAI Whisper API
Interactive player for sanzaru's video, audio and image files.
Missing parameter constraints and ranges. create_video's 'seconds' parameter lacks min/max bounds (e.g., 5-120?). list_* tools' 'limit' parameters lack stated bounds. Unbounded parameters invite absurd LLM inputs.
No error handling or recovery guidance. Tools that call external APIs (OpenAI Sora, Whisper, DALL-E) do not document failure modes, timeouts, rate limits, or what the LLM should do on failure.
Sparse parameter descriptions. Many parameters have trivial descriptions that add little value. E.g., 'filename' appears in 15 tools but nowhere does it explain accepted formats, path traversal rules, or length limits.
No permission checks documented. delete_video is destructive but no description of permissions, confirmation requirements, or audit trails.