MCP server for the Reelmotion platform that enables image, video, and speech generation through a Gemini chatbot interface with cost tracking, workflow state management, and moderation.
The server defines 4 tools with reasonable parameter schemas and descriptions. Tool naming follows verb_noun convention (generate_image, generate_video, generate_speech, craft_prompt). Descriptions are detailed and contextual (194 - 500+ chars, well above baseline avg 194 chars), providing cost information, async behavior, and error codes. However, there are significant gaps: (1) Output schemas are NOT documented, no explicit return type specifications for any tool; (2) Error handling descriptions exist but lack recovery guidance ('still processing (202)', 'failed job (422)', 'insufficient balance (402)') without telling the LLM what to do next; (3) Parameter descriptions lack formal constraints (e.g., aspect_ratio and quality accept free-form strings with only examples in descriptions, not enums); (4) Security: reference_image and reference_images parameters accept URLs including data: and blob: URIs, but no validation is documented; (5) Composition gaps: no batch variants offered; craft_prompt is a utility that refines other tool inputs but its relationship to the generation tools is not explicit. The tool definitions themselves are clear and verb-driven, with consistent parameter typing (strings, integers, arrays, numbers), but the lack of documented output schemas and formal input constraints significantly limits LLM safety and composability.
Refine and improve a generation prompt using the Gemini model to enhance clarity, detail, and generation quality. Returns an improved prompt text.
Generate or edit an image using the reelmotion backend. Supports text-to-image (type 1), image-to-image (type 2), and multi-image reference (type 3). COST: Seedream = 4, GPT = 6, Nano Banana 2 = 7, Midjourney = 9 tokens per image. Seedream and Midjourney deliver asynchronously (hybrid): the call may return the finished image (200) or, for slow jobs, 'still processing' (202) — in which case the user is notified when it's ready and we never retry. A failed job (422) is auto-refunded by the backend; insufficient balance is 402.
Generate speech/voiceover using ElevenLabs text-to-speech. Supports 100+ voices with configurable stability, similarity, and style parameters. Cost: 11 tokens per 1000 characters (6 tokens on the flash voice model).
Generate a video using the reelmotion backend. Supports various video models with configurable duration, resolution, aspect ratio, and optional motion/reference parameters. Videos are generated asynchronously; successful submissions return a processing marker (202) and the user is notified via push/realtime when ready. Failed jobs are auto-refunded by the backend.
Output schemas completely undocumented for all 4 tools. No return type specifications, field listings, or examples provided. LLMs cannot reliably extract data or plan downstream tool calls without knowing response structure.
Input validation constraints not formalized as enums or JSON Schema format/pattern. quality (e.g., '2K'), aspect_ratio (e.g., '16:9'), model (e.g., 'Seedream', 'Veo 3.1 Ultra'), resolution (e.g., '720p', '1080p'), and voice_id lack closed-set definitions. LLMs may hallucinate invalid values; descriptions alone are insufficient.
Error handling responses (202 'still processing', 422 'failed job', 402 'insufficient balance') are mentioned in descriptions but do not include recovery guidance for the LLM. No actionable next steps like 'retry after 30 seconds', 'inform user', or 'call refund_operation'.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | - | v1 |
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). craft_prompt is marked READ_ONLY in the risk field, but tool definitions lack annotations. generate_image and generate_video are WRITE operations but lack destructiveHint classification; generate_speech is WRITE (modifies external state) but unmarked. Protocol Readiness deduction applies.
Numeric parameters lack range constraints. generate_video duration has no stated bounds (1 - 300? 1 - 3600?). quantity and image_type in generate_image lack min/max. stability, similarity_boost, style in generate_speech are stated as 0.0 - 1.0 in descriptions but not enforced as numeric schema constraints.
No discovery/enumeration tool for ElevenLabs voices. generate_speech requires a voice_id parameter but offers no way for LLMs or users to discover valid voice IDs. Forced to require user input or hardcode IDs, poor composability.
Reference image handling (reference_image, reference_images in generate_image) accepts blob: and data: URIs per code comments, but no validation/sanitization is documented. potential for malformed/malicious input.
craft_prompt is a utility that refines inputs to generate_image/generate_video but the relationship is implicit. No tool composition guidance for LLMs (e.g., 'call craft_prompt before generate_image to improve quality').