AI video production pipeline: generate storyboards from scripts and batch-produce illustration images using Gemini AI. Manages full project lifecycle from script input to downloadable image assets.
The server provides 7 tools with basic descriptions and input schemas, but falls significantly short of production-quality standards. Major issues: (1) Parameter descriptions are minimal or absent, most parameters lack context for LLM reasoning. (2) Output schemas are completely undocumented, no structured return types specified anywhere. (3) Descriptions are too brief (10-30 chars for most tools) and lack actionable guidance on WHEN/WHY to use them. (4) No error handling guidance, agents have no recovery path on failures. (5) Chinese descriptions mixed with English parameter names create ambiguity. (6) Several parameters (e.g., 'shots' in edit_storyboard) have vague item types {'type': 'object'} with no schema detail. (7) No idempotency guarantees or rate-limiting guidance for concurrent image generation. Per the rubric baseline, average description should be 194 chars; these are 20-80 chars. Average tool should have 4-8 well-described parameters; most have 1-2 with minimal context.
No output schemas documented for ANY tool. Agents cannot know what fields to expect, forcing them to guess or waste tokens in follow-up discovery calls.
Parameter descriptions are minimal (20-35 chars for most) and lack LLM-actionable guidance. Descriptions like 'Project identifier (required)' are too vague, no format guidance, no context on when parameter is needed.
edit_storyboard 'shots' parameter uses bare {'type': 'object'} with no properties schema. LLM must reverse-engineer structure from prose description, violating schema formality rule.
edit_storyboard
Recommendations
For each tool, write a 100-200 char description that answers: What does it do? When should the LLM call it (after which prior steps)? What does it return? Example: 'create_storyboard: Analyzes a script text and breaks it into a shot-by-shot storyboard with timings and visual descriptions. Call after the user provides a script. Returns a project object with shots array and unique project_id for use in subsequent tools.'
Document output schemas for all tools. Example for list_projects: '{ "projects": [{ "project_id": string, "name": string, "status": string, "shots_count": integer, "created_at": ISO8601_datetime }], "total": integer }'
Expand project_id parameter descriptions: 'project_id: Unique identifier for the video project (returned by create_storyboard). Format: alphanumeric string, typically UUID or slug.'
Free-form string parameters (style in create_storyboard, style in generate_images) lack enum constraints. Descriptions list valid values ('AI科技/知识分享', 'default|tech|knowledge') as prose, not as JSON Schema enums. LLMs frequently hallucinate unlisted values.
No error handling guidance. Tools provide no recovery paths (e.g., 'if image generation fails, try generate_images with lower concurrency'). Per pattern, error responses must tell LLM what to do next.
Tool descriptions are 20-38 chars, far below 50-200 char LLM-optimized range. Missing context: WHEN to call (after script input? before editing?), WHY (vs similar tools), WHAT is returned (structure, format).
No pagination or result-limiting guidance. list_projects may return unbounded results. Per pattern, tools returning lists should accept limit/offset and document caps (recommend 20-50 items). Large result sets bloat context and degrade LLM reasoning.
generate_images is stateful (async image generation) but no polling/async semantics documented. Agents don't know: Does the call block until images are ready? Return a job ID? Should get_image_status be polled? This ambiguity will cause agent failures.
Mixed language (Chinese tool descriptions in mcp.json, English in server.py) creates ambiguity. LLMs may see different descriptions in different contexts. Descriptions should be consistent and in the agent's language.
all
Add error handling guidance to async/long-running tools. For generate_images: 'If generation exceeds timeout, call get_image_status to check progress. Partial results are returned if some images succeed and others fail, retry with remaining shots or lower concurrency.'
Add pagination to list_projects: 'Returns max 20 projects per call. Use optional limit (1-100) and offset parameters to retrieve more. Response includes total project count.'
Document idempotency and rate-limiting behavior. For generate_images: 'Calling with the same project_id and concurrency twice will not duplicate images, the tool skips shots that are already generated (idempotent). Concurrent calls are rate-limited to 3 simultaneous batch operations across all projects.'
Standardize language to English for all tool and parameter descriptions to prevent LLM confusion. Chinese descriptions in mcp.json should be translated to English in the inputSchema descriptions.
Add 'required' fields explicitly in tool input schemas. Example for create_storyboard, make it clear that script_text is required but style and duration have defaults.
Add 'createdAt' timestamps to all resource responses (projects, image statuses), LLMs use dates to infer recency and make sequencing decisions.