LEAP MCP has 6 tools with basic definitions but significant quality gaps. All tools have names starting with action verbs (create_, execute_, generate_, validate_, get_, list_), which is positive. However, tool descriptions are minimal (averaging ~100 chars), parameter descriptions are sparse or missing, input schemas lack validation constraints, and output schemas are not documented. The server uses fastmcp framework on STDIO transport, which limits remote accessibility. Tool compositions show reasonable separation of concerns (generation vs execution vs validation), but lack the chaining metadata (IDs, references) needed for multi-step agent workflows. No error handling guidance, no output schema documentation, and no tool annotations (readOnlyHint/destructiveHint) despite clear risk differences between tools.
Missing output schemas for all tools. Users and LLMs cannot predict the structure of responses, preventing proper downstream tool composition and error handling.
Input schemas lack validation constraints. 'topic' and 'scene_code' parameters accept free-form strings with no length limits, format validation, or enumeration of valid values. This invites invalid inputs and hallucinated content from LLMs.
No parameter descriptions for key inputs. 'scene_code' in execute_manim_scene and validate_manim_code lacks guidance on expected format, Python version constraints, or Manim library version compatibility.
execute_manim_scenevalidate_manim_code
Recommendations
Document output schemas for all tools. Define the structure returned by each tool (e.g., create_educational_video returns {video_path: string, duration: number, voice: string, created_at: string}). Use JSON Schema format and include in fastmcp tool definitions.
Add minimum/maximum constraints and enums to input schemas. Constrain 'topic' length (e.g., 1 - 200 chars), validate 'voice' as an enum (alloy|echo|fable|nova|onyx|shimmer), and require 'scene_code' to be valid Python with Manim imports.
Expand tool descriptions to 100 - 200 characters. Include WHAT the tool does, WHEN to use it, and any prerequisites. Example: 'Validates Manim Python code for syntax errors, import issues, and animation rendering compatibility. Call this before execute_manim_scene to catch errors early. Requires Python 3.8 - 3.11.'
Add parameter descriptions to all inputs. Explain what each parameter does, expected format, and constraints. Example for 'scene_code': 'Complete, executable Manim scene code (Python 3.8 - 3.11). Must include valid Scene subclass. Code is executed in a temporary sandbox.'
Implement error handling with recovery guidance. For execute_manim_scene, return structured errors: 'SyntaxError in scene code (line 42): missing colon. Call validate_manim_code to identify all errors before retry.' For missing API keys: 'OpenAI API key not configured. Add OPENAI_API_KEY to .env and restart the server.'
Add tool annotations (readOnlyHint/destructiveHint/idempotentHint) to fastmcp tool definitions. Mark execute_manim_scene and create_educational_video as destructiveHint=true. Mark validate_manim_code and get_scene_templates as readOnlyHint=true.
No error handling guidance. Tools offer no recovery instructions for common failures (missing API key, Manim installation failure, syntax errors in generated code, timeout from long-running animations).
Tool annotations missing. execute_manim_scene and create_educational_video are destructive (write files, execute code) but lack destructiveHint annotation. get_scene_templates and list_generated_videos are read-only but lack readOnlyHint. Idempotency not declared.
No pagination for list_generated_videos. If many videos are generated, the response could be large and blow context. No limit, offset, or cursor parameters.
Tool descriptions are too brief (55 - 70 chars average). 'Lists all previously generated videos with metadata' and 'Validates Manim scene code for syntax errors and common issues' lack context on when to use, prerequisites (API key setup), and what the response contains.
Chaining metadata incomplete. create_educational_video likely returns a video file path or ID, but this is not documented. Agents cannot know what to do next (e.g., retrieve the video, list it, share it).
Add pagination to list_generated_videos with limit (default=20, max=100) and offset/cursor. Return total_count and next_cursor in response. This prevents context window exhaustion when many videos exist.
Document response field names and include chaining IDs. If create_educational_video returns a video_path, document it. If possible, also return a unique video_id and timestamp so agents can retrieve, delete, or reference the video in subsequent calls.
Add a dry-run or confirm-before-execute pattern to destructive tools (create_educational_video, execute_manim_scene). Return {plan: '...', confirm_token: '...'} on first call; agent must call with confirm_token to execute.
Validate all user inputs on the server side (topic length, voice enum, scene_code syntax) and return clear, actionable error messages. Example: 'Invalid voice: must be one of: alloy, echo, fable, nova, onyx, shimmer. Got: "alto".'
Document the timeout and resource constraints. The code shows MANIM_TIMEOUT=300 and MANIM_QUALITY='medium', expose these as tool behavior in descriptions so agents understand why some calls may fail or take time.