A multimodal MCP server for product photo editing and video generation using Google's Generative AI APIs, built with FastMCP and Google ADK
Two tools with partially documented schemas and moderately detailed descriptions, but critical gaps in parameter type definitions, output schema documentation, and error handling. Tool naming is action-verb compliant but descriptions, while lengthy, lack clarity on error conditions and recovery paths. Parameter types are inferred from docstring context rather than formally declared. No explicit output schema documentation visible. Error handling returns dict structures but lacks actionable guidance for LLM recovery.
Modify an existing product photo or combine multiple product photos. This tool lets you make changes to product photos. You can: - Edit a single photo (change background, lighting, colors, etc.) - Combine multiple products into one photo (arrange them side by side, create bundles, etc.) IMPORTANT: - Make ONE type of change per tool call (background OR lighting OR props OR arrangement) - For complex edits, chain multiple tool calls together - BE AS DETAILED AS POSSIBLE in the change_description for best results!
Generates a professional product marketing video from text prompt and starting image using Google's Veo API. This function uses an image as the first frame of the generated video and automatically enriches your prompt with professional video production quality guidelines to create high-quality marketing assets suitable for commercial use. AUTOMATIC ENHANCEMENTS APPLIED: - 4K cinematic quality with professional color grading - Smooth, stabilized camera movements - Professional studio lighting setup - Shallow depth of field for product focus - Commercial-grade production quality - Marketing-focused visual style
Input parameter types not formally declared in schema. Parameters 'prompt', 'image_data', 'negative_prompt' for generate_video_with_image and 'change_description', 'image_artifact_ids' for edit_product_asset lack formal type definitions visible in code, inferred from docstring only.
Output schemas not documented. Both tools return dict with keys like 'status', 'tool_response_artifact_id', 'message', but LLMs receive no formal schema definition of expected structure, types, or field semantics. Responses are documented in code comments but not in schema form.
Error messages lack actionable recovery guidance. Both tools return error dicts with 'message' field describing the failure (e.g., 'Artifact {img_id} not found', 'No images provided') but do not suggest next steps (e.g., 'Call list_artifacts() to see available image IDs').
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Parameter 'image_data' for generate_video_with_image is described as 'Base64-encoded image data' but lacks format constraints (max length, character restrictions, base64 validation guidance). 'image_artifact_ids' for edit_product_asset has default=[] but no description of what an empty array means or whether this causes silent no-op behavior.
No permission gates or audit logging visible. Both tools call Google APIs (Veo, Gemini) with no explicit user/agent permission checks. No audit trail (who called what, when, with which params) is logged for compliance or security review.
Tool descriptions include prescriptive advice ('MAKE ONE TYPE OF CHANGE PER TOOL CALL', 'BE AS DETAILED AS POSSIBLE') but lack clarity on what LLM should do if advice is violated. No error condition explains what happens if user chains contradictory edits or provides ambiguous prompts.