Three tools with well-structured schemas and generally good descriptions. All tools use proper JSON Schema with typed parameters and enum constraints. However, descriptions lack depth on use cases and recovery guidance. The schemas are complete and visible in the source code. Parameter descriptions are adequate but could be more instructive about valid ranges and relationships.
Tools (3)
describe_imageread onlyauthsource verified77/100
Analyze and describe one or more images using Google Gemini. Returns a text description of the image contents.
edit_imagewriteauthsource verified80/100
Edit one or more images using Google Gemini. Provide images and instructions for how to modify them. Returns a base64-encoded image.
generate_imagewriteauthsource verified82/100
Generate an image using Google Gemini. Optionally provide reference images to guide the generation style or content. Returns a base64-encoded image.
Tool descriptions lack recovery guidance and error context. Descriptions do not explain what happens on API failure, rate limiting, invalid images, or how to retry. Agents receive no actionable recovery paths.
No documented output schema for any tool. Tool descriptions say 'Returns a base64-encoded image' or 'Returns a text description' but do not describe the actual response object structure, fields, or types. LLMs cannot plan downstream tool calls or extract the right data.
generate_image and edit_image accept 'images' as optional/required reference input but descriptions do not explain limits (14 total, 10 object + 4 person refs). Parameter description states the limit but lacks clarity on when and how to use reference images.
generate_imageedit_image
Recommendations
Add explicit output schema documentation to all three tools. Example for generate_image: 'Returns a JSON object: { image_data: string (base64), file_path?: string (if outputPath provided), model_used: string }'
Enhance tool descriptions with recovery guidance. Example: 'If the API returns rate_limit_error, retry after 30 seconds. If invalid_api_key, check your GEMINI_API_KEY environment variable. If image_rejected, the content may violate usage policies, try a different prompt.'
Document error conditions explicitly for each tool. Add a section in each description: 'Errors: Returns 400 if images are malformed or too large (>20MB). Returns 429 if rate limit exceeded, wait before retrying. Returns 403 if personGeneration=ALLOW_ADULT is not available in your region.'
Clarify the purpose and use case for thinkingConfig. Example: 'For complex scenes with multiple objects or precise text rendering, set thinkingLevel to HIGH to enable deeper reasoning. This may increase latency by 10-20 seconds but improves accuracy. Omit or set to MINIMAL for fast iterations.'
Explain useGoogleSearch behavior: 'When enabled, the model can access real-time web data to ground image generation (e.g., current weather, recent events, real products). This may add 2-5 seconds latency and requires internet connectivity. Use for time-sensitive or real-world subject matter; omit for fictional or privacy-sensitive content.'
Document the outputPath behavior: 'If outputPath is provided, the generated image is saved to disk and the response includes file_path. If omitted, the image is returned as base64 data only and not persisted to disk.'
Descriptions do not warn about potential risks: generate_image and edit_image can create images of people (controlled by personGeneration enum). Descriptions should note this is subject to Gemini's usage policies and may be restricted. edit_image lacks warning that editing could produce unexpected output.
No error classification or guidance. If API key is invalid, rate limit is hit, image is too large, or Gemini rejects the request, the tool description does not explain whether the error is retryable, user-fixable, or fatal, or what the agent should do next.
outputPath parameter is optional in generate_image and edit_image. No description explains what happens if omitted, does the image remain in memory, get deleted, or returned inline? Agents need to know the lifecycle and whether saving is the default behavior.
thinkingConfig in generate_image is underdocumented. What does thinkingLevel do? When should agents use HIGH vs MINIMAL? What does includeThoughts do, does it add to the response? Does it slow down the call? Descriptions are too vague for LLM decision-making.
useGoogleSearch parameter in generate_image is boolean with minimal guidance. When would an agent enable this? Does it slow the call? Does it cost extra? Agents need to know the tradeoff to decide whether to use it.
generate_image
Add limits and constraints to parameter descriptions. Example for images in generate_image: 'Up to 14 reference images total (10 object images + 4 person images). Each must be ≤10MB and in PNG, JPEG, WEBP, or GIF format.'
Add a note about idempotency and side effects. Example: 'Calling generate_image multiple times with the same prompt will produce different images (non-deterministic). If outputPath is provided, repeated calls will overwrite the file.'
Document model selection guidance: 'gemini-3.1-flash-image-preview is fastest and cheapest. Use gemini-3-pro-image-preview for higher quality. Nano Banana 2 models support ultra-wide/tall aspect ratios (4:1, 1:4, 8:1, 1:8). Ensure your API credentials support the selected model.'
Add sensitivity warning to generate_image and edit_image: 'Content generation is subject to Google Gemini usage policies. Images of people are controlled by personGeneration. Requests containing violence, explicit content, or misinformation may be rejected. Always set personGeneration=ALLOW_NONE if user age is unknown.'