This server has two tools with complete input schemas and descriptions, but the descriptions are generic and lack actionable context. Tool naming follows verb_noun convention, which is good. However, output schemas are not documented, error handling is minimal, and there are security concerns around credential management. The server operates via STDIO only, which is a hard constraint limiting remote usability. Per-tool analysis: generate_images has a basic but complete schema with appropriate constraints (numberOfImages 1-4), but the description lacks guidance on when to use it vs other image tools, what happens to generated files, or expected output structure. create_image_html has an adequate schema and description but offers no examples of output or error cases. Neither tool documents what data is returned beyond implicit inference.
Tools (2)
create_image_htmlread onlysource verified63/100
Create HTML img tags from image file paths with gallery view
API credentials exposed via environment variable without server-side secret injection pattern. GEMINI_API_KEY is read directly in index.ts line ~80 without wrapper or vault abstraction. If agent logs tool execution, API key could leak into traces.
Output schemas not documented. Neither tool describes what data is returned to the agent. generate_images returns file paths but no schema for the response object. create_image_html returns HTML but structure is undefined. LLMs cannot plan downstream steps without knowing response structure.
Tool descriptions lack recovery guidance and decision context. 'Generate images using Google Gemini AI' does not explain: when to use this vs other image tools, what happens to files, expected latency, rate limits, or how to handle generation failures. LLMs cannot make informed routing decisions.
generate_imagescreate_image_html
Recommendations
Add explicit output schema documentation to both tools. For generate_images: document return value as { filePaths: string[], message: string, totalGenerated: number }. For create_image_html: document return as { html: string, width: number, height: number, imageCount: number }.
Expand generate_images description to: 'Generate images from text prompts using Google Gemini Imagen-3. Creates PNG files on the user's desktop (AI-Generated-Images folder). Consumes API quota. Returns an array of file paths. Use this to create original artwork; for existing images, use image editing tools instead. Supports 1-4 images per call.'
Add 'prompt' parameter constraints: minimum length 5 characters, maximum 1000 characters, regex pattern forbidding shell metacharacters. Document: 'The prompt describes the desired image. Be specific (e.g. "a red fox in a snowy forest" not "fox"). Longer prompts (20-100 words) produce better results.'
Implement input sanitization for 'category' parameter. Validate against path traversal: reject '.', '..', absolute paths, and directory separators. Return error: 'Invalid category. Use only alphanumeric characters and hyphens (e.g. 'landscapes', 'character-designs').'
Refactor credential handling: move GEMINI_API_KEY to a dedicated secrets manager or environment loader with explicit permission checks. Never pass credentials through tool parameters or include them in logs.
Add actionable error responses: instead of 'No images were generated', return 'Image generation failed. Possible causes: (1) prompt blocked by safety filter, try less graphic descriptions, (2) API quota exceeded, wait before retrying, (3) network timeout, retry in 30 seconds. Current quota usage: X/Y.'
Error handling is minimal and non-actionable. McpError(2, 'No images were generated') provides no recovery guidance. LLMs see this error but don't know whether to retry, check the prompt, increase numberOfImages, or request user intervention.
No input validation or constraint enforcement for 'prompt' parameter in generate_images. The parameter has no minimum length, no pattern, no guidance on what prompts work vs fail. LLMs can pass empty strings, extremely long strings, or unsupported formats without feedback from the schema.
Tool descriptions lack state-change warnings. generate_images modifies filesystem and uses paid API quota, but the description does not explicitly state 'This tool creates files on disk and consumes Gemini API quota.' Agents need to know which calls have irreversible consequences.
No audit trail or permission gates. Any agent calling this server can generate unlimited images, consuming API quota and disk space without rate limits or user approval. No logging of who/what/when/why calls were made.
Path traversal vulnerability in 'category' parameter. The category string is joined directly to the output path without sanitization (path.join(baseDir, category)). An attacker could pass '../../../etc/passwd' to write files outside the intended directory. Input must be validated against path traversal patterns.
generate_images
Implement rate limiting: enforce max 10 calls per minute per agent session, max 100 images per day. Return 429 error with retry-after guidance when limits are hit.
Add per-call audit logging: log caller identity, timestamp, prompt length, numberOfImages, output directory, and result (success/failure). Do NOT log the full prompt or image data, only metadata.
Document the create_image_html tool with more context: 'Wraps image file paths in HTML img tags and optionally creates a CSS gallery view. Use this after generate_images to format the output for display. Returns raw HTML; embed in a document or web page. Set width/height to match your layout (default 512x512 pixels). Set gallery=true for a multi-column responsive layout.'
Add idempotency documentation to generate_images: 'This tool is NOT idempotent. Each call generates new images with unique filenames (timestamp-based). Retry the same prompt with the same settings will generate different images, not return cached results.'
Add discovery guidance: 'Call this tool to create original images from descriptions. If you need to edit, resize, or format existing images, use image editing tools instead. If you need to search for existing images online, use a web search tool.'
Validate numberOfImages input strictly: enforce that it is an integer (not float), within 1-4 range, with error message: 'numberOfImages must be an integer between 1 and 4 (got <value>).'
Document expected latency and retry behavior: 'Image generation typically takes 10-30 seconds per image. If the call times out, retry with exponential backoff (wait 5s, then 10s, then 30s). Failed images do not consume quota.'
Add imagePaths validation in create_image_html: verify all paths exist and are readable before generating HTML. Return error: 'Image file not found: <path>. Verified files: <list>. Check paths and retry.'
Add gallery layout documentation: 'When gallery=true, generates a responsive CSS grid that adapts to screen width (1-4 columns). Useful for viewing multiple generated images at once. When gallery=false, returns a simple vertical list of img tags.'