vision-mcp has two tools with significant quality gaps. Both tools have reasonable names and some description, but parameter schemas lack complete type definitions and descriptions are generic. Tool descriptions are under 100 characters and lack WHEN/WHY guidance. Parameters are sparsely documented, most lack type info or actionable constraints. No output schemas are documented. Error handling is absent from tool definitions. No evidence of input validation or recovery guidance. The server exposes external API integration (Segmind, Stable Diffusion) but provides no handling for API failures, rate limits, or auth errors.
Tools (2)
gen_imagewritesource verified46/100
Generates an image from a prompt using Stable Diffusion 3.5 and returns pre-signed URL
Parameter descriptions are missing or incomplete. 'blend_mode' has a description but no enum values or valid options. 'image_url', 'text', 'out_dir' lack guidance on format/constraints. 'prompt' in gen_image lacks actionable constraints. LLMs cannot determine valid inputs.
Input schemas lack complete type definitions. While some parameters have 'type' and 'description' fields, there are no format constraints (regex patterns, length limits), no enums for constrained fields (blend_mode should be an enum), and no min/max for numeric 'steps'. Schemas are minimal and insufficient for validation.
Output schemas are not documented in tool definitions. Callers have no way to know what fields gen_text_overlay returns (just a URL? dimensions? processing time?) or what gen_image returns (image URL? dimensions? metadata?). Absence of documented return types forces LLMs to guess downstream field names.
gen_text_overlaygen_image
Recommendations
Add complete parameter descriptions to all tools. For gen_text_overlay, clarify: blend_mode (enum: normal, multiply, screen, overlay, darken, lighten, etc.), image_url (must be publicly accessible HTTP/HTTPS URL, max 50MB), text (max 500 characters, supports ASCII and Unicode), out_dir (path where output image is written; must be writable; defaults to 'data/output.jpg'). For gen_image, clarify: prompt (max 1000 characters, describe desired image in detail), steps (integer 1 - 100, trade-off between quality and speed, default 10).
Define and document output schemas. gen_text_overlay should return: { 'output_path': string (local file path to processed image), 'width': int, 'height': int, 'processing_time_ms': int }. gen_image should return: { 'image_url': string (HTTPS pre-signed URL), 'expires_at': string (ISO 8601 timestamp), 'width': int, 'height': int, 'seed': int }.
Add structured error handling guidance to tool descriptions. Both tools should document: 'On invalid image_url: returns 'Image URL unreachable or unsupported format. Try a public HTTPS URL.' On auth failure: returns 'API key invalid or expired. Check server configuration.' On rate limit: returns 'API limit reached. Retry after 60 seconds.' On timeout: returns 'Processing took longer than 30s. Try smaller image or simpler text.'
Fix out_dir parameter: rename to 'output_file_path' to match the default behavior, or rename the default to 'data/output/' if it truly is a directory. Add constraint: 'Path must be writable; directory must exist or be created by the server.'
No error handling or recovery guidance in tool definitions. Both tools call external APIs (Segmind, Hugging Face), these can fail with auth errors, rate limits, invalid inputs, or timeouts. No tool description mentions what can go wrong or how to recover. LLMs will be blindsided by failures.
Tool descriptions are too short and lack WHEN/WHY context. gen_text_overlay: 'Overlay text on an image using Segmind API.' (55 chars) does not explain when to use it vs alternatives, what prerequisites exist, or what 'blend mode' means to an LLM unfamiliar with image processing. gen_image: 'Generates an image from a prompt using Stable Diffusion 3.5 and returns pre-signed URL' (87 chars) mentions pre-signed URL but does not explain what that is or how long it remains valid.
Ambiguous parameter naming. gen_text_overlay has 'out_dir' defaulting to 'data/output.jpg', a file path, not a directory. Parameter name says directory but default is a file. LLMs will be confused about whether to pass a folder path or file path. Also, no guidance on whether the directory must exist or be writable.
Missing required parameter descriptions. gen_image 'prompt' parameter has a description but no guidance on length, content restrictions, or what prompts are invalid. gen_text_overlay 'image_url' lacks any description of expected URL format, accessible origins, supported image formats, or max size. 'text' parameter has no guidance on length, fonts, or character restrictions.
No input validation constraints visible in schemas. gen_image 'steps' defaults to 10 but has no min/max bounds. Can it be 1? 500? 10000? No constraint. blend_mode defaults to 'normal' but what other modes are valid? The schema does not enforce or document valid enum values.
Parameters expose or rely on external service details. gen_text_overlay references 'Segmind API' in description, implying LLM must know what Segmind is. gen_image references 'pre-signed URL', a technical implementation detail. Tool descriptions should hide API internals and focus on user intent.
gen_text_overlaygen_image
Add enum constraint for blend_mode in gen_text_overlay schema. Enum values: ['normal', 'multiply', 'screen', 'overlay', 'darken', 'lighten', 'color_dodge', 'color_burn', 'hard_light', 'soft_light', 'difference', 'exclusion']. Update parameter description: 'Blending mode controls how text opacity/color merges with the image background. Use normal for simple overlay, multiply for darkening, screen for lightening.'
Add min/max constraints for numeric parameters. gen_image steps: minimum 1, maximum 100, default 10. Update description: 'Number of inference steps (1 - 100). Higher values improve quality but take longer.'
Enhance tool descriptions with WHEN/WHY context. gen_text_overlay: 'Overlays text on an image with customizable blending. Use this to add captions, labels, or annotations to images. Requires a publicly accessible image URL.' gen_image: 'Generates an image from a text description using Stable Diffusion 3.5. Returns a pre-signed HTTPS URL (valid for 24 hours). Use this to create new images, illustrations, or variations on a concept.'
Document prerequisites and dependencies. gen_text_overlay requires Segmind API key (server-side injected, not exposed to LLM). gen_image requires Hugging Face Diffusers library and GPU or sufficient CPU. If either requirement is missing, tool calls will fail, document this in server README with clear error messages for misconfiguration.
Add idempotency/retry guidance. Both tools call external services that may be temporarily unavailable. Clarify in descriptions: 'If the call fails with a timeout, it is safe to retry. If it fails with 'Invalid API key', do not retry, contact server administrator.'
Return metadata that enables chaining. If gen_text_overlay output_path is the result, a downstream tool should be able to pass it directly. If gen_image returns image_url, ensure the URL remains valid for the entire agent conversation so downstream tools can fetch/process it. Document expiration in the response.