Single tool 'generate-image' has a reasonable schema with proper type constraints (enums, min/max bounds) and descriptions for most parameters. However, the tool description is generic and lacks critical context about when to use it, what it costs, or how long it takes. Parameter descriptions are present but variable in quality, some are overly verbose (output_dir with OS-specific path examples), while others are minimal. The tool lacks error recovery guidance, confirmation mechanisms for an expensive operation, and documentation of expected output structure. Tool name 'generate-image' is acceptable but could be more specific (e.g., 'generate-flux-images'). Most critically, this is a STDIO-only server, which caps protocol readiness at 50 and severely limits production usability.
Tool description lacks critical context: does not explain what the tool does (flux-schnell model), when to use it instead of alternatives, cost/latency implications, or expected output structure. LLM cannot determine selection criteria.
No output schema documented. Tool returns ImageGenerationResponse but LLM only sees 'Content' in handler, doesn't know the structure of image_paths, metadata fields, or how to chain with downstream tools.
Error handling returns raw error.message wrapped in JSON. Does not categorize errors as retryable vs. fatal, does not provide recovery guidance, and does not validate inputs early. LLM receives 'Error: ...' with no actionable next step.
REPLICATE_API_TOKEN exposed as environment variable with fallback string 'YOUR API TOKEN HERE'. If this server is accidentally deployed without proper env setup, the token will be plaintext in logs. No secret injection pattern documented.
Recommendations
Expand tool description from current generic text to ~150 chars: 'Generate images using Replicate's flux-schnell model. Outputs URLs to generated images (1-4). Expect 10-30 second latency and ~$0.01-0.05 per image. Use when user requests image creation from text prompts.' This tells LLM WHAT (flux-schnell), WHEN (image creation), and CONSEQUENCES (latency/cost).
Document the complete output schema in a comment or separate field: { image_paths: string[], metadata: { model: string, inference_time_ms: number, cache_hit: boolean } }. LLM needs to know what fields exist for chaining.
Add descriptive text to 'filename' parameter: 'Base filename without extension. Output files will be saved as {filename}_1.{output_format}, {filename}_2.{output_format}, etc. If omitted, uses UUID-based names.' This eliminates path prediction uncertainty.
Restructure 'output_dir' to accept common paths like 'downloads', 'temp', or 'current', resolve internally using os.homedir() / process.cwd(). If absolute paths are required, validate with path.resolve() and reject traversal attempts (e.g., '../../../etc/passwd'). Provide clear error: 'output_dir must be an absolute path within your home directory or temp folder.'
Add error categorization in the catch block: Check for known Replicate errors (quota exceeded, model not ready) and return structured errors with retry guidance: { error: 'QUOTA_EXCEEDED', message: 'Daily image generation limit reached. Try again after 24 hours.', retryable: false } instead of raw error.message.
No confirmation mechanism for an expensive, time-consuming operation. Image generation can take 30+ seconds and incurs Replicate API costs. Agents should confirm before executing. Missing pattern:confirmation-request.
Parameter 'output_dir' requires absolute filesystem paths with OS-specific syntax (Windows double-backslash vs Unix forward-slash). This is unfriendly to LLMs and forces path construction logic into the agent. Should accept relative paths or use a safer default directory strategy.
Parameter 'filename' is optional but no guidance on naming collisions, file format handling (does it auto-append extension?), or what the actual saved filenames will be. LLM cannot predict output paths for downstream use.
No documentation of rate limits, latency expectations, or cost implications. LLMs need to know: Will this call succeed quickly? Is there a daily quota? Should I batch or spread requests? Missing operational context.
generate-image
Implement a confirmation step: Add optional parameter 'confirm: boolean' (default false). If false and cost_estimate > threshold, return { status: 'CONFIRMATION_REQUIRED', cost_estimate_usd: 0.03, message: 'This will cost ~$0.03. Call again with confirm=true to proceed.' }. This prevents accidental expensive calls.
Move REPLICATE_API_TOKEN to .env.example and validate at startup with a clear error message: 'REPLICATE_API_TOKEN environment variable is required. Set it in .env or via export REPLICATE_API_TOKEN=...'. Never fallback to a placeholder string.
Document latency and costs in the tool description: 'Note: Image generation typically takes 10-30 seconds. Each call costs approximately $0.01-0.05 depending on parameters.'
Add validation for 'num_outputs' with clear error message if user requests >4 images: 'Cannot generate more than 4 images per call. Requested: {num_outputs}. Use multiple calls to generate more images.'
In imageService.ts saveImages(), handle filesystem errors explicitly: check if output_dir exists and is writable before downloading images. Return clear error: 'Output directory does not exist or is not writable: {output_dir}. Create it first or use a different path.' This prevents silent failures.