AI-powered image stylization server with Model Context Protocol (MCP) integration using FastAPI and OpenAI
Single tool with decent parameter schema but critically lacks comprehensive documentation. The stylize_image tool has a populated input schema showing 6 parameters with types and descriptions, which is a positive sign. However, the tool description, while present, is overly long (447 chars, well above the 200-char baseline for LLM optimization) and reads more like a technical API doc than an agent-friendly guide. Output schema is documented in the description text but not formalized. The tool violates several naming and composition patterns: it performs multiple distinct operations (single-style application vs. random 4-style generation) in one tool, which should be split. Parameters like 'api_key' and 'session_id' suggest trial/paid authentication logic embedded in the tool, which complicates agent reasoning. No explicit error handling guidance, no recovery suggestions, and no validation of input constraints (e.g., base64 format, style_id enum values). The per-parameter descriptions are present but generic ('Base64 encoded image data' lacks format validation detail). No information about pagination, limits, or how to handle large responses. Risk categorization as WRITE is correct but no idempotency guarantees or confirmation patterns documented.
Stylize an image with the specified style or generate 4 random styles. Args: image_base64: Base64 encoded image data style_id: ID of the style to apply (e.g., 'van_gogh', 'pixel_art'). If not provided, generates 4 random styles. user_prompt: Optional custom prompt to guide the stylization project_context: Optional context with brand colors, mood, etc. api_key: API key for authentication (if provided, skips trial) session_id: Trial session ID for anonymous usage (if no api_key) Returns: Dictionary with: For single style (when style_id provided): - stylized_image_url: URL of the stylized image - style_applied: The style that was applied - prompt_used: The final prompt used for generation For multiple styles (when style_id not provided): - multiple_styles: True - images: Array of objects with style_id, style_name, stylized_image_url, prompt_used - total_images: Number of images generated For all modes: - trial_info: Trial usage information (for anonymous users) - upgrade_options: Available packages (if trial expired)
Tool performs two distinct operations (single style + random multi-style generation) in one tool. Should be split into stylize_image_with_style and generate_random_stylizations to enable independent composition and clearer agent reasoning.
Tool description is 447 characters (well above LLM-optimized 50-200 char baseline). Exceeds baseline by 2.2x, diluting signal and wasting tokens. Rewrite as: 'Apply an artistic style to an image, or generate 4 random styled variations. Provide style_id for a single style, or omit it for random exploration.'
Authentication parameters (api_key, session_id) exposed as tool parameters. Credentials must never be parameters, use server-side secret injection via environment variables. Agents log all parameters; secrets leak into traces and history.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No output schema formalized (documentation is in-text only). Return structure is described in prose but not as a JSON Schema. LLMs need structured output schema to plan downstream tool calls and type-check responses.
No error handling guidance or recovery instructions. What happens if style_id is invalid? If base64 decoding fails? If the API rate-limits? Tool provides no 'next steps' for the agent to recover from failures.
Parameter 'style_id' lacks enum constraint. Description mentions examples ('van_gogh', 'pixel_art') but no formal enum declaration. LLMs may hallucinate style values outside the valid set.
No idempotency guarantee or retry safety declared. Tool is WRITE (creates images), but agents don't know if retrying with same input is safe or will generate duplicate output.
Parameter descriptions lack actionable constraints. 'Base64 encoded image data' does not specify max size, supported formats, or dimensionality. 'Optional context with brand colors, mood, etc.' is too vague, what are valid fields?