This MCP server exhibits significant definition quality gaps across multiple dimensions. While tool names follow verb conventions (generate_image, edit_image, clear_cache, etc.), the implementations lack proper schema documentation, parameter descriptions are sparse or missing type information, and error handling guidance is absent. The source provided is compiled JavaScript rather than TypeScript source, making verification of actual implementation details impossible. Tool descriptions are present but generic. Only 2 of 7 tools have reasonably detailed parameter schemas visible. The server demonstrates basic tool structure but falls well below production quality standards for agent tooling.
Missing output schemas across all tools. Agents cannot plan downstream calls or extract required fields (e.g., what fields does generate_image return? Is there an image_id, image_url, metadata?). This forces agents to guess at response structure, causing mid-chain failures.
Parameter descriptions lack format constraints, ranges, and enum values. E.g., 'size' parameter in generate_image lists examples ('1024x1024', '1536x1024') but no enum or validation rule. 'format' lists 'png', 'jpeg', 'webp' as examples but not as enforced enums. LLMs will hallucinate unsupported values. 'output_compression' lacks min/max range (0-100 is stated but not in schema). 'background' is completely underdescribed.
Document output schema for every tool. For generate_image, explicitly define: { image_url: string, image_id: string, size: string, generation_time_ms: number, is_streaming: boolean, partial_images_count?: number }. For get_conversation, return: { conversation_id: string, messages: [{ role: string, content: string, timestamp: ISO8601 }], created_at: ISO8601, updated_at: ISO8601 }.
Convert parameter examples to enums with validation. Replace 'size (e.g., 1024x1024, 1536x1024)' with enum: ['1024x1024', '1536x1024', '1024x1792', '1792x1024']. Replace 'quality (e.g., standard, high)' with enum: ['standard', 'high']. Replace 'format (e.g., png, jpeg, webp)' with enum: ['png', 'jpeg', 'webp'].
Add min/max constraints and format descriptions to numeric/string parameters. E.g., 'output_compression: integer, range 0-100 (0=maximum compression/lowest quality, 100=no compression/highest quality). Defaults to 75.' For 'background', define enum values: ['transparent', 'white', 'black'] with description of each.
Expand tool descriptions to guide agent selection. E.g., generate_image: 'Generate new images from text prompts. Use this when the user wants to create images from scratch. If you have existing images to modify, use edit_image instead. Supports streaming for real-time partial image previews.' edit_image: 'Modify existing images by providing base64-encoded images and a text description of changes. Use this when the user has an image and wants to refine, edit, or transform it. For new image creation, use generate_image.'
Score history
Overall score trend
↑ 9 points across a rubric change (v1 → v2)
50/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
50
2026-07-28+
v2
2026-03-09
F
41
-
v1
No error recovery guidance. Tools lack descriptions of failure modes, retryability, and next steps. E.g., if generate_image fails due to invalid prompt, no guidance offered. If get_conversation returns empty, no guidance. Pattern requires actionable error messages (e.g., 'User not found. Try search_users() first.').
Tool descriptions are inconsistently detailed. Some (generate_image, edit_image) claim to support streaming but return structures are undocumented. Others (clear_cache, cache_stats) are too terse. No guidance on tool selection, agents cannot distinguish when to use generate_image vs edit_image based on descriptions alone.
List tool (list_conversations) provides no pagination parameters or limits. Can return unbounded results, risking context window exhaustion. Pattern baseline requires page/offset, limit, and total_count for list operations.
Composition issue: generate_image and edit_image both require conversationId and useContext, but no clear guidance on how conversation context flows between them. No documented relationship or chaining mechanism. Agents may misuse these tools in sequence.
Source provided is compiled JavaScript (server.js), not TypeScript source. Cannot verify actual parameter validation, error handling, or output formatting logic. Evaluation forced to rely on tool metadata and descriptions, which may not reflect implementation reality. Must review src/server.ts for complete assessment.
Add pagination to list_conversations: accept 'limit' (default 20, max 100) and 'cursor' parameters. Return { conversation_ids: string[], total_count: number, next_cursor?: string }. Document in description: 'Returns paginated conversation history. Use cursor for fetching additional results.'
Document error modes and recovery. Add to generate_image description: 'Returns error if prompt is invalid, API limit exceeded, or image cannot be generated. If rate-limited, retry after 60 seconds. If prompt rejected, try a simpler or more specific description.' Similar guidance for edit_image (e.g., 'base64 decoding failure, image dimensions unsupported').
Add dry-run/confirmation for destructive tools. Document: 'clear_conversation has a confirm_clear_conversation_first tool that previews what will be deleted without executing. Call that first, then call clear_conversation to complete the operation.' Implement a two-step pattern.
Define conversationId format. Describe whether it's UUID, numeric, or arbitrary string. Specify max length and character restrictions. E.g., 'conversationId: string, 1-256 characters, alphanumeric and hyphen only (UUID v4 format recommended).'
Document streaming behavior in generate_image. Clarify: 'When stream=true, returns multiple partial image updates via Server-Sent Events. partialImages controls how many intermediate previews are sent (1-10). Use for real-time feedback. Defaults to stream=false for simple single-image response.'
Add composition guidance in parameter descriptions. E.g., generate_image: 'useContext: boolean. If true, uses conversation context to inform image generation. Requires conversationId to be set. Useful for generating follow-up images related to prior context.' Clarifies the dependency relationship.