MCP-compliant image augmentation server that bridges natural language processing with computer vision using the Albumentations library
The server provides 18 tools with schemas and descriptions, but exhibits significant quality gaps. Most tools have schemas and descriptions present, but many descriptions are generic or task-focused rather than LLM-optimized. Parameter descriptions are sparse or missing entirely for complex inputs. Error handling is present but not consistently integrated into tool definitions. The tool naming follows verb-noun conventions reasonably well, but lacks clarity in distinguishing between informational/reference tools (tools 9-18) which seem to serve prompt-engineering purposes rather than direct user actions. No tool annotations (readOnlyHint/destructiveHint) are declared. Security validation is implemented in server.py but not surfaced in tool descriptions. The server architecture shows thoughtful input validation and error recovery, but this quality is not reflected in the MCP tool interface itself.
Apply image augmentation transforms to an image using natural language prompts or presets
Generate a parsing prompt for a user augmentation request
Get examples and usage patterns for available transforms
Generate a policy prompt from a base preset with optional tweaks
Get recovery suggestions and error analysis for specific error types
Generate natural-language explanation of a pipeline JSON
Get current pipeline status and registered hooks
Tool descriptions lack LLM optimization and context. Descriptions like 'Get a condensed reference of transform keywords and their effects' and 'Get a comprehensive JSON guide for all available transforms with examples' do not explain WHEN to use the tool, what distinguishes it from similar tools, or what structure the agent should expect. The tool set includes 10 highly similar reference/discovery tools (tools 9-18) with overlapping purposes (transforms_guide, get_quick_transform_reference, available_transforms_examples, troubleshooting_common_issues, policy_presets, compose_preset, explain_effects, augmentation_parser, vision_verification, error_handler) that should either be consolidated or have clear, distinct descriptions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Get a condensed reference of transform keywords and their effects
List all available augmentation presets with their configurations
List all available image augmentation transforms with descriptions and parameters
Load an image from file path, URL, or base64 and prepare it for processing
Lightweight health check for the MCP server.
Get JSON of built-in augmentation presets with their policies
Set or clear the default random seed for reproducible augmentations
Get a comprehensive JSON guide for all available transforms with examples
Get solutions for common augmentation issues and errors
Validate an augmentation prompt for syntax and feasibility
Compare original and augmented images for consistency and verify requested transforms were applied
Parameter descriptions are missing or generic across multiple tools. For example, augment_image has parameters like 'image_path', 'image_b64', 'session_id', 'prompt', 'preset', 'seed', 'output_dir' with descriptions that are too brief to guide LLM usage. No parameter descriptions explain constraints: what constitutes a valid prompt? What image formats are accepted? What does session_id contain? No mention of parameter relationships (e.g., 'image_path and image_b64 are mutually exclusive; provide exactly one') required by pattern:constrained-input.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are declared in the visible schema. Tools like set_default_seed, augment_image, and load_image_for_processing modify state (WRITE risk), but this is not annotated in the tool definition itself, only in the external risk metadata. The MCP spec (2026-07-28) requires tools to declare destructiveHint for operations that modify state or have side effects. This forces clients to infer write-safety from tool name or external metadata rather than a structured declaration.
Reference/helper tools (tools 9-18) blur the boundary between user-facing tools and internal prompt-engineering utilities. Tools like 'augmentation_parser', 'explain_effects', 'vision_verification', and 'error_handler' appear designed to support the LLM's own reasoning rather than expose Albumentations functionality to users. These should be either removed from the tool list (handled internally), renamed to clarify they support agent planning, or significantly expanded with descriptions that explain when an LLM should call them vs. calling core augmentation tools directly.
Output schemas are not documented for any tool. The tool definitions show input schemas but do not declare what each tool returns. For example, augment_image description does not specify whether it returns a file path, base64-encoded data, or a URL. list_available_transforms does not specify the structure of transform definitions (are they nested objects, flat arrays, do they include examples?). Without documented output schemas, LLMs cannot plan downstream tool calls or extract data reliably.
Validation error messages may not meet pattern:recovery-guide requirements. The server implements validate_mcp_request() with good constraint checking (range validation, enum checks, string length limits), but these error messages are returned as tuples and may not include actionable recovery guidance. For example, if preset validation fails, the error should suggest 'preset must be one of: segmentation, portrait, lowlight. Try list_available_presets to see all options.' rather than a bare constraint message.
Tool naming ambiguity for reference/helper tools. Names like 'explain_effects', 'augmentation_parser', 'vision_verification' do not follow clear verb-noun conventions that make it obvious when to call them. 'explain_effects', is this a debug/introspection tool, or a user-facing explainability tool? 'augmentation_parser', does this parse user input or parse policy JSON? Without clarity, LLMs waste reasoning cycles guessing tool purpose.
Parameter mutually-exclusive relationships not documented. augment_image accepts both image_path and image_b64, the code validates that both can be provided, but the parameter descriptions do not state which is preferred, whether both can be provided simultaneously, or what happens if both are present. This forces LLMs to guess or experiment. Per pattern:constrained-input, mutually exclusive or interdependent parameters must be explicitly documented.