MCP server exposing pipeline monitoring, quality assessment, and fashion-specific tools for the JSON Ghost-Mannequin Pipeline to assist Warp agents with intelligent workflow assistance
This MCP server exposes 5 tools for a ghost-mannequin image pipeline. Tool naming is reasonably clear (verb_noun pattern for 4/5 tools, though 'recommend-parameters' is weak). Descriptions are present but generic, they lack the specificity LLMs need to choose between similar tools. Parameter schemas are visible and mostly typed, but lack actionable descriptions for constraint guidance. Output schemas are not documented. No error handling or recovery guidance is present. The server is STDIO-only, which is a hard transport limitation (capped at 50 for protocol readiness), but definition quality issues are significant: descriptions average ~60 chars (well below the 194-char baseline), parameters lack guidance on format/range, and there is no output schema documentation.
Analyze the quality of a garment image using the pipeline's QA system
Analyze failures and issues in a specific pipeline step
Get current status of all pipeline steps and processing queue
Analyze which rendering route (SDXL vs Gemini) to use for optimal quality
Get parameter recommendations based on garment type and historical performance
Descriptions are too generic and lack LLM selection guidance. 'Analyze the quality of a garment image using the pipeline's QA system' (77 chars) does not explain WHEN to call this vs debug-pipeline-step, what QA metrics are returned, or what happens on failure. Baseline for production tools is 194 chars; these average ~50.
Output schemas are not documented. LLMs cannot plan downstream calls or extract required fields (e.g., does get-pipeline-status return queue_length, step_timings, error_counts?). Without output schema, the LLM must guess what data is available.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 17 | - | v1 |
Parameter descriptions lack constraint guidance. 'step' (integer 0 - 6) has a range but no guidance on what each step does or how to pick one. 'garment_type' has an enum but no explanation of when to use 'shirt' vs 'jacket' or what quality adjustments each triggers. 'detailed' (boolean) provides no hint on cost/latency trade-off.
'recommend-parameters' is a weak verb. The tool selects parameter recommendations based on garment and fabric type, better names: 'get-parameters-for-garment' or 'suggest-processing-params'. Current name does not signal that this is a discovery/lookup tool, not a planning tool.
No error handling or recovery guidance. If analyze-garment-quality fails (e.g., invalid image format, corrupted JSON), what does the LLM do next? Are there retry-safe conditions? Should it call debug-pipeline-step? Should it ask the user for a different image? No guidance provided.
Composition flaw: 'get-pipeline-status' and 'debug-pipeline-step' overlap. Both inspect the pipeline, but the distinction is unclear. Does get-pipeline-status return per-step details? If so, when would you call debug-pipeline-step? If not, should get-pipeline-status include a 'step' parameter to focus on one stage?