An MCP server that generates ChaosBlade YAML configurations from natural language instructions using AI models. It provides both a web interface and API endpoints for creating chaos engineering experiments.
This MCP server exhibits fundamental definition quality gaps. While 7 tools are registered, most lack proper parameter descriptions and output schemas. Tool naming is inconsistent (mix of verb-noun and abbreviated patterns), and descriptions are either absent or minimal. No input validation rules, error recovery guidance, or security considerations are documented. The server appears to be a Flask wrapper around ChaosBlade chaos engineering with AI-powered YAML generation, but the tool definitions do not meet production standards for agentic tool composition. Critical issues include missing parameter descriptions (especially for 'instruction' and 'model' params in generate_yaml/batch_generate_yaml), no documented output schemas, and no error handling guidance.
批量生成YAML API - generates multiple YAML configurations from a list of natural language instructions
生成YAML API - generates a single YAML configuration from a natural language instruction
获取文件内容 - retrieves the content of a generated YAML file
获取已生成的文件列表 - retrieves the list of previously generated YAML files
获取可用模型列表 - retrieves the list of available AI models with their configuration status
获取模板列表 - retrieves predefined experiment templates
健康检查 - health check endpoint that verifies server status
Missing output schemas for all tools. No documented return types, field names, or structures. LLMs cannot plan downstream actions without knowing what fields to extract.
Parameter 'instruction' in generate_yaml/batch_generate_yaml lacks detailed description. No guidance on expected format, length constraints, or example patterns. Description should explain what constitutes a valid chaos experiment instruction.
Parameter 'model' in generate_yaml lacks description entirely. Is it optional? What are valid values? How does it interact with get_models? Undocumented dependencies prevent correct LLM usage.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 34 | - | v1 |
No error handling or recovery guidance. Tools return bare success/failure without categorizing errors as retryable, user-fixable, or fatal. Example: what does the LLM do if generate_yaml fails due to invalid instruction vs. API timeout vs. malformed YAML output?
Tool naming inconsistency. 'get_models' and 'get_templates' use standard verb_noun pattern, but 'generate_yaml' and 'batch_generate_yaml' are longer. More critically, parameter names like 'instructions' (in batch_generate_yaml) should be singular for consistency with other list-taking tools, or the tool should document why plural is preferred.
No documented constraints on 'instruction' parameter (max length, character restrictions, required fields). LLMs will pass arbitrary free-form text without validation, risking malformed input to the chaos engine.
get_file_content accepts a 'filename' parameter with no description of valid format, path traversal protection, or available file list. Should users pass full paths? Bare filenames? How does this integrate with get_generated_files?
No pagination support documented for get_generated_files. If hundreds of YAML files are generated, does the tool return all of them at once, risking context window exhaustion? Tool description does not mention limits or pagination params.
No security gate on destructive tools. generate_yaml and batch_generate_yaml create files without documented permission checks, rate limiting, or audit trails. No mention of secret injection for OpenAI API keys.
Batch operation (batch_generate_yaml) does not document partial failure semantics. If 5 of 10 instructions fail, what is returned? Per-item success/failure? Total count vs. failed count? LLM cannot plan recovery without this clarity.