Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This is a chaos engineering server with 14 tools targeting Kubernetes fault injection. While tool definitions are visible in the source code, there are consistent and significant quality gaps. Most tools have basic descriptions (60-150 chars) but parameter descriptions are often generic and lack actionable constraints. Input schemas are present but underspecified, many parameters accept free-form strings where enums or format constraints would be more appropriate. Error handling exists but is minimal; responses provide basic recovery hints but do not categorize errors or guide retry logic clearly. The security posture is weak: the server handles Kubernetes credentials through environment variables, but there is no evidence of per-tool permission gating or audit logging. Most concerning: many high-risk tools (pod_kill, container_kill, delete_experiment) lack confirmation/dry-run patterns despite being explicitly marked IRREVERSIBLE. These are production-grade chaos experiments that should never run without explicit confirmation.
No enum constraints on mode/direction/op parameters despite documented valid values. These parameters accept free-form strings, inviting hallucinated values from LLMs. The pod_kill docstring lists 'one', 'all', 'fixed', 'fixed-percent', 'random-max-percent' for mode, but the input schema defines mode as type:string with no enum. network_partition and network_bandwidth accept direction as free-form string when it should be enum [to, from, both]. host_disk_fault.op should be enum [read, write, append].
IRREVERSIBLE operations (pod_kill, container_kill, pod_failure, delete_experiment) lack confirmation or dry-run patterns. These are high-risk chaos experiments that should never execute without explicit user confirmation. Agents can invoke them without safety barriers, risking unintended cluster disruption. The source code includes error handling for failures but no pre-execution confirmation flow.
Recommendations
Convert mode, direction, op, and experiment_type parameters to JSON Schema enums. For pod_kill.mode, use enum: [one, all, fixed, fixed-percent, random-max-percent]. For network operations, use enum: [to, from, both]. For delete_experiment.experiment_type, enumerate all valid chaos experiment types.
Add regex patterns or format constraints to string parameters. In the JSON schema, add pattern: '^\d+(\.\d+)?(ms|s|m|h)$' for duration parameters (e.g., '5m', '100ms'). For size parameters, use pattern: '^\d+(\.\d+)?(B|KB|MB|GB|%)?$'. For rate, use pattern: '^\d+(mbps|gbps|kbps)$'.
Implement confirmation/dry-run patterns for IRREVERSIBLE tools. Before executing pod_kill, container_kill, pod_failure, or delete_experiment, call a confirm_chaos_experiment tool that returns the impact summary and requires explicit user approval. Or add a dry_run boolean parameter (default false) that shows what would happen without applying the experiment.
Document parameter interdependencies explicitly. For tools with mode+value pairs, add to each parameter description: 'When mode is fixed, value must be an integer. When mode is fixed-percent, value must be a percentage (0-100). When mode is random-max-percent, value must be a percentage.' Use this pattern for all mode+value tools.
Expand tool descriptions to 100 - 200 chars with WHEN/WHY context. Example: 'Kill pods of a service to test resilience to pod failures. Use when you want to verify that the application recovers from random pod restarts. Do NOT use if the application is stateless; use pod_cpu_stress instead to test performance degradation.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 59 points across a rubric change (v1 → v2)
59/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
59
<=2025-11-25
v2
2026-03-09
F
0
-
v1
72/100
Inject network delay fault.
network_lossreversiblesource verified69/100
Inject network packet loss fault.
network_partitionreversiblesource verified66/100
Inject a network partition fault.
pod_cpu_stressreversiblesource verified72/100
Apply CPU stress on pods.
pod_failureirreversiblesource verified64/100
Inject a failure into pods of a service.
pod_killirreversiblesource verified62/100
Kill pods of a service with improved error handling.
Parameter interdependencies undocumented. The 'value' parameter's meaning depends on 'mode' (e.g., fixed vs fixed-percent interprets value differently), but neither parameter description states this relationship. LLMs cannot infer cross-parameter constraints and will guess, leading to invalid experiments.
Format constraints on string parameters are not machine-parseable. Parameters like latency ('100ms'), jitter ('100ms'), size ('256MB' or '50%'), rate ('1mbps'), loss ('10%') accept format strings but lack regex patterns or explicit format descriptions in the JSON schema. LLMs cannot validate these before calling, risking invalid syntax errors.
No per-tool permission gating or scope declaration. Chaos Mesh tools are high-risk, they modify cluster state and can cause availability issues. The server lacks checks for caller authorization, and tool definitions do not declare required permissions (e.g., 'chaos.io/read', 'chaos.io/write'). Any authenticated user can invoke any tool.
No audit logging. The source code includes basic logging (logger.info, logger.error) but does not log who called what tool, with which parameters, at what time, or what the outcome was. For chaos experiments, audit trails are essential for incident response and compliance.
Descriptions are too brief and lack WHEN/WHY context. Many tool descriptions are 60 - 75 characters and state only WHAT the tool does. They lack guidance on WHEN to use this tool instead of a similar one (e.g., pod_kill vs pod_failure vs pod_cpu_stress). LLMs cannot select the right tool without explicit context. Baseline for A+ tools: 50 - 200 chars with clear WHEN/WHY context.
Output schemas not documented. Tool docstrings do not specify the structure of returned dicts. For example, pod_kill returns a dict described as 'The applied experiment's resource in Kubernetes', but the structure is not documented. LLMs cannot plan downstream tool calls without knowing what fields are available.
Ambiguous external_targets parameter. In network_partition, network_bandwidth, network_delay, and network_loss, the external_targets parameter accepts 'service names or IPs' but does not specify how to distinguish them. Is 'adservice' a service name and '10.0.0.1' an IP? What if a service name looks like an IP? This ambiguity forces LLMs to guess.
No rate limiting or concurrency safeguards. The server accepts batch operations (e.g., host_cpu_stress takes an array of IPs) but does not document or enforce limits on batch size. An agent could pass 1000 IPs and overwhelm the Chaos Mesh API.
Error handling lacks recovery guidance and error classification. The pod_kill function catches exceptions and returns a dict with 'error' and 'suggestion' fields, but most other tools do not. Errors are not categorized as retryable, user-fixable, or fatal. LLMs cannot determine if they should retry, ask the user, or give up.
Document output schemas in docstrings. Add to each tool: 'Returns a dict with keys: {name (str), namespace (str), kind (str), status (str), message (str)}.' Match the actual structure returned by fault_inject.pod_fault().
Add per-tool permission checks. Before executing, verify that the calling principal has a scope like 'chaos.io/write:pod_kill' or 'chaos.io:admin'. Log the permission check result. Return a clear error if the principal lacks permission: 'Permission denied: principal john@example.com lacks scope chaos.io/write:pod_kill. Contact admin to grant access.'
Implement structured audit logging. For each tool invocation, log: {timestamp, caller_id, tool_name, parameters, result_status, error_message}. Write to a dedicated audit log file or send to a centralized logging system (CloudWatch, Datadog, etc.).
Add error classification and recovery guidance to all tools. When an error occurs, return: {error (str), error_type (str: retryable|user_fixable|fatal), recovery_hint (str)}. Example: {error: 'Service not found', error_type: 'user_fixable', recovery_hint: 'Try get_services() to list available services.'}.
Clarify external_targets format. Add parameter description: 'Array of target service names (e.g., ["adservice", "backend"]) or IP addresses (e.g., ["10.0.0.1", "10.0.0.2"]). Mixing service names and IPs is allowed.'
Document batch size limits for array parameters. Add to tools with array params: 'Maximum array size is 50 items. For larger operations, split into multiple calls.'
Add dry_run parameters to high-risk operations. For pod_kill, add dry_run: bool (default false). When dry_run=true, return the experiment spec that would be applied without actually submitting it to Chaos Mesh.
Reduce generic parameter descriptions. Replace 'Mode of pod selection' with 'Pod selection mode: one (random), all (all eligible), fixed (N pods), fixed-percent (N% of pods), random-max-percent (max random %).' This is machine-parseable and eliminates ambiguity.