Governed Kubernetes operations for AI agents with 55 MCP tools including audit, budget, undo, and risk-tier audit labels
The server implements 11 Kubernetes operations with explicit JSON Schema definitions and clear descriptions. Naming follows verb_noun conventions (scale_deployment, delete_job, cordon_node). Most tools include parameter descriptions and schemas. However, several critical gaps reduce the score: (1) output schemas are not documented, tools return results but the structure is not specified in the submission; (2) error handling descriptions are minimal, no guidance on recovery steps or error classification; (3) parameters lack format constraints (e.g., no enum for namespace, no regex for names); (4) descriptions average ~60 chars but lack WHEN/WHY context required by pattern:tool-description; (5) destructive tools (delete_deployment, delete_namespace, drain_node) lack comprehensive confirmation workflows despite acknowledging HIGH RISK. The governance harness (audit, undo, risk-tier) is mentioned but not surfaced in tool descriptions, making it invisible to LLM decision-making.
Flag CrashLoopBackOff, image-pull failures, OOMKilled, unschedulable, high restarts.
Flag Deployments/StatefulSets/DaemonSets with ready<desired or stuck rollouts.
Mark a node unschedulable (destructive — double confirm).
Create a namespace.
Delete a deployment and its pods (HIGH RISK — double confirm).
Delete a job and its pods (destructive — double confirm).
Delete a namespace and EVERYTHING in it (HIGH RISK — double confirm).
Output schemas not documented. Tools return results (pod rows, workload status, operation outcomes) but the response structure is not visible in the submission. LLMs cannot plan downstream tool calls or extract chaining IDs without knowing what fields to expect.
Descriptions lack WHEN/WHY context. Most descriptions state WHAT the tool does (60 chars avg) but not WHEN to use it or how it differs from similar tools. E.g., 'scale_deployment' → 'Scale a deployment to a replica count (medium risk, single confirm)' does not explain when scaling vs rollout-restart; 'collect_pod_rows' flags failures but does not say what threshold triggers a flag.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
Cordon a node and evict its pods (HIGH RISK — double confirm).
Trigger a rolling restart of a deployment (medium risk — single confirm).
Scale a deployment to a replica count (medium risk — single confirm).
Mark a node schedulable again.
No input validation guidance. Parameters like 'namespace', 'name', 'label_selector' lack format constraints. E.g., are namespace names regex-validated? Does label_selector accept arbitrary strings or must it follow Kubernetes label syntax? No enum, pattern, or minLength constraints visible. Invites hallucinated invalid values.
Destructive tools lack confirmation patterns. delete_deployment, delete_job, delete_namespace, drain_node all mark HIGH RISK but only accept a 'dry_run' flag. No explicit confirmation step, user override, or approval gate is visible. pattern:confirmation-request requires LLM to verify intent before irreversible operations.
Error handling not documented. No description of what exceptions each tool raises, how failures are classified (retryable vs fatal), or what guidance is returned to the LLM. E.g., does 'deployment not found' return a recoverable error with suggestions to search for similar names? Or a raw 404?
Governance harness (audit, undo, risk-tier, budget) is implemented but invisible to tool descriptions. LLMs cannot reason about audit implications, undo availability, or budget constraints if these features are not surfaced in the tool interface. Descriptions should explicitly mention that operations are audited and undoable.
Parameter interdependencies not documented. E.g., 'namespace' and 'target' are both optional but may have a relationship (target could imply a default namespace). The interaction is not explained, forcing LLMs to guess which to supply.
Diagnostic tools (collect_pod_rows, collect_workload_rows) lack pagination and result limits. No description of max rows returned, no cursor/offset params for large clusters. Returning unbounded results risks context window exhaustion and token waste.