MCP server for interacting with Kubernetes clusters, providing tools to manage and inspect Kubernetes resources
The server defines 8 tools with consistent naming conventions (verb_noun pattern: get_, list_, scale_). All tools have descriptions and input schemas with parameter types and descriptions. However, there are critical gaps in output schema documentation, error handling guidance, and parameter descriptions lack specificity (no enums, ranges, or format constraints). The tools follow a predictable pattern for Kubernetes resource operations but lack LLM-specific optimizations like enum constraints for Kubernetes-specific fields and actionable error messages. No tool annotations (readOnlyHint/destructiveHint) are present despite clear risk categorization (READ_ONLY vs WRITE). Descriptions are adequate but generic; they state WHAT the tool does but rarely explain WHEN to use it relative to similar tools or what dependencies exist.
Get details of a specific configmap
Get details of a specific deployment
Get details of a specific node
List configmaps in a namespace
List deployments in a namespace
List namespaces
List all nodes in the cluster
No output schemas documented. Tools return structured Kubernetes objects but LLMs cannot plan subsequent calls or extract required IDs without knowing what fields are present in responses (e.g., does get_deployment return status.replicas? conditionStatus array? Does list_deployments return total_count for pagination?). This violates the pattern:tool requirement to document return types and violates pattern:tool-chain, agents cannot compose tools without knowing output field names.
No tool annotations despite clear risk categorization. The server metadata indicates scale_deployment is WRITE and others are READ_ONLY, but these are not exposed as toolAnnotations (readOnlyHint/destructiveHint per MCP spec). LLMs cannot distinguish safely-retryable tools from potentially destructive ones without this metadata. scale_deployment should carry destructiveHint:true or at least idempotentHint:false to signal that retries require caution.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 55 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Scale a deployment to a specified number of replicas
Parameter descriptions lack Kubernetes-specific constraints and enums. fieldSelector and labelSelector parameters appear in list_* tools but descriptions say 'Selector to restrict the list of returned objects by their fields|labels', LLMs cannot infer valid selector syntax (e.g., 'status.phase=Running' for fieldSelector or 'app=frontend' for labelSelector). Add format hints or enum examples. Namespace parameter doesn't hint at 'default' as a common value. Replicas parameter in scale_deployment has no min/max constraints, LLMs could try passing 0 or 999999.
No pagination metadata in list tools. list_configmaps, list_deployments, list_namespaces, and list_nodes lack limit, offset, or next_cursor parameters. If a cluster has many resources, responses could be very large, blowing context windows. At minimum, tools should accept a limit parameter (default 20-50) and document result caps. Current descriptions do not mention pagination or result limits, violating pattern:paginated-result.
No error handling or recovery guidance. Descriptions do not explain what happens if a namespace does not exist, a deployment is not found, or scaling fails due to resource limits. Per pattern:recovery-guide, error responses should be actionable ('Deployment prod-api not found in namespace production. Available deployments: prod-app-v1, prod-app-v2, staging-app'). Current tool definitions offer no recovery paths.
Descriptions lack WHEN guidance. Descriptions state WHAT each tool does but rarely explain when to call one tool vs a similar one (e.g., when to use get_deployment vs list_deployments with a label selector). Per pattern:tool-description, descriptions should include 'WHEN to use it' and 'dependencies'. LLMs often misselect between similar-sounding tools (get_* vs list_*) when guidance is absent.
scale_deployment lacks confirmation or dry-run pattern. Scaling is an irreversible state change that can disrupt services. Per pattern:confirmation-request, destructive tools should support a dry-run mode or explicit confirm step. Current definition offers no such safeguard, an LLM could mistakenly scale a production deployment to 0 replicas, causing an outage.
No batch operations for common multi-resource workflows. If an LLM wants to list deployments and their configmaps, it must call list_deployments, then loop through results calling get_configmap repeatedly. Per pattern:tool-composition, batch variants (e.g., get_configmaps_by_names taking an array) reduce token waste and latency. Not critical for 8 tools, but limits agent efficiency.