Production-grade MCP server for Kubernetes cluster diagnostics, workload health, and pod event inspection. Enforces strict read-only cluster observation.
This Kubernetes diagnostic server demonstrates solid naming conventions and reasonable descriptions, but has moderate gaps in schema documentation, parameter validation, and error handling. All 5 tools follow verb_noun naming (get_*, list_) which aids discoverability. Descriptions are domain-specific and contextual (10-200 chars each, within baseline range). However, input schemas are only partially visible in the source, the code shows Pydantic Field definitions but the complete JSON Schema representation is not explicitly documented in the response. Output schemas are not documented at all, forcing LLMs to infer field structures. Error handling exists in get_pod_logs but is minimal elsewhere. No tool annotations (readOnlyHint, etc.) despite all tools being read-only operations. Overall, the server is functional and domain-competent but lacks polish in agent-facing documentation.
List all cluster nodes with internal IP addresses, roles, Kubernetes versions, and ready states. ### Usage Guidelines - Diagnostic discovery tool for inspecting node capacity and Kubernetes control plane versions.
Inspect pod health across a namespace, detecting restart counts, OOMKills, and CrashLoopBackOffs. ### Usage Guidelines - Evaluates container state transitions and abnormal restart loops for incident debugging.
Retrieve real-time or post-crash container logs from a specific pod. ### Usage Guidelines - Crucial for diagnosing root causes during CrashLoopBackOff or application startup failures.
Inspect HTTP/HTTPS routing configurations, ingress classes, hosts, and TLS certificates. ### Usage Guidelines - Useful for network troubleshooting, DNS routing audits, and TLS certificate inspection.
Query recent Warning events across pods, PVCs, nodes, and deployments. ### Usage Guidelines - Surfaces FailedScheduling, FailedMount, BackOff, and Unhealthy probe warnings.
Output schemas not documented. Tools return complex nested structures (e.g., get_cluster_nodes returns capacity dict, get_pod_diagnostics returns containers array with state objects) but LLMs have no declared schema to plan downstream processing or validate field existence.
No tool annotations despite all operations being read-only. Spec 2026-07-28 supports readOnlyHint in tool metadata. Adding 'readOnlyHint: true' to all tools clarifies they are safe for unsupervised execution and improves agent reasoning.
Incomplete parameter descriptions for optional parameters. get_cluster_nodes has NO input schema at all (takes {} as input). get_pod_logs documents 'namespace' as required but the Field definition has no default, creating ambiguity about optionality. list_warning_events and list_ingresses have null defaults but descriptions don't explicitly state 'cluster-wide if omitted'.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | <=2025-11-25 | v2 |
Minimal error handling and recovery guidance. Only get_pod_logs catches ApiException and returns structured error. Other tools will surface raw Kubernetes exceptions (e.g., ConnectionError, ApiException) with no guidance for LLM recovery. Missing pattern: tell agent 'If pod not found, try list_pods() first' or 'Namespace may not exist, verify with get_namespaces()'.
No pagination support. list_warning_events and list_ingresses iterate over all items returned by the Kubernetes API with no limit. In large clusters, returning hundreds of events or ingresses wastes tokens and risks context window exhaustion.
get_cluster_nodes returns raw response without filtering. The 'capacity' dict includes both valid fields (cpu, memory, pods) and potential None values. For LLM reasoning, strip None values or document which fields may be absent. High-token-count responses (os_image, kubelet_version, all capacity fields for every node) dilute signal.