A Kubernetes MCP server for troubleshooting and monitoring Kubernetes clusters with tools for analyzing deployments, pod health, metrics, network policies, and resource management.
Server provides 24 well-organized READ_ONLY tools with consistent naming (all verb_noun pattern: Get*). Tool descriptions are present and moderately detailed (averaging ~110 chars, within the 10-1024 guideline). However, schema completeness varies significantly. Input schemas are visible and include type definitions and descriptions for all parameters. Critical gaps: (1) NO output schemas documented, tools return structured text via TextContent but the LLM has no visibility into response structure; (2) NO pagination support despite tools returning lists (GetDeployments, GetPods, GetNamespaces, GetSecrets, etc.) that could exceed reasonable limits; (3) Parameter descriptions are generic and lack format/constraint details (e.g., 'Kubernetes namespace (defaults to 'default')' provides no guidance on valid namespace patterns); (4) No error recovery guidance, errors are returned via fmt.Errorf() with minimal context for LLM self-correction; (5) No tool annotations (readOnlyHint present implicitly via Risk tags in metadata, but not in schema). Naming is strong throughout, all tools start with action verbs (Get*) and distinguish clearly between list operations (GetDeployments, GetNamespaces) and detail operations (GetDeploymentDetails). Parameter names match across tool chains (Namespace, Name used consistently). Tool composition is good, single responsibility per tool, appropriate granularity. However, output chaining is weak, response structure is not documented, making downstream tool composition difficult for LLMs.
Gets detailed YAML representation of a Kubernetes resource by type and name.
Provides a summary of cluster health including node status, pod distribution, and recent critical events.
Lists all ConfigMaps in a namespace with data key counts.
Gets detailed information about a specific deployment including replica status, strategy, pod selector, template, conditions, and recent events.
Lists all deployments in a specified Kubernetes namespace with their replica status and age information.
Gets detailed endpoint information including subsets with addresses and ports.
Lists endpoints in a namespace showing endpoint addresses and port information.
No output schemas documented. All 24 tools return unstructured TextContent (markdown tables or free-form strings). LLMs cannot parse response structure, cannot extract IDs for downstream tool calls, and cannot validate response completeness. Example: GetDeployments returns a markdown table; downstream tools need deployment names and namespace, but the response structure is invisible to the agent.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Lists events in a namespace optionally filtered by resource type and name.
Gets detailed health check configuration for a pod including liveness, readiness, and startup probes for all containers.
Lists all namespaces in the cluster with status and age information.
Lists network policies in a namespace with pod selectors and policy types.
Gets detailed network policy information including ingress and egress rules with selectors and ports.
Gets disk usage and pressure status for all nodes in the cluster.
Lists node health status for all cluster nodes showing ready status and pressure conditions.
Gets resource usage metrics for all nodes including CPU and memory usage and capacity.
Lists all cluster nodes with status, roles, version, IPs, OS image, and kernel information.
Gets detailed information about a specific pod including labels, container status, conditions, and recent events.
Lists pod health status in a namespace showing status, ready containers, restarts, and health check probe status.
Retrieves logs from a specific container in a pod with configurable number of tail lines.
Gets detailed resource metrics for a specific pod including per-container CPU and memory usage with requests and limits.
Gets resource usage metrics for all pods in a namespace including CPU and memory usage.
Lists all pods in a namespace with status, ready containers, restarts, and age information.
Lists resource quotas for a namespace showing CPU and memory request/limit quotas and usage.
Lists all Secrets in a namespace with type and data key counts.
No pagination support for list operations. Tools like GetDeployments, GetPods, GetNamespaces, GetSecrets, GetEvents accept no limit or offset parameters. A cluster with 500+ deployments would return all of them in a single unstructured response, exhausting context windows and degrading LLM reasoning. Production baseline: list tools should accept limit (1-100, default 20-50) and offset/cursor parameters, and return a total count.
Parameter descriptions lack format and constraint details. Example: 'Kubernetes namespace (defaults to \'default\')' does not specify valid namespace patterns (alphanumeric, hyphens, max 63 chars per RFC 1123). 'Pod name' does not state the format. 'TailLines' (integer) has no min/max bounds. Rubric baseline: 100% of A+ tools specify format, range, and allowed values directly in parameter descriptions.
No error recovery guidance. Error responses are returned via fmt.Errorf() with raw messages (e.g., 'failed to list deployments: <error>'). Rubric baseline: error responses must tell the LLM what to do next. Examples: 'Namespace \'invalid-ns\' not found. Call GetNamespaces to list available namespaces.' or 'Pod \'mypod\' not found in namespace \'default\', verify the pod exists and the namespace is correct.'
No tool annotations in schema. Tools are READ_ONLY (stated in metadata Risk field), but JSON Schema input definitions lack readOnlyHint annotations. Rubric baseline: A+ tools include tool annotations (readOnlyHint, destructiveHint, idempotentHint) in schema for discoverability and safety validation.
GetPodLogs has no container parameter validation or error guidance. Container defaults to 'first container if not specified', but if the first container is wrong, the LLM gets logs from the wrong source with no indication of the error. Missing: enumerate containers in response or return error 'Pod \'mypod\' has 3 containers: app, sidecar, init. Requested: unknown. Did you mean one of the above?'
Parameter 'ResourceType' in GetEvents and DescribeResource expects a string but has no enum constraint or validation guidance. Example: user might pass 'pods' (plural) instead of 'pod' (singular), or 'Pod' (capitalized) instead of 'pod'. Rubric baseline: declare enums for known sets of values to prevent hallucinated invalid values.