A Kubernetes MCP server providing tools for cluster management, resource operations, metrics retrieval, and optional Prometheus/Loki integration
This Kubernetes/observability MCP server has 23 tools with variable definition quality. Strengths: all tools have descriptions (passing the critical check); tools are mostly verb-noun named and leverage familiar Kubernetes/Prometheus/Loki APIs. Weaknesses: many descriptions are generic and brief (avg ~80 chars, below the 194-char production baseline); parameter descriptions are sparse or absent for several tools; output schemas are not documented in the tool definitions; error handling and recovery guidance are not evident in the tool specs; no tool annotations (readOnlyHint/destructiveHint) despite having both read-only and destructive operations (delete_resource, create_or_update_resource_yaml/json). Security concern: send_to_feishu accepts feishu_webhook_url as a parameter, violating secret-injection patterns. Conservative estimate: definition quality averages ~58 across all tools, passing on description presence but losing points for lack of depth, schema documentation, and error guidance.
Create or update a Kubernetes resource from JSON manifest
Create or update a Kubernetes resource from YAML manifest
Delete a Kubernetes resource
Get detailed description of a Kubernetes resource
Get active alerts from Prometheus
Get API resources available in the Kubernetes cluster
Get events from the Kubernetes cluster
Get ingress resources from the cluster
Webhook URL exposed as parameter: send_to_feishu accepts 'feishu_webhook_url' as a required parameter instead of injecting it server-side. This violates secret-injection patterns and risks credential leakage into logs and agent traces.
Output schemas not documented in tool definitions. Agents cannot predict response structure, making chaining and data extraction error-prone. All 23 tools lack documented return types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Get available values for a specific log label from Loki
Get available log labels from Loki
Get log streams matching a selector from Loki
Get available metric names from Prometheus
Get resource metrics for a specific node
Get resource metrics for a specific pod
Get logs from a Kubernetes pod
Get a specific Kubernetes resource by kind, name, and namespace
List Kubernetes resources of a specific kind with optional filtering
Execute an instant query against Prometheus
Execute an instant query against Loki
Execute a range query against Loki
Execute a range query against Prometheus
Perform a rollout restart on a Kubernetes resource
Send a message to Feishu via webhook
Missing tool annotations: destructive/write operations (delete_resource, rollout_restart, create_or_update_resource_*) lack readOnlyHint/destructiveHint/idempotentHint tags. LLMs cannot distinguish safe reads from risky writes without hints.
Error handling not visible in tool definitions. No error response guidance or recovery instructions. Agents will have no context for handling timeouts, API errors, or invalid inputs.
Generic/short descriptions: get_metric_names ('Get available metric names from Prometheus'), get_alerts ('Get active alerts from Prometheus'), and several log tools have descriptions under 70 chars. Insufficient context for LLM tool selection.
Sparse parameter descriptions: get_metric_names and get_alerts have no input parameters, so descriptions cannot guide LLM usage. Many parameters like 'TailLogsLen' in get_pods_logs lack usage context (e.g., 'default: 100' but no explanation of tradeoffs).
No pagination/result limits documented. list_resources and similar discovery tools may return hundreds of items, bloating context. No mention of max result count or pagination parameters.
Parameter naming inconsistency: get_pods_logs uses 'Name' and 'TailLogsLen' (capitalized), while most other tools use snake_case. This inconsistency confuses agents and violates the verb_noun convention.