Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server has significant definition quality gaps. While it exposes 17 tools via HTTP with basic descriptions, the definitions lack depth, structure, and LLM-optimization. Most parameter descriptions are generic, schemas are incomplete (many parameters lack type specifications beyond string), and output formats are undocumented. Error handling is minimal. The server relies on raw API responses without shaping, which will waste LLM context. Naming is acceptable (verb-noun convention mostly followed), but overall quality falls in the D-range (poor). The destructive operations (delete_resource, scale_deployment, apply_yaml, exec_pod_command) lack confirmation patterns or dry-run options, creating safety risks.
Tools (17)
apply_yamlwritesource verified55/100
Apply a Kubernetes YAML configuration to the cluster.
automate_remediationwritesource verified48/100
Automatically remediate common Kubernetes issues detected in a pod or deployment.
create_deploymentwritesource verified58/100
Create a new deployment in Kubernetes cluster.
delete_resourcedestructivesource verified52/100
Delete a Kubernetes resource (pod, deployment, service, etc).
describe_podread onlysource verified62/100
Get detailed information about a pod for troubleshooting.
exec_pod_commandwritesource verified53/100
Execute a command inside a running pod.
get_cluster_inforead onlysource verified63/100
Get basic Kubernetes cluster information and health status.
Parameters lack type specifications and constraints. Most params (namespace, replicas, lines, target, port, etc.) are defined as type 'string' with minimal validation hints. This invites invalid input from LLMs, e.g., 'replicas' should be an integer with range 1-100, not a free-form string.
Output schemas are not documented. Tools return raw Kubernetes API responses or unstructured text without describing what fields the LLM should expect. This forces LLMs to parse and reason about unstructured data, wasting tokens and increasing errors.
Convert all string parameters to properly typed inputs. Use 'integer' for replicas, lines, port; 'enum' for namespace (dropdown of available namespaces), show_all (true|false), and resource_type (pod|deployment|service|etc). Add min/max constraints: lines 1-1000, port 1-65535, replicas 0-1000.
Document output schemas for every tool. Create a reference like: 'Returns: {pods: [{name, namespace, status, ready_replicas, desired_replicas}], total_count, last_updated}'. This guides LLM parsing and enables tool chaining.
Add confirmation patterns to destructive tools. E.g., delete_resource should accept a 'confirm=true' parameter and return 'This will permanently delete <resource_type>/<resource_name>. This cannot be undone. Are you sure?' unless confirm=true. Alternatively, implement a separate 'confirm_delete_resource' tool that the agent must call first.
Enrich error messages with recovery hints. E.g., 'Pod "web-123" not found in namespace "default". Available pods: web-1, web-2, web-prod. Call list_pods(namespace="default") to explore.' This enables self-correction without extra calls.
Move GCP project_id, cluster_name, and zone from parameters to environment variables or a server config file. Update get_gke_cluster_metrics to accept only optional overrides, defaulting to env vars.
Add permission-gate middleware: before executing delete_resource, scale_deployment, exec_pod_command, or apply_yaml, check the calling agent's RBAC token or identity. Return a clear 'Insufficient permissions: you lack write:pods scope' error if unauthorized.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 35 points across a rubric change (v1 → v2)
48/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-21
F
48
<=2025-11-25
v2
2026-03-09
F
13
-
v1
read onlysource verified60/100
Get status of deployments in a namespace.
get_deploymentsread onlysource verified62/100
Get deployments in a specific namespace with their status and replica information.
Destructive operations (delete_resource, scale_deployment, exec_pod_command, apply_yaml, automate_remediation) lack confirmation patterns or dry-run capability. An agent mistake can permanently delete resources. No 'confirm_before_execute' pattern is visible.
Parameter descriptions are generic and often vague. E.g., 'Name of the pod to analyze' (suggest_troubleshooting) doesn't explain what 'analysis' returns or when to use this vs describe_pod. Descriptions should state WHAT, WHEN, and WHAT IT RETURNS.
Error handling is not visible or actionable. No evidence of error categorization (retryable, user-fixable, fatal) or recovery guidance. Tools that fail (e.g., pod not found, network timeout) likely return raw exceptions rather than 'Try search_pods() with partial name' style guidance.
GCP credentials and cluster info (project_id, cluster_name, zone in get_gke_cluster_metrics) are passed as parameters instead of server-side injection. If the agent logs or forwards these params, sensitive infrastructure details leak.
No visible permission checks or scope declarations. Tools like delete_resource, scale_deployment, and exec_pod_command can be invoked by any caller without RBAC validation. This violates the principle of least privilege.
Parameters like 'lines' (for log retrieval) and 'port' (for connectivity test) lack min/max bounds. LLMs can request 1 million log lines or test port 999999, causing DoS or timeouts.
suggest_troubleshooting and automate_remediation are vague about what 'analysis' and 'remediation' actually do. Are they heuristics? Machine learning? Static rules? Output format is not documented. LLMs cannot decide when to use these vs direct tools.
No evidence of pagination support in list_pods, get_deployments, or other bulk-retrieval tools. If a namespace has thousands of pods, the server may return all of them, bloating the LLM context and inviting timeout.
list_podsget_deploymentsget_service_status
Implement result limiting and pagination. Add 'limit' (default 20, max 100) and 'offset' parameters to list_pods, get_deployments, get_service_status. Return {items, total_count, has_more, next_offset}.
Clarify suggest_troubleshooting and automate_remediation. Document exactly what heuristics/rules are applied, what the output fields mean, and example scenarios. If they're AI-driven, state the model and version.
Add a dry_run parameter to create_deployment, scale_deployment, apply_yaml, and exec_pod_command. When true, return what would happen without executing. This lets agents preview changes.
Implement input validation and sanitization. Check pod_name against allowed characters (alphanumeric, dash, underscore), namespace against a whitelist, and command against shell injection patterns. Return actionable errors like 'Invalid pod_name: must match [a-z0-9-]+'.
Add audit logging: log every tool call with caller identity, tool name, parameters (sanitized), timestamp, and result. This enables compliance tracking and incident response.
Strip raw Kubernetes metadata from responses. Remove apiVersion, kind, uid, resourceVersion, and other API internals that the LLM never uses. Return only: name, namespace, status, image, replicas, containers, events (summary).