Prometheus/Thanos metrics query server for AI assistants. Provides tools to query metrics, discover metric names, inspect labels, and run pre-configured queries.
kubectl-metrics MCP server provides 2 tools for Prometheus/Thanos metric queries. The tools have reasonable length descriptions (180+ chars each) and cover distinct responsibilities (query execution vs. help/discovery). However, the input schemas are severely underspecified: the 'flags' parameter is typed as a generic object with no formal schema constraints, no enum definitions for subcommands, no type restrictions on individual flags, and no description of flag semantics. Parameter descriptions exist but are embedded in the tool description rather than formally documented in the schema. Output schemas are not documented at all. Error handling guidance is absent. The metrics_help tool is essentially a documentation tool, not an action tool, which is a composition smell. The server follows Go SDK best practices for structure but lacks the rigor expected of production tools.
Get detailed help for metrics_read subcommands and PromQL query language. WHEN TO USE: Before calling metrics_read, call metrics_help("<command>") to learn the available flags and their meaning. Call metrics_help("promql") for PromQL syntax reference. Commands: query, query_range, discover, labels, preset, promql Omit command for an overview of all subcommands and available presets.
Query Prometheus / Thanos metrics. Use metrics_help for flag details and PromQL reference. Subcommands (pass as "command"): query Instant PromQL query (flags: query, output, filename, name, group_by, no_pivot, selector) query_range Range PromQL query over a time window (flags: query, name, start, end, step, output, filename, group_by, no_pivot, selector) Supports multiple queries: pass query and name as arrays (e.g. query: ["rate(container_cpu_usage_seconds_total[5m])", "container_memory_working_set_bytes"], name: ["cpu", "mem"]). Each query's results are labeled with the corresponding name (auto-generated q1, q2, ... if omitted). Pass filename to write output to a temp file and return only a summary with the full path (e.g. filename: "data.tsv"). Useful for large results intended for gnuplot or other tools. discover List available metric names (flags: keyword, group_by_prefix) labels List labels or label sets for a metric (flags: metric) preset Run a pre-configured named query (flags: name, namespace, start, end, step, output, filename, group_by, no_pivot, selector) Every preset works as both instant (default) and range query. Pass start to get a time-series trend. Range queries use a pivot table by default (one column per label combination). Set no_pivot: true to revert to the traditional row-per-sample format. Use selector to filter results by labels post-query (e.g. "namespace=prod,pod=~nginx.*"). Supported operators: = (equal), != (not equal), =~ (regex), !~ (negative regex).
Input schema 'flags' parameter is untyped object with no formal constraints. No enum for 'command' values, no type or description for individual flags (query, output, filename, name, group_by, etc.). LLMs cannot validate inputs or discover valid flags without calling metrics_help first.
Output schemas not documented. LLMs do not know what fields to expect from metrics_read (instant query result structure, range query time-series format, discover result fields, labels result format, etc.). This forces LLMs to infer structure, risking parsing errors.
metrics_help is a documentation/discovery tool, not an action tool. It does not modify state or retrieve data, it returns formatted help text. This violates single-responsibility: agents should discover available metrics via metrics_read discover, not call a separate help tool. Consider moving help into inline descriptions and making discover the primary discovery mechanism.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 40 | - | v1 |
No error handling guidance. Tool descriptions do not indicate what errors agents should expect (invalid PromQL, no metrics found, Prometheus unreachable, timeout, etc.) or how to recover. Error responses are likely raw server errors that LLMs cannot act on.
Subcommand-specific flag documentation is embedded in the tool description as free text. No formal schema separates query flags from query_range flags from discover flags. LLMs must parse natural language to understand which flags apply to which subcommand.
No pagination support documented. If metrics_read discover or labels returns hundreds of items, no limit or offset mechanism is visible. Tool description mentions large results can be written to files, but pagination API is not formalized.
Tool naming is reasonable (verb + noun) but not fully verb-first. 'metrics_read' is clear; 'metrics_help' is ambiguous (is it help about metrics, or help from metrics?). Better: 'get_metrics' and 'describe_metrics_schema'.