A Model Context Protocol (MCP) server that provides access to Prometheus metrics backends, enabling AI models to query metrics, execute PromQL queries, and list available metrics from multiple Prometheus/Mimir/Thanos backends
Prometheus MCP presents 3 tools with mostly complete schemas and descriptions, but multiple issues prevent a higher score. All tools follow verb_noun naming (prometheus_query, prometheus_range_query, prometheus_list_metrics) which is correct. Descriptions are present (72-88 chars) and address the core function, but fall short of LLM-optimized guidance on WHEN to use each tool vs. alternatives. Schemas include typed parameters with descriptions, but lack explicit output/response documentation, critical for chaining and LLM planning. Parameter validation is implicit (e.g., RFC3339 format mentioned in description) rather than formally expressed via JSON Schema patterns/formats. Error handling is not visible in the provided code snippet, and no recovery guidance is evident. The server accepts backend names, org_id for multi-tenancy, and sensible defaults (step='1m', limit=100), which are strengths. However, the lack of documented output schemas, absence of pagination documentation for list_metrics, and no visible error categorization or recovery patterns significantly limit production readiness.
List all available metrics from a metrics backend
Execute a PromQL query against a metrics backend
Execute a PromQL range query against a metrics backend
Output schemas not documented. The tools return Prometheus responses (instant vectors, range vectors, metric metadata) but no JSON Schema is provided for LLM planning. LLMs cannot determine what fields to extract or how to chain results to downstream tools.
Parameter format constraints not formalized. Descriptions mention 'RFC3339 format' and 'glob pattern', but JSON Schema does not declare format: 'date-time', format: 'date-time', or pattern: regex. LLMs cannot validate input without formal constraints in the schema.
No error handling guidance visible. The code snippet truncates at error handling, and no documentation shows how the server categorizes errors (retryable, user-fixable, fatal) or provides recovery hints. An LLM receiving a PromQL syntax error has no guidance on how to fix it.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
prometheus_list_metrics lacks pagination documentation. The schema includes limit and offset parameters, but the description does not explain pagination semantics: does offset apply before or after filtering? Is there a total_count in the response? Do results have a next_cursor?
Descriptions lack WHEN guidance. Each tool describes WHAT it does, but not WHEN to use prometheus_query vs prometheus_range_query. An LLM with a user asking 'Get CPU usage over the last hour' must infer to call prometheus_range_query; the descriptions do not make this distinction explicit.
backend parameter vague. Description says 'Available: [list of backends]' but does not show the actual list. An LLM cannot determine valid values without seeing concrete examples or an enum. If not provided, the tool 'defaults to single backend', but what is that default? Undocumented behavior invites wrong tool calls.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). All three tools are read-only, but this is not declared in the schema. Modern MCP servers use tool annotations to signal to clients which tools are safe to retry, which are destructive, and which are idempotent, enabling smarter LLM planning.