MCP server for Prometheus metrics and queries that provides access to Prometheus metrics through standardized MCP interfaces, allowing AI assistants to execute PromQL queries, discover metrics, and analyze metrics data
mcp-prometheus demonstrates solid fundamentals across 15 well-named tools with consistent action-verb prefixes (execute_, get_, query_, check_). Descriptions are present and adequately detailed (mostly 60-150 chars, within production baseline of 34-392 chars). All tools follow a verb_noun naming convention. Input schemas are properly typed with JSON Schema format. However, several gaps prevent a higher score: (1) Output schemas are NOT documented in the provided source, only input schemas are visible. (2) Parameter descriptions lack specificity about format constraints, ranges, and examples of valid values (e.g., 'timeout' accepts strings like '30s' or '5m', but this is only hinted in description, not formalized as a pattern or enum). (3) No visible error handling guidance, tools do not document what errors to expect or how to recover. (4) Pagination/limiting strategy is inconsistent: 'execute_query' and 'execute_range_query' expose an 'unlimited' parameter with a performance warning, which is an anti-pattern, caps should be enforced server-side, not exposed to the LLM. (5) The 'org_id' parameter is repeated across all 15 tools, suggesting opportunity for server-level configuration rather than per-call injection. (6) Several tools ('get_config', 'get_flags') expose sensitive Prometheus internals without clear guidance on what the agent should do with that information.
Check if Prometheus server is ready
Execute an instant PromQL query against Prometheus
Execute a range PromQL query against Prometheus over a time interval
Get all active alerts from Prometheus
Get Prometheus server build information
Get Prometheus server configuration
Get Prometheus server command-line flags
Output schemas not documented. The source shows input parameter schemas but no declaration of what fields each tool returns. LLMs cannot plan downstream calls or extract results without knowing the response structure.
Parameter 'unlimited' exposes performance risk as an LLM-controllable flag. Anti-pattern: limits should be enforced server-side. 'unlimited': 'true' allows agents to request unbounded results, risking context window exhaustion and API abuse. Seen in execute_query and execute_range_query.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 71 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 55 | - | v1 |
Get all label names from Prometheus
Get all values for a specific label name in Prometheus
Get metadata (type, help, unit) for Prometheus metrics matching a name pattern
Get alert and recording rules from Prometheus
Get time series matching label matchers from Prometheus
Get all Prometheus scrape targets (active and dropped)
Get TSDB statistics from Prometheus
Query exemplars from Prometheus
Format constraints not formalized. Parameters like 'timeout', 'start', 'end', 'step' accept strings in specific formats ('30s', '5m', RFC3339 dates). Descriptions mention examples but lack formal regex patterns or explicit enum lists. LLMs may pass invalid formats.
No error recovery guidance. Tools do not document expected error conditions (e.g., 'If query_timeout is exceeded, retry with a larger timeout', 'If metric_name not found, call get_metric_metadata with partial name'). Errors are likely raw HTTP status codes with no actionable next steps.
Repetitive 'org_id' parameter across all 15 tools. This suggests the org context should be configured at the server/client level rather than repeated in every call. Current design increases parameter count, wastes tokens, and invites mismatches.
Sensitive introspection tools lack use-case guidance. Tools like 'get_config', 'get_flags', 'get_build_info' expose internal Prometheus state. Descriptions do not explain when an agent should call these or what to do with the returned data. Risks over-exposure without clear value.