Enhanced MCP server for Prometheus observability with monitoring, alerting, RCA, incident management, and natural language query interface for DevOps/SRE teams
The server defines 16 monitoring tools with fastmcp. Most tools have basic descriptions and schemas visible in server_test.py, but several critical quality issues reduce the score. Naming follows action-verb patterns well (get_, detect_, identify_), and all tools have descriptions. However, schemas are incomplete or missing for many tools, descriptions lack depth or actionable context, and error handling is minimal. Parameter descriptions are sparse. The server lacks output schema documentation, pagination guidance, and recovery hints. Most tools would benefit from richer descriptions that explain WHEN and WHY to use them, not just what they do.
Correlate two metrics to identify relationships.
Get the current metric values for a list of pods.
Get overall cluster health status.
Detect pods in crash loops.
Detect pod metric anomalies using statistical methods.
Get resource usage summary per namespace.
Get summary of node conditions.
Output schemas not documented. No tool explicitly declares what fields it returns. LLMs cannot infer downstream data structure and cannot chain tool calls effectively. For example, current_metric_for_pods returns {'metric', 'pods_current_cpu_per_prometheus', 'timestamp'} but the structure of 'pods_current_cpu_per_prometheus' (is it a dict of dicts? a list?) is not formally documented.
Descriptions for tools with empty input schemas (describe_cluster_health, pod_status_summary, node_condition_summary) lack actionable context. Descriptions are generic ('Get overall cluster health status', 'Get summary of pod statuses across cluster') and do not explain WHEN to call them, what data they reveal, or how to act on the results.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | <=2025-11-25 | v2 |
Get node disk usage for important mount points.
Get timeline of events for a specific pod.
Get network I/O statistics for specified pods.
Analyze pod restart trends.
Get summary of pod statuses across cluster.
Identify pods exceeding CPU threshold.
Retrieve recent pod events.
Find nodes with highest disk pressure.
Find top N pods by metric value over a time window.
No error handling or recovery guidance in tool definitions. Tools return raw {'error': 'message'} but do not categorize errors as retryable, user-fixable, or fatal. No guidance on what the LLM should do next (e.g., 'Try calling search_pod() first', 'Check Prometheus connectivity'). For example, current_metric_for_pods returns {'error': 'Prometheus client not initialized'} but does not tell the LLM how to fix this.
Parameter descriptions are sparse or missing context. For example, 'window' parameter (e.g., in top_n_pods_by_metric) is documented as 'Time window (e.g., "30m", "1h")' but does not specify the required format (does '1hour' work? '30 minutes'?) or valid range. Descriptions should state the format explicitly, e.g., 'Time window in Prometheus duration format (e.g., "30m", "1h", "7d"; valid range 1m - 30d)'.
No pagination or result limits documented. Tools like top_n_pods_by_metric, recent_pod_events, and correlate_metrics do not declare max result counts or offer offset/limit parameters. A query returning 10,000 events could blow the context window. Descriptions should state 'Returns up to 50 items; use limit and offset for pagination.'
Sensitive configuration (Prometheus base_url, headers, SSL settings) loaded from yaml config file without secrets injection pattern. If the config contains API keys or auth tokens, they could be logged or exposed. No evidence of environment-based secret injection or vault integration.
Tool descriptions do not explain prerequisites or dependencies. For example, detect_pod_anomalies uses z_threshold but does not explain what z-score is, why 3 is the default, or when to adjust it. Descriptions should guide LLMs: 'Z-score threshold for detecting statistical outliers. Default 3 flags values >3 standard deviations from mean. Increase to 4-5 for less aggressive detection.'
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) declared. All 16 tools are read-only and idempotent, but this is not formally declared via tool annotations. MCP 2026-07-28 encourages annotations to help LLMs reason about side effects and retry safety.
No discovery or dependency hints in descriptions. For example, if detect_pod_anomalies requires pod_name or namespace to be specified by the user but current_metric_for_pods lists available pods, the descriptions should link them: 'Call current_metric_for_pods first to discover available pods, then pass pod_names here.'
namespace parameter in current_metric_for_pods and other tools is documented as 'Optional namespace filter' but is never used in the implementation (not included in the PromQL query). This is misleading, either remove the parameter or implement the filtering.