MCP server for querying observability platforms including Prometheus, Alertmanager, Tempo (distributed tracing), Loki (log management), and OpenTelemetry Collector configuration assistance.
obs-mcp demonstrates solid definition quality with consistent verb-noun naming, meaningful descriptions, and proper parameter schemas across all 22 tools. All tools follow the action-first naming convention (list_, get_, execute_, search_) which helps LLMs quickly infer intent. Descriptions average ~80-120 chars, meeting the 10-1024 char baseline. All tools have typed input schemas with parameter descriptions. However, there are notable gaps: output schemas are not documented in the visible code, error handling guidance is minimal, and some parameter descriptions lack format/constraint details. The tool definitions appear well-structured but lack the polish of A-grade servers that document return types and include dependency hints.
Execute an instant query against Prometheus at a specific point in time
Execute a range query against Prometheus over a time interval
Get active alerts from Alertmanager
Get the JSON schema for an OpenTelemetry Collector component
Get all label names from Prometheus
Get all label values for a given label name
Get series data matching given label matchers
Output schemas not documented in source code. While all tools have input schemas with typed parameters, return types are not visible. This prevents LLMs from planning downstream chains and extracting correct fields.
Error handling guidance missing. Tools have no documented error conditions, recovery paths, or error classification (retryable vs user-fixable vs fatal). Example: execute_instant_query could fail with invalid PromQL but the description provides no guidance on what error the agent should expect or how to recover.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Get active silences from Alertmanager
Retrieve a trace by its ID from Tempo
Get available OpenTelemetry Collector versions and their components
Get all label names available in Loki
Get all values for a specific label in Loki
List available OpenTelemetry Collector components
List all Tempo instances available
List all Loki instances available
List all metrics available in the Prometheus instance
Execute a range query against Loki logs
Search for values of a specific tag in Tempo traces
Search for available trace tags in Tempo
Search for traces in Tempo using tag and duration filters
Show timeseries matching given label matchers
Validate an OpenTelemetry Collector configuration
Parameter constraints not specified in descriptions. Many parameters lack format details: 'time' in execute_instant_query says 'Unix timestamp or RFC3339 format' but should explicitly document which format is preferred and what happens if format is ambiguous. Parameters like 'query' in execute_range_query should mention character limits, complexity bounds, or rejected patterns.
Duplicate tool names. 'list_instances' appears twice (line 10 in Tempo context, line 15 in Loki context). LLMs cannot disambiguate identical names across different backends. Should be 'list_tempo_instances' and 'list_loki_instances' or namespace them differently.
Pagination not mentioned for list/search tools. Tools like list_metrics, search_traces, and query_range could return large result sets but no limit, offset, or pagination parameters are visible. Without pagination, large responses could exhaust context or breach API rate limits.
Parameter descriptions too generic for 'filter' objects. Tools get_alerts and get_silences accept 'filter' as object but describe it only as 'Optional filters for alert/silence selection' with no detail on what keys filter accepts, what values are valid, or format. This forces LLMs to guess structure.