MCP server for ScanWarp production monitoring, providing tools to query incidents, events, traces, analytics, and application status
ScanWarp MCP server has 10 well-named tools with consistent verb_noun patterns (get_*, resolve_*) and documented descriptions. However, significant gaps exist in parameter descriptions, output schema documentation, and error handling guidance. All tools have type definitions for input parameters and enums where appropriate (e.g., status, severity, granularity), which is a strength. Parameter descriptions are present but often generic. No output schemas are documented in the source, forcing LLMs to infer response structures. The single write operation (resolve_incident) lacks confirmation/dry-run patterns. Error handling appears basic with no recovery guidance visible.
Query application metrics and analytics data with time range and dimension filtering
Get current application health status, including active incidents, monitor health, and provider status
Analyze error patterns and anomalies with grouping and correlation
Query events with filtering by type, source, severity, and limit
Get the AI-generated fix prompt for an incident that is ready to use in Claude or other AI tools
Get detailed information about a specific incident, including root cause diagnosis, suggested fix, fix prompt, event timeline, and trace data
List incidents with filtering by status, severity, and limit
Output schemas not documented. Tools return complex nested structures (incidents with root cause analysis, event timelines, trace waterfalls, service dependency graphs) but no field-level schema is visible in source. LLMs cannot reason about response structure, forcing inference and risking downstream tool failures.
Write operation (resolve_incident) lacks confirmation/dry-run pattern. Agents may accidentally resolve critical incidents. No explicit 'this modifies state' warning in description. No undo/compensation tool provided.
Parameter descriptions are minimal and lack actionable format/constraint guidance. E.g., 'projectId' lacks context on format (UUID vs alphanumeric); 'timeWindow' in get_error_analysis states 'e.g. 1h, 24h, 7d' as example rather than formal constraint; 'dimensions' in get_analytics has no guidance on valid keys (service, endpoint, status_code listed but not exhaustive).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 62 | - | v1 |
Get the service dependency graph and topology from OpenTelemetry span data
Get OpenTelemetry trace data and build a waterfall visualization for a specific trace ID
Mark an incident as resolved and send resolution notifications to configured channels
Error handling not visible. No recovery guidance documented for common failures (incident not found, invalid time range, insufficient permissions, service unavailable). Pattern: recovery-guide missing.
Tool descriptions do not clarify when to select between similar read tools. E.g., get_events vs get_incidents overlap in monitoring scenarios. No dependency hints (e.g., 'Call get_incidents first, then get_incident_detail for root cause'). Agents waste cycles guessing which to call.
Pagination and result limits not documented. No visible limit parameters for list-like tools (get_incidents, get_events, get_analytics). Returning unbounded results risks context window exhaustion. Baseline expectation: 20-50 item default with configurable pagination.
get_analytics 'dimensions' parameter is object-typed with no schema. LLMs cannot determine valid keys or values. Should be structured (typed object with defined keys) or converted to multiple specific parameters (filter_by_service, filter_by_endpoint, filter_by_status_code).