Runbook-driven backend incident investigation framework for AI agents
The 'investigate' tool has a clear, actionable description and well-defined input schema with proper enum constraints. However, the server exposes only a single tool with no output schema documentation, limiting composability. The tool description (186 chars) falls within the acceptable range (10-1024 chars), and all four input parameters have types and descriptions. Zod validation is present server-side. The lack of pagination, result limits, and explicit error recovery guidance are notable gaps for a tool that performs complex backend investigation.
Performs Runbook-driven automated investigation for backend incidents, returning a structured incident report (root cause, evidence, recommended next actions).
No output schema documented for the 'investigate' tool. LLMs cannot anticipate the structure of the incident report response, forcing them to parse unstructured JSON and making downstream chaining difficult.
Tool description lacks context on when to use this tool vs alternatives, what prerequisites are needed (e.g., configured runbooks, adapter setup), and recovery guidance if no suitable runbook is found.
Error handling is minimal. The tool throws a generic error if no runbook supports the context_type, but does not guide the LLM toward recovery (e.g., 'Supported context types are: trace_id, request_id, order_id, task_id, message_id, user_id. Try one of these instead.').
No limit on investigation scope, complexity, or execution time documented. A runaway agent calling investigate repeatedly on the same context_id could generate excessive backend load without rate limiting or timeout guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 64 | 2026-07-28+ | v2 |
Input parameters 'symptom' and 'expected' accept free-form strings with only length validation (min 1 char). No guidance on format, detail level, or examples. LLMs may pass vague symptoms like 'bad' or 'wrong', reducing investigation quality.
Tool name 'investigate' is generic and could be confused with other debug/logging tools. A more specific name like 'investigate_incident' or 'run_incident_runbook' would clarify that this tool specifically performs runbook-driven incident investigation.