Query agent telemetry and AI/ML features from Splunk or Django backend. Supports natural language search, anomaly detection, failure analysis, and alert summarization.
AgentGuard MCP server has 8 tools with basic descriptions and input schemas, but significant gaps in parameter documentation, output schema clarity, and error handling guidance. All tools are READ_ONLY (good for safety), but descriptions lack LLM-optimized detail. Parameter descriptions are present but generic. No tool annotations (readOnlyHint, etc.). Output schemas are not formally documented, responses are inferred from code. Error handling is minimal; no recovery guidance. Naming is verb-forward (search_, explain_, check_) which is good, but some names are vague (nl_search, anomaly_detection lack specificity about what anomalies or what NL→SPL fallback means).
Aggregate pass/fail rates and top error types by agent_name.
Summarize fired AgentGuard Splunk alerts (or FAILED-span proxy when alert log unavailable).
Detect latency/event anomalies using MLTK DensityFunction or built-in anomalydetection.
Report availability of Splunk AI Assistant, MLTK, and fallback modes.
Return span details and error context for a failed agent trace_id.
Compute failure rates and error types per agent over a time window.
Output schemas not formally documented. Code infers responses (e.g., search_agent_traces returns {source, results, spl, warning}) but no JSON Schema or structured type hints in tool definitions. LLMs cannot plan downstream calls without knowing what fields to expect.
Parameter descriptions are generic and lack constraints. E.g., 'minutes' has no min/max bounds (default 15 or 60, but no stated range). 'status' accepts 'SUCCESS/FAILED/TIMEOUT' but is not declared as an enum in schema. LLMs may pass invalid values.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). All tools are READ_ONLY but this is not declared in the tool definition. Clients cannot infer safety properties without annotations.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | 2026-07-28+ | v2 |
Convert natural language to SPL and run the search. Falls back to rule-based SPL if AI Assistant is unavailable.
Search recent agent spans in Splunk. Filter by agent_name and status (SUCCESS/FAILED/TIMEOUT).
Error handling lacks recovery guidance. Code catches SearchError and falls back to Django, but tool responses do not tell LLM what to do on failure. E.g., 'Splunk unavailable; try Django backend' is not returned to the agent.
Vague tool names reduce clarity. 'nl_search' does not convey 'convert natural language to Splunk query'. 'anomaly_detection' does not specify latency vs event anomalies. LLMs may misselect tools when names are ambiguous.