An MCP server that analyzes Claude session data, agent performance, tool usage patterns, anomalies, context pressure, and cost attribution across projects.
The claudewatch server exposes 5 analytics tools with basic descriptions and sparse parameter documentation. Most tools lack input schema clarity, parameter descriptions are minimal or absent, and output schemas are not documented. Tool naming follows verb conventions (get_*), but the overall definition quality is below production standards. No evidence of error handling guidance, and schema validation is minimal.
Agent performance metrics across all session transcripts.
How much of the context window has been consumed. Reports token usage, compaction count, and utilization status (comfortable/filling/pressure/critical).
Break down token cost by tool type for a session. Answer 'which tool calls consumed most of my budget?' Defaults to most recent session.
CLAUDE.md effectiveness scores for each project with before/after session data.
Detect anomalous sessions for a project using z-score analysis against a historical baseline. Returns sessions with cost or friction deviating beyond the threshold.
Output schemas not documented for any tool. LLMs cannot plan downstream operations or extract fields without knowing the response structure.
Tools get_agent_performance and get_effectiveness have empty input schemas ({}). Unclear whether they accept filters, pagination, or optional parameters. If intentional, the empty schema is correct; if parameters exist but undocumented, schema is incomplete.
No pagination support declared. Tools returning analytics across 'all session transcripts' or 'each project' likely produce large result sets. No limit, offset, or next_cursor parameters visible, context window risks.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
No error handling guidance. Tools do not document recovery actions (e.g., 'if session not found, try get_agent_performance to discover available sessions'). Error responses would not guide LLM's next step.
Parameter descriptions are minimal. 'project' parameter in get_project_anomalies is described as 'Project name (e.g. 'commitmux'). Omit to use the current session's project.', this is good, but 'threshold' lacks range guidance. No minimum/maximum specified for the z-score threshold (default 2.0).
Tool descriptions do not explain WHEN to use them relative to each other. All five tools are read-only analytics, the server provides no guidance on recommended analysis sequence (e.g., 'start with get_agent_performance to identify trends, then drill into get_project_anomalies').