Agent-first event tracking — ingest events, hand notifications to agents who relay via their own gateways
Cairo provides 11 tools with explicit schemas and descriptions. Strengths: all tools have descriptions (50-300 chars), all have input schemas with typed properties, and clear distinctions between read/write operations. Weaknesses: descriptions lack LLM-optimization guidance (e.g., when to call track_generation vs track_tool_call), parameter descriptions are minimal/generic, no output schemas documented, error handling lacks recovery guidance, and no tool annotations (readOnlyHint/destructiveHint). The tool definitions are well-structured but lack the depth needed for A-grade quality.
Capture an error event (Sentry-like). Supports stack traces, severity levels, and context.
End the current tracking session
List error groups (like Sentry issues). Filter by status.
Look up a user by email. Returns profile and recent activity.
Query CDP events. Returns recent events matching filters.
Start a new tracking session for an agent task
No output schemas documented. LLMs cannot infer what fields to expect from tool responses, breaking downstream planning and field-mapping logic.
Parameter descriptions are minimal and generic. 'Tool input parameters', 'Additional context', and 'Tags for filtering' do not explain constraints, formats, or when parameters are required. LLMs lack actionable guidance.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). LLMs cannot distinguish safe read operations from destructive writes. This is a current MCP 2026-07-28 pattern that should be present.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | 2026-07-28+ | v2 |
Get Cairo system health: database, uptime, event counts, error counts.
Report a routing or branching decision the agent made
Report an error encountered during agent execution
Report an LLM generation event including model, token counts, latency, and cost
Report a tool invocation with its name, success status, latency, and any errors
Missing error handling guidance. Tools like capture_error and track_error report failures but do not tell the LLM how to recover or what to do next.
Parameter optionality unclear. start_session lists 'task' and 'agentType' without 'required' declaration in the schema snippet; end_session's 'exitReason' requirement is unstated. LLMs guess which params are mandatory.
query_events has no documented limits or pagination. 'Returns recent events' is vague, does it return 10? 100? 1000? Without a stated cap and pagination, large result sets could blow context windows.
Tool naming could better disambiguate intent. 'track_error' vs 'capture_error' overlap conceptually (both report errors). Descriptions don't explain when to use each. Pattern suggests singular verb_noun, 'report_error' and 'capture_error' is clearer than 'track_error'.