Record, replay and fork debugger for AI agents. Exposes six MCP tools to query, replay, and compare agent runs.
OrcaReplay MCP server has well-crafted tool descriptions (avg 180 chars, within 10-1024 baseline) and clear naming conventions (verb_noun: orca_list_runs, orca_show_run, etc.). All 6 tools have descriptions explaining WHAT they do and WHEN to use them. Input schemas are present with type definitions. However, parameter descriptions are sparse, most parameters lack detailed guidance on format, constraints, or dependencies. Output schemas are not documented. Error handling is minimal; no recovery guidance or actionable error messages visible. Tool composition is sound (each does one thing), but missing pagination guidance for list_runs and missing dry-run/confirmation for destructive tools (orca_replay, orca_compare). Risk annotations (READ_ONLY, WRITE, IRREVERSIBLE) are present but not formalized as tool annotations in the schema.
The points in a run a fork can start from — where the conversation prefix is complete and the workspace was snapshotted. Use before orca_compare to pick a fork point.
Fork one recorded run onto several models from the same checkpoint — same files, same conversation prefix — and grade each with a command you choose. SPENDS REAL TOKENS and reaches the network: every model named is actually called. Ask before using it.
What caused what in a run, as a list of edges. Each edge says which event produced which, and whether it is `recorded` — the recorder watched it happen and wrote it into the trace — or `inferred`, meaning this derived it just now from the rule it names and the trace does not vouch for it. Pass `to` to get only the chain that produced one event, which is the shape of an answer to "why did this fail" rather than "what happened".
List every agent run recorded in this project, newest first, with the run it was forked from where there is one. Start here when you do not already know which run to look at.
Re-run a recording exactly, with the network blocked and no tokens spent, and report what could not be reproduced: divergences, and requests the recording could not serve. Free and repeatable. It shows the recorded decisions still reproduce; it cannot show a fresh run would fail the same way, because the model is not asked again — its recorded responses are served back.
Parameter descriptions are minimal or absent. 'run' parameter appears in 5 tools but lacks guidance on format (UUID? filename? timestamp?). 'models' in orca_compare has no description of valid model IDs or how to discover them.
Output schemas are not documented. LLMs cannot plan downstream calls or extract required fields. orca_list_runs should document return structure (array of run objects with fields: id, timestamp, forked_from, etc.).
Destructive/irreversible tools (orca_replay, orca_compare) lack dry-run or confirmation steps. orca_compare explicitly 'SPENDS REAL TOKENS' but no mechanism to preview cost or confirm before execution.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
The full timeline of one run: every model turn with its token counts and stop reason, every tool call with its arguments and result, every shell command with its exit code, and every file the run changed. This is what tells you why an agent did something, rather than what it cost.
Risk annotations (READ_ONLY, WRITE, IRREVERSIBLE) are metadata comments, not formalized as tool annotations in the MCP schema. Should use readOnlyHint, destructiveHint, idempotentHint in tool definitions.
No pagination guidance for orca_list_runs. If many runs exist, response could exceed context window. Should document limit, offset/cursor, and total count in output schema.