Tamper-evident audit trail for OpenClaw agent runs. A flight recorder system that records, replays, and verifies agent actions with cryptographic hash chain integrity.
Clawprint demonstrates solid definition quality with consistent naming, clear descriptions, and complete input schemas across all 7 tools. All tools follow verb_noun naming conventions (list_, get_, verify_, extract, replay_), have substantive descriptions (70-100+ chars), and include properly typed input parameters with descriptions. The server is read-only, which simplifies error handling concerns. However, output schemas are not explicitly documented in the source code provided, preventing higher scores. The tool suite is well-composed for audit trail operations: list/get/verify/replay form a coherent query/analysis pipeline. Parameter naming is consistent and self-documenting (run_id, output_dir, limit, offset, kind). Default values are sensible (limit=50-100, offset=0). No security risks identified, all tools are read-only and accept filesystem paths (no secrets exposed as parameters). The main limitation is the absence of explicit output schema documentation, which prevents confident scoring of the 'schema' dimension and forces a conservative ceiling.
List agent conversation runs from the continuous ledger
Retrieve events from a recorded run with optional filtering
Retrieve detailed information about a specific recorded run
Extract and list all tool calls from a recorded run
List all recorded agent runs with statistics
Generate a transcript or replay of a recorded run
Verify the cryptographic integrity of a recorded run
Output schemas not explicitly documented in source code. Tool definitions show input schemas clearly, but return types/response structures are not visible in mcp.rs. LLMs cannot reliably plan downstream tool chains or extract specific fields without knowing what each tool returns.
Parameter 'kind' in get_events is a free-form string filter, not an enum. Although the description gives examples (AGENT_EVENT, TOOL_CALL, OUTPUT_CHUNK), valid values should be declared as an enum constraint to prevent LLMs from hallucinating invalid event types.
Parameter 'format' in replay_run is correctly an enum, but tool descriptions do not explain what the transcript vs. json formats contain or when to choose each. LLMs need guidance on which format serves which use case.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No pagination documentation. Tools like list_runs and get_agent_runs accept limit/offset but do not document whether they return a total_count, next_cursor, or has_more flag. Large result sets risk context window exhaustion without clear pagination behavior.
Error handling patterns not visible in source. Tool descriptions do not indicate what errors are possible, whether they are retryable, or what recovery actions are available. LLMs will not know whether to retry, ask the user, or escalate.