A commitment tracker and cognitive state manager that detects promises, manages follow-ups via nudges, and applies constitutional safety gates to prevent burnout. Integrates with Claude AI for commitment detection and agentic tool use.
OTTO has well-structured tool definitions with consistent naming (verb_noun pattern: otto_list_*, otto_add_*, otto_mark_*) and clear descriptions (avg 150 chars, within 10-1024 baseline). All 10 tools have input schemas with proper JSON Schema format and type declarations. However, output schemas are completely undocumented, no tool describes what it returns, forcing LLMs to infer structure. Parameter descriptions are present but minimal (avg 40 chars); several lack actionable constraints (e.g., 'duration' in otto_snooze_commitment has no format spec like '30m, 4h, 2d' in description). Error handling is absent, no recovery guidance, no categorization of retryable vs fatal errors. No tool annotations (readOnlyHint/destructiveHint) despite clear risk levels (READ_ONLY vs WRITE). Missing idempotency guidance for write operations.
Add a new commitment. Requires the commitment text. Optionally specify who it's to and a deadline (YYYY-MM-DD).
Add a work-in-progress note to a commitment. Notes are appended and help track progress over time.
Get your current cognitive state: energy level, burnout, momentum, nudge counters, and suppression count.
Show commitment statistics: active count, done count, parked count, and average follow-ups before completion.
List commitments. By default shows active commitments ordered by deadline. Use 'filter' to change: 'active' (default), 'due' (overdue only), or 'all' (including done/parked).
Mark a commitment as done. Takes the short numeric ID shown in otto_list_commitments output.
No output schemas documented. Tools return JSON but LLMs cannot plan downstream calls without knowing response structure (e.g., otto_list_commitments returns what fields? otto_get_stats returns what metrics?). This violates pattern:tool and forces LLMs to guess.
Missing tool annotations. All tools have explicit risk levels (READ_ONLY vs WRITE) but no readOnlyHint or destructiveHint in schema. LLMs cannot distinguish safe from unsafe operations without annotations.
Parameter 'duration' in otto_snooze_commitment lacks format specification in description. Code shows '30m, 4h, 2d' format but description only says 'Snooze duration (e.g. 30m, 4h, 2d).', example values in descriptions are anti-pattern; use formal constraints instead.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | <=2025-11-25 | v2 |
Park a commitment guilt-free. It won't generate nudges but stays in history. Takes the short numeric ID.
Run the follow-up nudge check. Returns nudge messages for overdue and stale commitments. Constitutional gating may suppress this if you're in a depleted state.
Set your energy level. Valid values: high, medium, low, depleted. This affects how OTTO behaves -- in depleted state, nudges are suppressed by the constitutional layer.
Snooze a commitment for a duration. The commitment won't appear in nudges until the snooze expires. Duration format: 30m (minutes), 4h (hours), 2d (days).
No error handling guidance. Tools provide no recovery hints (e.g., 'short_id not found, call otto_list_commitments to see valid IDs'). Agents cannot self-correct on failures.
Idempotency not documented. Write operations (otto_add_commitment, otto_mark_done, otto_snooze_commitment) do not state whether repeated calls with same input are safe. Agents may retry on ambiguous failures, causing duplicates.