Add your description here
ExpenseTracker has 4 tools with explicit schema definitions, descriptions, and proper async implementations. However, multiple critical gaps reduce quality significantly. All tools lack parameter-level descriptions in the actual JSON schema (only tool descriptions exist). Output schemas are completely undocumented, callers cannot plan downstream operations. Error messages are generic ('Database error: ...') and provide no recovery guidance. The summarize tool's 'category' parameter defaults to null but lacks clear documentation of what that means. Tool naming follows verb_noun convention (add_expense, list_expenses, delete_expense, summarize) which is good, but descriptions are terse (10-34 chars) and lack WHEN/WHY guidance. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk profiles (WRITE, DESTRUCTIVE). The resource 'categories' is exposed but not referenced by any tool description. No pagination on list_expenses despite potentially unbounded result sets. Overall: adequate for a basic prototype but falls short of production standards.
Add a new expense entry to the database.
Delete an expense from the database using its ID.
List expense entries within an inclusive date range.
Summarize expenses by category within an inclusive date range.
All tools lack per-parameter descriptions in JSON schema. Only tool-level descriptions exist. This violates pattern:tool-description which requires EVERY parameter to have a non-empty description so LLMs can infer meaning from names alone.
Output schemas completely undocumented. add_expense returns {status, id, message} but spec is silent. list_expenses returns [{id, date, amount, category, subcategory, note}] but undocumented. delete_expense and summarize similar. LLMs cannot plan downstream operations or extract correct fields without code inspection.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Tools have explicit Risk tags in metadata (WRITE, READ_ONLY, DESTRUCTIVE) but spec does not expose these as annotations. LLM cannot infer safety/retry semantics from the spec alone.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
Error handling is generic and provides no recovery guidance. 'Database error: [exception text]' gives LLM no hint on retry, fallback, or next action. Violates pattern:recovery-guide.
Tool descriptions are terse (35-50 chars, vs baseline 194 chars). None answer WHEN/WHY to use the tool vs similar tools. add_expense vs summarize, how does an LLM choose? No guidance.
list_expenses has no pagination parameters (limit, offset, cursor) or result limit. Could return thousands of expenses for a multi-year user, exhausting context window. Violates pattern:paginated-result.
delete_expense is destructive but has no dry-run, confirmation, or undo mechanism. An LLM bug deletes real data permanently. Violates pattern:confirmation-request.
Parameter 'category' in summarize defaults to null with description 'Optional category filter'. Undocumented what null means, return all categories, or group by category? Violates pattern:constrained-input and rule on defaults.