Remote MCP server for expense tracking with async support
The server implements 8 tools with consistent verb_noun naming and basic schemas. However, descriptions lack depth for LLM optimization, parameter descriptions are minimal or missing, output schemas are undocumented, and error handling provides no recovery guidance. This is a solid foundation that falls short of production-grade agent integration.
Add a new expense to the database.
Delete an expense by its ID.
Check budget vs actual spending for current month.
Get monthly spending trends over time.
Get spending summary grouped by category.
List expenses with optional filters for category and date range.
Search expenses by description keyword.
Missing output schemas for all tools. LLMs cannot predict response structure, plan downstream calls, or extract required fields (e.g., expense_id from add_expense). This violates the core pattern that documented output drives composition.
Destructive tool (delete_expense) description does not state irreversibility.
No parameter constraints on numeric fields. 'amount' in add_expense has no min (can be negative? zero?). 'limit' in list_expenses and search_expenses have no documented max (can LLM pass 1M?). 'months' in get_monthly_trend unbounded.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Update an existing expense's details.
Error handling provides no recovery guidance. Code returns bare success/false flags and generic messages ('Expense with ID X not found').
Parameter descriptions are minimal or absent. Example: 'start_date' in list_expenses is described only as 'Start date for filtering (YYYY-MM-DD format)', does not explain behavior if end_date is omitted, or interaction with 'category' filter.
Tool descriptions lack depth for LLM selection. Average length is ~70 chars (below 194-char baseline for A-grade tools). Example: 'Get spending summary grouped by category' does not explain when to call vs list_expenses, what fields summary contains, or whether filtering by date is useful.
Idempotency not documented. update_expense accepts partial updates; unclear if calling twice with same params is safe. add_expense creates new record each time (non-idempotent by design), but this is not stated.
Missing confirmation/dry-run for irreversible operations. delete_expense has no preview or confirmation step.
Field naming inconsistency may confuse LLM chaining. update_expense returns 'updates' dict (the fields changed), but no 'expense_id' in response. If next tool needs to fetch the updated record, LLM must retain expense_id from the request context rather than extracting from response. Per pattern, chaining requires response to include next tool's required params.