An MCP server for tracking personal expenses with SQLite — supports full CRUD, budgets, analytics, and export.
This server has 16 total tools (with 4 duplicates: add_expense appears twice, list_expenses appears twice), reducing logical tool count to 12. Tools have clear action-verb naming (add_, list_, get_, update_, delete_, search_, set_, remove_, export_) which is good. Descriptions are present for all tools and are 10-200+ characters. However, critical issues reduce the score: (1) Tool duplication, add_expense and list_expenses are defined twice with different input schemas, creating ambiguity about which definition the server actually registers. (2) Input schemas are complete with type annotations and descriptions for most tools, but output schemas are entirely undocumented, the rubric requires documentation of what fields and structure tools return so LLMs can chain them. (3) Some parameters lack detailed constraints: add_expense accepts a 'description' string with no length/format guidance; set_budget accepts 'monthly_limit' (a number) with no min/max bounds (users might pass 0 or negative values). (4) Error handling is not visible in the provided code, no guidance on retryability, user-fixable errors, or recovery steps. (5) The export_expenses tool claims to return 'CSV or JSON text' but the actual response structure is not documented. (6) Several tools like update_expense say 'Only provide fields you want to change' but do not document which fields are optional or what partial updates return. Overall, the naming and basic schemas are solid, but missing output documentation, sparse parameter constraints, and duplicate definitions keep this in the 'fair' range.
Add a new expense. Returns the created expense record with its ID.
Add a new expense entry to the database.
Permanently delete an expense by ID.
Export expenses as CSV or JSON text.
See how much of each budget has been spent this month (or a specified month).
Break down spending by category with totals, counts, averages, and percentages.
Get a single expense by its ID.
Output schemas entirely undocumented across all tools. LLMs cannot know what fields and structure tools return, forcing them to make assumptions about response chaining and downstream tool parameters.
Tool duplication: add_expense is defined twice with different input schemas (first with optional date/payment_method/tags/description, second with date/category/subcategory/note). list_expenses is defined twice with different filtering capabilities (first with rich filters, second with only date range). Causes ambiguity about which schema is actually registered and which tool to call.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
Get a summary of expenses: total count, total amount, averages, top category, etc.
See month-over-month spending trends.
List expense entries within an inclusive date range.
List expenses with optional filters. Returns matching expenses sorted by date (newest first).
Remove a budget for a category.
Search expenses by keyword in description, category, or tags.
Set a monthly spending budget for a category. Updates if already exists.
Summarize expenses by category within an inclusive date range.
Update an existing expense. Only provide fields you want to change.
Numeric parameters lack bounds. set_budget's 'monthly_limit' has no min constraint (should be > 0). get_monthly_trend's 'months' has no max (should cap at e.g. 120). list_expenses' 'limit' has default 50 but no stated maximum. Unbounded numbers invite LLMs to pass absurd values.
Category parameter is not enum-constrained in set_budget, remove_budget, and filters. Tools like add_expense suggest examples ('food, transport, entertainment...') but provide no enum. This forces LLMs to guess at valid category values and increases likelihood of invalid calls.
Destructive operations (delete_expense, remove_budget) lack confirmation or dry-run support. Descriptions do not emphasize permanence or offer recovery guidance. No mention of confirmation flow or undo capability.
Error handling and recovery guidance not visible in tool definitions. No descriptions of what errors might occur, which are retryable, or how an LLM should respond to failures. This violates the recovery-guide pattern.
Redundant tools: get_expense_summary and summarize both appear to summarize expenses by category. Similarly, multiple list_expenses variants with different filter sets. LLMs will struggle to disambiguate and may call the wrong tool.
Parameter naming inconsistency: add_expense uses 'description' and 'payment_method', but the duplicate add_expense (from Python) uses 'note' and no payment_method. List_expenses uses generic 'limit' for pagination, but export_expenses does not support pagination at all. Inconsistent naming confuses LLMs about which parameter names to use.
Date format constraints are mentioned in descriptions ('YYYY-MM-DD') but not enforced in schema patterns. LLMs may pass dates in other formats, leading to parsing failures. Descriptions should state: 'ISO 8601 format required (YYYY-MM-DD), e.g. 2024-01-15'.