A Model Context Protocol server for managing personal expenses with support for PostgreSQL and SQLite backends, multiple user accounts, expense tracking, categorization, and summarization
This expense tracker MCP server has significant quality gaps across naming, descriptions, and parameter documentation. While tool names follow verb-noun conventions (add_expense, list_expenses, etc.), parameter descriptions are often missing or too terse. Input schemas are partially defined but lack rigor in several tools. Output documentation is absent, the tools return unstructured strings or raw database rows rather than documented schemas. Error handling is minimal; many tools return plain exception strings without guidance for recovery. The server implements basic functionality but falls well short of production-grade quality expected of an A/B-tier tool.
Add a new expense to the database
Delete an expense by ID.
List expenses within the inclusive Date Range (YYYY-MM-DD)
Summarize expenses by category within a date range
Update an expense by ID. Only fields that are provided will be updated.
Output schemas are undocumented. Tools return raw strings (e.g., 'Expense added successfully. ID: {new_id}') or unstructured database row lists. LLMs cannot reliably extract fields or plan downstream calls. add_expense returns a string message instead of a structured {id, status} object; list_expenses returns str(rows) instead of {expenses: [{id, date, amount, ...}], count: int}.
Parameter descriptions are incomplete or missing. user_id appears in every tool but descriptions are minimal ('User ID for authentication', 'User ID to filter expenses', 'User ID for the expense owner'). These do not explain the semantic role, default behavior, or format. category in summarize_expenses is described as 'Optional category filter' with a null default, but the parameter description does not explicitly state it is optional or that null means 'all categories'.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 8 | - | v1 |
Tool descriptions lack context and recovery guidance. delete_expense is described as 'Delete an expense by ID.', does not explain consequences, confirm step, or recovery path. list_expenses says 'List expenses within the inclusive Date Range' but does not explain pagination, result limits, or what happens if no expenses are found. summarize_expenses does not document what structure is returned (category + total?) or whether results are sorted.
Error handling is minimal and non-recoverable. Tools return raw exception strings like 'Database error: {str(e)}' with no guidance on what the LLM should do next. No distinction between retryable (connection timeout), user-fixable (invalid date format), or fatal (database corrupted) errors. No suggestions for next steps (e.g., 'Try again' or 'Check date format YYYY-MM-DD').
Destructive operation (delete_expense) lacks confirmation step. No dry-run, no confirmation_required response, no undo mechanism. An agent could delete the wrong expense and the server would silently commit the deletion. The description does not warn of irreversibility.
Schemas use inconsistent and under-specified parameter types. add_expense has 'amount' as number with no min/max bounds, an LLM could pass 1000000000 or negative values. date parameters accept 'YYYY-MM-DD format' as a string description, but there is no format validator in the schema (no 'pattern' field or 'format': 'date'). expense_id in delete_expense is integer with minimal description. subcategory and note default to empty string ('') or null without explanation of the semantic difference.
Parameter documentation violates the 10 - 1024 character guideline. Several parameter descriptions are under 20 characters: 'The ID of the expense to delete' (32 chars, acceptable), but 'User ID for authentication' (26 chars) is too terse and does not explain what happens with invalid IDs or how authentication is enforced. Tool descriptions like 'List expenses within the inclusive Date Range (YYYY-MM-DD)' (60 chars) lack context on when to use this tool or what output structure to expect.
No pagination or result limits. list_expenses and summarize_expenses return all matching rows without limit or offset parameters. If a user has thousands of expenses, the response could exhaust context window. No documentation of expected result size or guidance to cap at 20 - 50 items per call.
Default user_id='guest' silently creates a shared namespace. All unauthenticated calls operate under the same 'guest' user, potentially mixing data from different clients. No description warns of this behavior. If the server is intended to be multi-tenant, user_id must be mandatory or validated against the caller's credentials.