MCP server for managing expenses with SQLite database backend, supporting add, list, update, delete, and summarize operations
The server defines 6 tools with visible schemas and descriptions in main.py. However, critical gaps in parameter documentation, output schema clarity, and error handling prevent a higher score. Tool names follow verb_noun convention (add_expense, list_expenses, etc.), which is correct. Descriptions exist but are minimal (10-50 chars for most tools), below the baseline of 194 chars. Parameters have type definitions but lack depth in describing constraints, ranges, and valid values. Output schemas are inferred from return statements but not formally documented. Error handling is present but minimal, error responses lack recovery guidance. The resource tool 'categories' is underdescribed (only 50 chars). Overall, this server demonstrates functional basic patterns but falls short of production quality; it lands in the D/low-C range.
Add a new expense to the database.
Resource providing access to expense categories
Delete an expense by ID.
List expenses filtered by category and an optional date range. If start_date or end_date are missing, the function will return results across the full DB range.
Get total expenses by category within a date range.
Update an existing expense.
Output schemas are not formally documented. All 6 tools return dict objects with 'status', 'message'/'data' fields, but the structure is not declared in tool metadata. LLMs cannot infer what fields to extract or plan downstream operations. list_expenses returns an array of expense objects (id, amount, category, subcategory, date, description), but this structure is only visible in the implementation, not in the tool registration.
Parameter descriptions lack constraint details. 'amount' (float) has no min/max; 'category' (string) has no enum or list of valid options; 'date' (string) specifies YYYY-MM-DD format in docstring but not in parameter description. LLMs cannot validate inputs before calling, invalid amounts (negative, extremely large) or unknown categories will be passed, causing silent failures or server-side rejections.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Error messages are minimal and non-actionable. add_expense returns {"status": "error", "message": "date must be YYYY-MM-DD"} but does not suggest how to recover or what the current date is. list_expenses silently defaults missing dates to '0001-01-01' and '9999-12-31', an LLM cannot know this behavior, leading to unexpected full-range results. Error responses lack recovery guidance (pattern:recovery-guide).
Resource tool 'categories' is severely underdescribed (only ~50 chars: "Resource providing access to expense categories"). No explanation of what the JSON contains, whether it is a list of strings or objects, or how to use it. The MIME type is declared but the schema is not. The FileNotFoundError fallback returns '[]' without warning.
Destructive tools (delete_expense) lack confirmation or dry-run. An agent can immediately delete records without a second step. No idempotency check, calling delete_expense twice with the same ID will succeed the first time and silently fail the second (no record to delete), but the response is the same ({"status": "ok"}). LLMs cannot tell if deletion succeeded or the record was already gone.
Tool composition could be improved. list_expenses and summarize both filter by date range but return different structures (full records vs aggregates). An agent may not know which to call for a given intent. Consider clarifying use cases in descriptions (e.g., 'Call list_expenses to inspect individual transactions; call summarize for high-level totals').
update_expense accepts optional float/string fields with None defaults, but the description and behavior are unclear. Passing {"expense_id": 5} with no other fields returns {"status": "error", "message": "no fields to update"}. An LLM may be unsure whether this is a recoverable error or a sign of incorrect tool usage. The parameter descriptions do not clarify that at least one field must be provided.