A multi-user expense and income tracking MCP server with SQLite backend, designed to run on Horizon with gateway-managed authentication
ExpenseTracker has solid naming and parameter coverage for its four tools (add_expense, edit_expense, delete_expense, list_expenses). All tools are registered with input schemas and descriptions present. However, there are notable gaps: (1) Output schemas are not documented, callers don't see what fields to expect from responses, forcing inference; (2) Descriptions lack specificity about return values and when to use each tool vs others; (3) Error handling is minimal, no recovery guidance in tool descriptions; (4) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk stratification (READ_ONLY, WRITE, DESTRUCTIVE); (5) Enum constraints are missing (e.g., category and subcategory are free-text strings without a predefined list, inviting hallucinated values). The server is competent but falls short of production polish.
Add a new expense entry to the database. Args: date: Date of the expense in YYYY-MM-DD format. amount: Expense amount (positive number). category: Top-level category (e.g. 'food', 'transport'). subcategory: Sub-category within the chosen category. note: Optional free-text description.
Delete expense entries. Only deletes expenses owned by the current user. Two modes: • Single delete — provide id. • Bulk delete — provide start_date + end_date, optional category filter.
Update one or more fields of an existing expense. Only edits expenses that belong to the current user. Args: id: ID of the expense to edit. date: New date (optional). amount: New amount (optional, must be positive). category: New top-level category (optional). subcategory: New sub-category (optional). note: New note text (optional).
Retrieve the current user's expenses within a date range. Args: start_date: Start date (YYYY-MM-DD), inclusive. end_date: End date (YYYY-MM-DD), inclusive. category: Optional category filter.
No output schemas documented for any tool. Callers (LLM agents) cannot infer what fields to expect in responses, breaking tool chaining and forcing wasteful inference or follow-up discovery calls.
Missing enum constraints for category and subcategory parameters across all expense tools. Free-text strings invite LLM hallucination (e.g., 'food-dining' instead of 'food'). Should declare a fixed set of valid categories and enforce at schema level.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk stratification in the metadata (READ_ONLY, WRITE, DESTRUCTIVE). Agents cannot determine which tools are safe to retry or which modify state without reading descriptions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Destructive operation (delete_expense) lacks confirmation or dry-run support. No pattern offered to prevent accidental deletion. An agent mistake could delete all expenses in a date range with no undo.
list_expenses has no pagination parameters (limit, offset, cursor) and returns all results. Large expense histories could exceed context windows. Rubric baseline: list tools should accept page/offset and limit, capped at reasonable limits (20 - 50).
Parameter relationships underdocumented in delete_expense. Whether 'id' and 'start_date/end_date' are mutually exclusive is stated in docstring but not in parameter descriptions, forcing LLMs to infer intent.
Error handling is minimal. Tool descriptions do not include recovery guidance (e.g., 'if amount is invalid, provide a positive number'). Responses return status/error messages, but no guidance on what the LLM should do next.