An expense tracking assistant that helps manage and summarize expenses with category validation.
The Expense Tracker server has 4 tools with complete JSON schemas and descriptions. Naming conventions are appropriate (verb-noun pattern: add_expense, list_expenses, summarize, get_categories). However, there are significant gaps in parameter descriptions, output schema documentation, and error handling guidance. Three of four tools lack complete parameter descriptions, and none document their output schemas. The add_expense tool lacks validation hints for the date parameter. Error handling is minimal, no recovery guidance or retryability classification. Tools are well-composed (single responsibility) and use appropriate async patterns. The instructions in the lifespan mention the expense://categories resource but don't guide LLMs on error scenarios (e.g., what to do if an invalid category is passed).
Add a new expense to the database.
Get a list of unique expense categories.
List all expenses in the database.
Summarize total expenses by category for a particular period.
Output schemas not documented. Tools return list[dict] or dict, but the structure of those dicts is not formally documented. LLMs must infer field names from the tool names or rely on inspection. add_expense returns {'status': 'success', 'id': <int>}, but this is not stated in the tool definition. list_expenses returns list with {id, date, amount, category, subcategory, notes}, but undocumented.
Date parameter format not specified. Both list_expenses and summarize accept start_date and end_date as strings, but the format is not documented. This forces the LLM to guess: ISO 8601, YYYY-MM-DD, MM/DD/YYYY, or a natural-language string? The SQL query suggests a simple string comparison, so format consistency is critical.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 60 | - | v1 |
No error handling or recovery guidance. If an agent passes an invalid category (not in the categories.json resource), add_expense silently inserts it into the database. The tool should validate category membership against the resource and return actionable errors like: 'Invalid category: "shopping", valid categories are: food, entertainment, transportation. Call the get_categories resource to see all valid options.'
No pagination or result limits. list_expenses and summarize can return unbounded results. If the database grows to thousands of expenses, returning all of them in a single response exhausts context and wastes tokens. Tools should accept limit and offset parameters, document the default limit (e.g., 20), and support pagination.
Parameter constraints missing for numeric and string inputs. amount (float) has no min/max bounds, can the agent pass -999999 or 0? date fields have no format constraints. subcategory and notes default to empty string but don't document max length or allowed characters.
Ambiguous categorization: get_categories is registered as a resource (expense://categories), not a tool, but the evaluation prompt lists it as a tool. This creates ambiguity in the server's interface. Either register it as @mcp.tool and make it callable from tool invocation, or document clearly that it's a resource and not available through tool calling.
Server instructions mention using the expense://categories resource but don't guide LLMs on what to do when category validation fails. Instructions should say: 'If an add_expense call fails due to invalid category, check the expense://categories resource and suggest a valid category to the user.'