An MCP server for tracking personal expenses
Server has 4 tools with mixed quality. Two tools (add_expense, show_transactions) are domain-relevant but lack critical documentation. Two tools (add_two_numbers, generate_random_number) are test utilities that don't belong in a production expense tracker. Schema definitions are present but sparse. Descriptions are minimal (10-25 chars) and fail to explain WHEN or WHY to use each tool. No parameter descriptions for show_transactions. No output schemas documented. No error handling guidance. No tool annotations. This server demonstrates a foundational MCP implementation but lacks the depth required for reliable agent interaction.
Adds the expense details into the Database
Adds two numbers a and b
Generates a random number between 0 and 100
Lists down all the data in expenses table
show_transactions missing parameter descriptions and type hints. 'start' and 'end' parameters have no type declarations in code, only inferred as strings in schema. Description says 'Lists down all the data' but doesn't explain what data structure is returned or what date format is expected.
Tool descriptions are under 30 characters and lack context. 'Adds the expense details into the Database' (44 chars) doesn't explain when to use this vs alternatives, what happens if date/amount are invalid, or whether calls are idempotent. 'Lists down all the data in expenses table' (42 chars) doesn't specify pagination, result limits, or date format expectations.
No output schemas documented. add_expense returns {'status': 'ok', 'id': <int>} but LLM has no formal schema. show_transactions returns a list of dict objects with fields [id, date, amount, category, subcategory, note] but structure is undocumented, forcing LLMs to infer field types and presence.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Test utilities (add_two_numbers, generate_random_number) mixed with domain tools. These belong in a separate test server or should be removed from production. They dilute the tool namespace and confuse agents about the server's actual purpose.
No error handling or recovery guidance. If date format is invalid, amount is negative, or category is unrecognized, no error response tells the LLM what to do. No validation, no constraint documentation (e.g., 'amount must be > 0' or 'date must be ISO 8601').
add_expense accepts 'date' as string with no format specification. LLMs will guess: ISO 8601? US MM/DD/YYYY? Locale-dependent? Query will fail silently if format doesn't match. Expected format must be stated in parameter description.
add_expense has no validation constraints documented. Negative amounts should be rejected. Empty category should be rejected. No mention of whether duplicate entries are allowed or how they're detected. Idempotency is unclear, will multiple identical calls create duplicates?
show_transactions returns potentially unbounded result set. No pagination, no limit, no cursor. If expenses table has 10,000 rows, entire dataset is returned, exhausting context and tokens. Baseline best practice: cap at 20-50 items per page with offset/limit params.