Comprehensive expense tracking MCP server with analytics, visualizations, and German tax tracking
This expense tracker server demonstrates solid fundamentals with comprehensive tool coverage (16 tools) and consistent schemas across all tools. All tools have descriptions and properly typed input parameters. However, the server exhibits moderate gaps in parameter descriptions, error handling guidance, and output schema documentation that prevent it from reaching 'good' (70+) territory. Parameter descriptions are often generic and lack the constraint details (ranges, patterns, formats) that help LLMs avoid invalid inputs. Output schemas are largely undocumented in the tool definitions shown. Error handling is minimal, most tools lack recovery guidance. While naming is generally clear and verb-driven, the descriptions are functional but not optimized for LLM decision-making. The code structure is clean and separated into services, which is a positive architectural choice.
Add a new expense row to the database.
Trend analysis grouped by time period (day, week, or month).
Return category analytics: counts, totals, averages, min/max and share (%) per category.
Compare spending between two months.
Delete an expense row by id.
Update one or more fields of an existing expense row by id. Only fields that are not None are updated.
Export data to a file (CSV, JSON, or Excel); returns the file path.
Output schemas not documented in tool definitions. Tools like generate_html_report, generate_charts, export_data, and import_expenses return file paths or success status, but the exact response structure (field names, types, whether a list or object) is not documented in the visible tool registrations. LLMs cannot plan chained calls without knowing what fields are returned.
Parameter descriptions lack constraint details. Numeric parameters (amount, limit, offset, months_ahead, based_on_last_months, year) have no documented ranges or limits. String parameters (date, month, format, category) lack format specifications or enum declarations. Date format is mentioned in the tool docstring but not in individual parameter descriptions. LLMs will pass invalid values without clear constraints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 53 | 2026-07-28+ | v2 |
Simple moving-average forecast based on historical monthly totals.
Generate PNG charts (Matplotlib) and returns generated file paths.
Generate an interactive HTML report with Plotly charts and returns the file path.
Quick stats (total spent, avg expense, top category, most expensive day).
Import expenses from a file path on disk (CSV or JSON format).
List expenses in an inclusive date range. Defaults: If dates are omitted, returns current month-to-date.
Flexible search/filter over expenses (pagination supported). Note: Unlike list_expenses, this tool does NOT auto-fill missing dates. If you omit both start_date and end_date, it searches across all rows.
Summarize expenses by category in an inclusive date range. Defaults: Month-to-date if start_date/end_date omitted.
Summary for tax-deductible expenses grouped into common German tax buckets.
Delete and destructive operations lack confirmation/dry-run support. delete_expense has no dry-run mode or confirmation step. generate_html_report and generate_charts write to disk with optional output_path but no validation or recovery guidance if the path is invalid or the directory doesn't exist.
Error handling is minimal. Tool descriptions do not explain what errors can occur, what the LLM should do on failure, or how to recover. For example, delete_expense does not explain what happens if the id doesn't exist. search_expenses, import_expenses, and export_data don't document failure modes.
Tool descriptions are functional but not LLM-optimized. Most descriptions (65-75 chars) fall below the production baseline of 194 chars average. They state WHAT the tool does but often lack WHEN to use it or key prerequisites. For example, 'Summarize expenses by category' doesn't explain when to call it vs. category_analytics, or that it groups results differently.
Pagination documentation is incomplete. search_expenses accepts limit and offset but does not document the default page size, max page size, or whether there's a total_count in the response. Tools like list_expenses don't mention if results are paginated or capped at a limit.
No permission or scope declarations. Tools like delete_expense and import_expenses do not declare what permissions they require (e.g., 'admin', 'write:expenses'). For security and auditability, each tool should document its required scope.
Missing batch operations. Tools like add_expense operate on single records. If an agent needs to import or add 50 expenses, it must call add_expense 50 times. No bulk_add_expenses variant exists to reduce round-trips and token waste.
Resource 'expense://categories' is mentioned but no explicit @mcp.resource() decorator is visible in the provided source code snippet. If this is inferred rather than explicitly registered, tool scoring may need to be capped.