A remote MCP server for tracking and managing expenses with database storage and categorization
ExpenseTracker has 3 tools with adequate basic definitions and clear naming. All tools start with action verbs (add_, list_, summarize) and have descriptions present. Input schemas are properly defined with types and parameter descriptions. However, output schemas are not documented, error handling is minimal (generic messages without recovery guidance), and descriptions lack specificity about when to use each tool or what distinguishes them. Parameters are reasonably named but lack format constraints (date fields accept free-form strings without ISO 8601 validation hints). The server handles READ_ONLY and WRITE operations but provides no guidance for error recovery, and responses return raw dict/list structures without documented output schema contracts.
Add a new expense entry to the database.
List expense entries within an inclusive date range.
Summarize expenses by category within an inclusive date range.
Output schemas not documented. Tools return raw dicts/lists without specifying field names, types, or required fields. LLMs cannot plan downstream calls or extract structured data reliably.
Date parameters accept free-form strings without format constraints. Descriptions state 'Date of the expense' but do not specify ISO 8601 or 'YYYY-MM-DD'. LLMs will pass ambiguous formats (e.g. '01/15/2024' vs '2024-01-15') causing silent failures or misinterpretation.
Error handling is generic and non-recoverable. Errors return {'status': 'error', 'message': '...'} with no guidance on what the LLM should do next (retry, ask user, call another tool). Missing error classification (retryable vs. fatal).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Tool descriptions are brief and do not explain when to use each tool vs. alternatives. 'Summarize expenses by category' does not clarify: Does it group only by category, or also by date? Can it filter by multiple categories? When should the LLM call this vs. list_expenses?
No pagination or result limits documented. list_expenses and summarize do not specify max results, cursor handling, or pagination strategy. A large date range could return thousands of items, exhausting the context window.
add_expense returns lastrowid as 'id' but list_expenses returns 'id' as part of the dict. Consistency is present but output schema is undocumented, leaving LLMs to infer field names.