A remote MCP server for tracking and managing expenses with SQLite database backend
ExpenseTracker is a basic CRUD server with 3 tools. All tools are explicitly defined with schemas and descriptions. However, multiple definition quality gaps limit its production readiness: (1) Tool descriptions are present but generic and lack actionable guidance ('Add a new expense entry' vs. context on when/why to use it); (2) Input schemas exist but parameters lack meaningful constraints, dates are strings with no format guidance, category accepts free-form strings instead of enums, amounts have no bounds; (3) Error handling is minimal (generic error returns) with no recovery guidance; (4) Output schemas are partially documented but lack richness (list_expenses returns raw dict zip, no field descriptions); (5) No pagination despite list_expenses potentially returning many records. Naming follows verb_noun convention correctly (add_, list_, summarize), a positive. Parameters are well-typed (string, number) but descriptions are sparse and lack format/constraint details. The server functions but falls short of production-grade tool design.
Add a new expense entry to the database.
List expense entries within an inclusive date range.
Summarize expenses by category within an inclusive date range.
No enum constraints on 'category' parameter across all tools. Free-form strings invite hallucinated invalid categories. Should declare known categories as an enum.
Date parameters (date, start_date, end_date) are typed as 'string' with no format constraint. Description does not specify ISO 8601 or any other format. LLMs will guess and may pass invalid dates like '2024-13-45'.
amount parameter (number type) has no min/max bounds documented. Unbounded numerics let LLMs pass negative, zero, or absurdly large values. Schema should declare minimum: 0 and suggest a maximum (e.g., 999999).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Tool descriptions are generic and under 100 characters, do not explain WHEN to use each tool vs. alternatives, WHAT downstream actions are enabled, or WHAT the return value structure contains. 'Add a new expense entry to the database' is vague; should explain: 'Adds a single expense and returns the new expense ID for later retrieval or updates.'
list_expenses returns raw dict list with no pagination (no limit, offset, or total count). For large expense databases, this can return thousands of records, exhausting context. Should add limit (default 20) and offset/next_cursor parameters.
Error responses are bare strings or dicts without recovery guidance. 'Database error: ...' tells LLM nothing actionable. Should return structured errors with 'error_code', 'message', and 'suggested_action' (e.g., 'check_date_format', 'retry_later', 'contact_support').
Output schemas not explicitly documented. add_expense returns {status, id, message}; list_expenses returns list of dicts; summarize returns category summary dicts. LLMs cannot plan downstream chaining without knowing field names and types. Add explicit output schema annotations.
add_expense is destructive (WRITE risk) but has no dry-run, confirmation, or idempotency hint. Agents should know: is this safe to retry? Should I confirm first? No idempotent hint means agent cannot safely re-attempt on transient failures.
'category' parameter in summarize marked with default: null (which is unusual for JSON Schema and may not parse). Should clarify: is category optional? If omitted, summarize ALL categories. Explicit description needed.