Add your description here
The server provides 3 tools with basic definitions and input schemas. Naming follows verb_noun convention (add_, list_, summary). Descriptions exist but are minimal (10-40 chars, below the 194-char baseline). Input schemas are present with JSON Schema structure, but descriptions are sparse and lack actionable guidance. No documented output schemas, no error handling guidance, no parameter constraints (enums for category), no pagination despite list_expenses potentially returning unbounded results. Security: uses SQLite with parameterized queries (good), but hardcoded DB paths and no input validation for date formats. Composition: tools are single-responsibility and chainable (add_expense → list_expenses workflow possible). Overall: meets minimum viability but falls short of production quality.
add a new expense
list all expenses entries from the database between a date range
Summarize expenses by category within an inclusive date range
Tool descriptions are critically short (10-40 chars vs. 194-char baseline). 'add a new expense', 'list all expenses entries from the database between a date range', 'Summarize expenses by category within an inclusive date range' lack context about WHEN to use each tool, dependencies, or what the LLM should expect in return.
No documented output schemas. list_expenses and summary return Python dicts with fields inferred from cursor.description, but the schema is implicit. LLMs cannot plan downstream tool calls or extract correct fields without explicit documentation of response structure (e.g., [{id: int, date: string, amount: float, ...}]).
Parameter 'category' in add_expense and summary lacks enum constraint. Free-form string invites hallucinated categories. Should define valid enums (loaded from categories.json) or document the constraint explicitly in parameter descriptions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 46 | - | v1 |
No date format validation or documentation. Parameters accept 'date', 'start_date', 'end_date' as strings with no documented format (ISO 8601? YYYY-MM-DD? natural language?). SQL query uses BETWEEN, which works for string dates only if they are ISO format, but LLMs frequently misformat dates. Parameter descriptions must state: 'ISO 8601 format (YYYY-MM-DD)'.
list_expenses and summary have no pagination limits. Returning all expenses from a date range could be thousands of rows, exhausting context window and degrading LLM reasoning. Should implement limit/offset parameters with a default (e.g., limit=20) and cap results per the rubric baseline of 20-50 items.
No error handling or recovery guidance. If date parsing fails, if category is invalid, or if the database is unavailable, the LLM receives raw exceptions (or silent failures). Errors must categorize as retryable/user-fixable and guide the LLM: 'Invalid date format. Use ISO 8601 (YYYY-MM-DD). Try again.' or 'Category not found. Call /expense://categories resource to see valid options.'
Input validation is missing. add_expense(amount) accepts any float, no validation that amount > 0 or reasonable bounds (e.g., max 1M). date field is not validated against format or reasonable ranges. Malformed input causes silent DB failures or nonsensical data.
category parameter in summary defaults to None, but add_expense requires it (non-optional). Inconsistent naming/requirements: is it optional or required? Parameter descriptions should clarify mutual exclusivity or relationships.
Resource 'expense://categories' exists but is not documented in tool descriptions. LLMs won't know to call it to discover valid category/subcategory options. Parameter descriptions should hint: 'Valid categories are listed in the expense://categories resource. Call it first if unsure.'
Parameter 'subcategory' in add_expense is optional with empty string default. No documentation of how subcategory relates to category (is it a free-form note, or constrained by category?). Undocumented dependencies invite misuse.