A remote MCP server for tracking and managing expense entries with database storage
ExpenseTracker has three well-named tools with clear action verbs (add_, list_, summarize). All tools have descriptions and input schemas with typed parameters. However, several quality gaps prevent a higher score: (1) parameter descriptions are minimal and lack constraint details (e.g., date format, amount range); (2) output schemas are not formally documented, tool responses are inferred from code, not explicitly declared; (3) error handling returns generic error dicts without recovery guidance; (4) no input validation documentation or examples of expected formats. The server implements async/await correctly and uses fastmcp framework properly, but falls short on the 50-200 character description baseline (current descriptions average 60 chars, acceptable but brief) and lacks structured guidance for LLM tool selection.
Add a new expense entry to the database.
List expense entries within an inclusive date range.
Summarize expenses by category within an inclusive date range.
Date format not specified in parameter descriptions. Parameters 'date', 'start_date', 'end_date' accept strings but lack format guidance (ISO 8601? MM/DD/YYYY?). LLMs will guess and pass invalid formats, causing silent failures or database errors.
Numeric parameter 'amount' lacks range constraints. No min/max specified. LLMs may pass negative, zero, or absurdly large values. Should document: 'amount: positive number, typically 0.01 - 999,999.99'.
Output schemas not formally documented. Each tool's return structure is inferred from code (e.g., add_expense returns {status, id, message} or {status, message}; list_expenses returns list of dicts; summarize returns list of dicts). No explicit output schema in tool registration means LLMs cannot plan downstream operations or validate responses.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 54 | - | v1 |
Error responses are generic and non-actionable. All tools return {status: 'error', message: <str>} with no recovery guidance. Pattern requires: 'Error responses must tell the LLM what to do next.' E.g., database error should suggest 'Check file permissions or database connectivity.' Read-only database error is specific but others are vague.
Parameter descriptions are too brief (avg ~40 - 50 chars). Rubric baseline for param annotations is 72 chars. Examples: 'The date of the expense' (26 chars), 'The amount of the expense' (26 chars). Should expand with format/constraint info: 'The date of the expense (ISO 8601 format, e.g., 2024-01-15).'
Optional parameters ('subcategory', 'note', 'category' in summarize) have default='' or default=null but lack explanation of behavior when omitted. Should document: 'If omitted, all subcategories are included' or 'If omitted, filter is not applied.'
No input validation or constraint documentation. Rubric requires: 'Describe the expected format, range, and allowed values directly in the parameter description.' Category parameter accepts any string; no enum of valid categories is enforced at the tool level (categories.json exists but is not referenced in schema).
add_expense: returned 'id' field (expense_id) may not match naming convention if downstream tools expect 'expense_id' not 'id'. Rubric pattern:response-field-naming requires field names to match what tool parameters expect.