Add your description here
This expense tracker MCP has clear tool names and basic schemas, but significant gaps in parameter descriptions, output documentation, and error handling guidance. Of 5 tools, 3 are domain-relevant (add_expense, list_expenses, summarize) and 2 are demo artifacts (roll_dice, add_number) that should not be in a production server. Tool schemas exist but parameter descriptions are minimal (10-20 chars), falling below the 72-char baseline for production tools. Output schemas are not documented, LLMs cannot predict return structure. Error handling is present in main_remote_mcp.py but absent in main_exp.py (the primary entry point). The server conflates demo tools with real ones, and tool descriptions lack context on when to use each tool vs. alternatives.
Add a new expense entry to the database.
Add two numbers and return the result.
List expense entries within an inclusive date range.
Roll n six-sided dice and return the results.
Summarize expenses by category within an inclusive date range.
Demo tools (roll_dice, add_number) included in production server, violates single-responsibility pattern
Parameter descriptions too short (10-20 chars), below 72-char baseline. E.g., 'Date of the expense' for 'date' param in add_expense does not explain format (YYYY-MM-DD?). LLMs cannot infer constraints.
No output schema documentation. add_expense returns {status, id}, list_expenses returns array of objects, summarize returns {category, total_amount, count}, but LLM cannot see this structure. LLMs must guess what fields to extract for downstream use.
main_exp.py (primary entry point) has no error handling or recovery guidance. Only main_remote_mcp.py provides error messages. Users deploying the default config get silent failures with no LLM-actionable error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
No input validation or constraints on date format. 'date' param accepts any string, LLM might pass 'Jan 1', '1/1/2024', or 'today'. Query will silently return no results if format mismatches DB.
Tool descriptions do not explain when to use one vs. another. Why call 'summarize' instead of calling 'list_expenses' and computing locally? Distinction unclear.
No pagination or result limits on list_expenses and summarize. If the DB grows to 10k+ expense records, list_expenses could return thousands of items, exhausting context and degrading LLM reasoning.