MCP Server that exposes financial management tools via Model Context Protocol. Allows AI assistants to interact with financial data using natural language for expense tracking, income logging, budget management, financial summaries, and AI-powered advice.
The Finance Coach MCP server demonstrates reasonable structure with 18 tools across expense, income, budget, summary, and advice domains. However, it suffers from significant quality gaps: (1) Most tool descriptions are extremely short (10-20 chars), well below the 50-200 char LLM-optimization baseline; (2) Input schemas are visible and properly typed, but parameter descriptions are minimal; (3) No output schemas are documented; (4) Error handling and recovery guidance are absent; (5) No examples in descriptions violate the 'do not include example values' rule (e.g., 'log $50 for groceries' teaches the LLM to pass '$50' literally). The tools themselves are well-named with clear verb prefixes (log_, get_, delete_, set_), but the presentation lacks the depth required for robust LLM reasoning. This is a typical C-grade community server, functional but not production-ready.
Delete an expense by ID. Example: delete expense expenses-1-20260307
Delete an income entry by ID. Example: delete income income-1-20260307
Get all previously generated advice entries. Example: show me my advice history
Generate fresh AI financial advice for a given month. Example: give me financial advice for March 2026
Get budgets that are over or close to their limit. Example: show me budget warnings
Get all budgets with current spending status. Example: show me all my budgets
Critical: Descriptions contain literal example values that LLMs will reuse. E.g., 'log $50 for groceries on 2026-03-07' teaches the LLM to pass '$50' and '2026-03-07' literally in real calls, not adapt them to user context.
High: Tool descriptions are extremely brief (10-35 characters), well below the 50-200 char LLM-optimization baseline. Descriptions lack context for tool selection and do not explain WHEN to use each tool vs. similar ones (e.g., difference between get_expenses and get_expenses_by_category is unclear).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Get all expenses grouped by category with totals. Example: show me spending by category
Get all expenses for a given month and year. Example: get expenses for March 2026
Get financial health score out of 100. Example: what is my financial health score?
Get all income grouped by source with totals. Example: show me income by source
Get all income for a given month and year. Example: get income for March 2026
Get the most recently generated advice. Example: show me my latest financial advice
Get income vs expenses trend over last N months. Example: show me my spending trend for last 6 months
Get complete financial summary for a given month. Example: give me my financial summary for March 2026
Log a new expense. Example: log $50 for groceries on 2026-03-07
Log a new income entry. Example: log $3000 salary income on 2026-03-01
Set or update a budget limit for a category. Example: set $300 monthly budget for groceries
Get complete financial context for AI decision making. Example: give me full financial context for March 2026
High: No output schemas documented. LLMs cannot plan chaining operations because they don't know what fields will be returned. For example, log_expense_tool returns {'...': '...'} with no documentation of what fields to expect, preventing the LLM from extracting expense_id for use in delete_expense_tool.
High: Parameter descriptions are minimal. E.g., 'Category of expense' does not explain valid categories, valid lengths, or format. 'Date of expense in YYYY-MM-DD format, defaults to today' is good, but most others lack constraints. Enum constraints (categories, sources, budget periods) are not enforced, free-form strings invite hallucinated invalid values.
High: No error handling or recovery guidance. If log_expense_tool fails (e.g., invalid category, negative amount, date in future), there is no indication of what went wrong or what the LLM should try next. No validation of inputs (amount >= 0, category in enum, date format valid) is visible.
Medium: Destructive tools (delete_expense_tool, delete_income_tool) lack confirmation/dry-run patterns. Agents can permanently delete records without a safety gate. A simple dry-run or explicit confirmation step would prevent accidental data loss.
Medium: No pagination support visible in tools that return lists (get_expenses_tool, get_income_tool, get_advice_history_tool). Without limit/offset or cursor params and a total count in response, large result sets can exhaust the context window.
Medium: Tool naming inconsistency with respect to 'tool' suffix. All tools are named with '_tool' suffix (log_expense_tool, get_expenses_tool), which is redundant and violates verb_noun convention. Prefer 'log_expense', 'get_expenses'. The suffix adds no semantic value and wastes characters.