An offline-first Rust-based financial AI stack powered by MCP (Model Context Protocol)
This server has 11 tools with basic structure and mostly present descriptions, but significant quality gaps prevent a higher score. All tools have descriptions (meeting a baseline requirement), but they are often generic and lack actionable context. Input schemas are present and typed, but lack constraint details (enums, ranges, formats). Parameters have descriptions, but many are terse or incomplete. No tool annotations (readOnlyHint/destructiveHint) are present despite clear semantic differences (READ_ONLY vs WRITE vs DESTRUCTIVE). Error handling is not documented. The tool composition is reasonable (one concern per tool), but parameter naming could be clearer (e.g., 'filter' as an object is reasonable but lacks nested field constraints). No evidence of pagination support for list operations. Output schemas are not documented in the code provided.
Add a new monthly recurring spending entry
Add a new tax deduction entry
Record a cash flow transaction without an explicit date (defaults to today)
Record a cash flow transaction with an explicit date
Remove a monthly recurring spending entry by ID
Remove a tax deduction entry by ID
Scan spending transactions based on a time range filter (today, this month, this year, lifetime, or custom date range)
Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint). Tools are marked with Risk: READ_ONLY, WRITE, DESTRUCTIVE in metadata, but MCP server does not expose these as tool annotations. LLMs cannot determine which tools are safe to retry or have side effects.
Filter parameter uses nested object with incomplete constraint documentation. 'scan_spending' and 'visualize_spending' accept a 'filter' object with 'range' enum and optional 'start'/'end' fields, but the description lacks clarification on: (1) which combinations are valid (if range=Custom, are start/end required?), (2) date format expectations (ISO 8601?), (3) timezone handling. Undocumented dependencies cause LLM errors.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Simulate tax calculation for a given year based on total income and tax deductions
View all monthly recurring spending entries
View all tax deduction entries
Visualize spending aggregated by category based on a time range filter
No pagination support documented or visible for list operations. 'view_all_monthly_spending_list' and 'view_all_tax_deductions_list' return all entries without limit/offset/page parameters. If a user has hundreds of monthly spending entries, this could exhaust context window and degrade LLM reasoning.
Output schemas not documented in code. Tool descriptions and parameter schemas are present, but there is no evidence of documented return types. LLMs cannot plan downstream tool calls or extract required fields (e.g., does 'scan_spending' return item IDs needed by other tools?) without seeing the output schema.
Generic/incomplete parameter descriptions. Examples: 'Transaction description' (what should be included? length limits?), 'Tax year to simulate' (what range of years is valid?). Descriptions under 50 characters lack actionable context for LLM parameter selection.
No error handling guidance documented. Tools do not describe what happens on invalid input (e.g., invalid category, negative amount, invalid date format, year out of range). No recovery instructions for LLM (e.g., 'If category is invalid, call list_categories first').
Destructive operations lack confirmation/dry-run support. 'remove_monthly_spending' and 'remove_tax_deduction_list' are irreversible, but no pattern exists to confirm before executing or see what would be deleted. Agents making mistakes could cause unrecoverable data loss.
'due_date' parameter in 'add_monthly_spending' lacks format specification. Is it a day-of-month (1-31)? An ISO date? A cron expression? LLMs will guess, leading to invalid input.