MCP server for Beancount plain-text accounting: double-entry transactions written by an AI agent, validated with bean-check and committed to git.
Countbean MCP demonstrates strong foundational quality with excellent descriptions and clear naming conventions. All 14 tools are explicitly registered with comprehensive parameter documentation and risk classifications. However, output schemas are not formally documented in the source code provided, and error handling lacks structured recovery guidance. Tool names follow verb_noun patterns consistently (connect_book, start_device_authorization, await_device_approval, create_book, etc.), aiding LLM disambiguation. Most descriptions are well-written (100-300 chars), explaining WHAT the tool does and WHEN to use it. Parameter descriptions are present and mostly clear, though some lack type constraints (e.g., 'pattern' in show_report accepts regex but this is only mentioned in prose, not as a formal constraint). The rubric baseline for A-grade tools is 80+; this server falls short of that due to missing output schema documentation and incomplete error recovery patterns, but exceeds typical community averages (50-60) significantly.
Every account in this book, sorted by name and grouped by type. For each account, the list includes open date, close date (if any), allowed currencies, and the sum of postings (a bare number for a single currency, a dict for multiple). Accounts are returned in a tree structure that mirrors the account hierarchy.
Add transactions to the book, validate with bean-check, and commit to git. Take a list of Proposal dicts (from import_statement or your own design) and write them to the ledger as beancount transactions. bean-check validates them immediately; if validation fails the write is rolled back and you get the error. If validation passes, the transactions are committed to git with the message you provide, or a default. The commit includes a copy of every input, so a later import can detect duplicates. Every transaction MUST have: - date: ISO string (YYYY-MM-DD) - flag: '*' or '!' (posting flag; typically '*' for normal txns) - narration: string description - account: the account money came from or went to - amount: decimal string or plain number - currency: 3-letter code - counter_account: the other side of the double-entry (can be a placeholder like Expenses:Unclassified; the model can change it) Optional: - import_id: string for duplicate detection (auto-generated if omitted) Return value is a commit hash and a summary line.
Step 2: wait for the user to approve the code from `start_device_authorization`. Blocks until they approve, decline, or the code expires. On success the connection is saved and every countbean tool switches to that book.
The current balance of every account in every currency, in a flat list.
Output schemas not formally documented in source. While tool descriptions mention what is returned (e.g., 'returns monthly totals, run rate, cash on hand'), the LLM cannot parse returned field types, required vs optional fields, or nested structures. This violates pattern:tool and forces LLMs to infer structure from narrative descriptions.
Error handling lacks recovery guidance. Tools like add_transactions (IRREVERSIBLE) and import_statement document risks but do not provide explicit error categories (retryable, user-fixable, fatal) or next-step guidance. When validation fails, the LLM receives 'error' but no hint whether to ask the user to edit the transaction or try a different approach.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 66 | 2026-07-28+ | v2 |
Connect this plugin to a hosted Countbean book, permanently. Give it the key shown once on your book's page (`cbk_…`) and the book id (`bok_…`). Verifies the pair against the live book BEFORE saving, then stores it in ~/.countbean/credentials.json (0600). Takes effect immediately — no restart, no environment variables. Paste both on one line and this tool sorts them out; the `bok_…` id can be omitted if you have already connected to that book before.
Create a new local book at the given path, or at ~/.countbean/main.
A list of every currency this book mentions, alphabetically: three-letter codes, nothing else. Use run_query for breakdowns.
Export a report to an Excel (XLSX) file. All report_types work (tax_summary, account_history, budget_summary, category_breakdown), and the Excel file includes formulas, formatting, and charts where sensible. Returned as base64-encoded bytes you can hand back to the client.
Parse a statement (CSV, OFX, or pasted table) into proposed transactions. Returns a list of Proposal dicts, each with a date, amount, description, and counter-account placeholder. Every row that could not be read with confidence is flagged with ambiguities, and duplicates against the existing book are marked. The proposals are NEVER WRITTEN. Call add_transactions to write them, and optionally edit them first (you can, the proposal is just JSON). This tool decides parsing only: column mapping, date formats, debit/credit direction, and whether the row is already in the book.
A digest of facts about this book, suitable for summary prose. If the book has income and expenses, returns monthly totals, run rate (average monthly net), cash on hand, liabilities, spending by category, month-over-month changes and outliers, and warnings if the data is thin. Every derived number — a ratio, a percentage, a spending delta — comes back as a dict with 'sufficient': true or false and a 'why' that explains what it needs to trust the figure. Do not report a number without reporting its sufficiency. If you are Claude, pretend every 'sufficient': false means the customer told you not to answer the question. Do not invent a number from thin data, and do not report a number without checking sufficiency — a confident "burn rate: $9000 per month" off three weeks of data is how this feature breaks trust.
Run a BeanQuery SELECT statement. This is the escape hatch for anything the named tools do not cover: custom reports, category analysis, date ranges, regex filters on account names, anything. The full BeanQuery manual is at query.beancount.org. The returned rows are dicts, not arrays, so you can address a cell as row['name'] not row[3]. Currency values are Decimal objects so they are exact and do not round like floats.
Return a report of a specific account or category in a human-readable format. Report types: - tax_summary: summarize which accounts matched a regex, by category, income vs. expense - account_history: debit/credit activity on a specific account - budget_summary: compare actuals to a budget you provide - category_breakdown: spending by subcategory under a parent (e.g., Expenses:*) All but budget_summary work on the whole book. budget_summary takes explicit budget amounts as a dict.
Step 1 of connecting a hosted book: get a code for the user to approve. Returns IMMEDIATELY with a short code and a link. Show both to the user, then call `await_device_approval` to wait for them to approve it. Deliberately two tools and not one. An MCP tool returns a single result, at the end — so a tool that fetched the code and then waited for approval could never show the code to the person who has to type it. It could only ever expire. That was the first version of this, and it was unusable.
Validate a piece of beancount syntax without writing it. Parse it, run bean-check, and return OK or an error. Useful for showing the customer that a transaction you are about to ask them to confirm is actually valid beancount.
Parameters with constrained values lack enum declarations. 'sign_convention' in import_statement accepts 'debit_positive', 'credit_positive', or 'auto', documented in prose but not as a JSON Schema enum. 'date_format' and 'content_encoding' have the same issue. LLMs cannot auto-complete or validate against these constraints without parsing the description text.
No confirmation or dry-run pattern for irreversible operations. add_transactions (IRREVERSIBLE) writes to a git-backed ledger. While validation is performed, there is no confirm_before_execute or dry-run mode to let LLMs preview the impact before committing. An agent error leads to a git commit that requires manual reversal.
Pagination and result limits not documented. Tools like accounts, balances, currencies, and run_query do not specify maximum result sizes or pagination. If a user's book has thousands of accounts, accounts() could return massive output, blowing context and degrading LLM reasoning. Missing limit and pagination guidance.
Parameter 'pattern' in show_report and download_report described as regex but lacks format specification or examples. An LLM may pass an invalid regex and receive a parse error instead of clear guidance on expected syntax.