Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The Budget Tracker MCP server has 13 tools with explicit schemas and descriptions visible in the source code. Tool naming follows verb_noun convention (add_, list_, search_, update_, delete_, bulk_, get_, import_). Most tools have descriptions between 34-194 chars, within the baseline range. However, there are significant gaps: (1) No output schemas are documented anywhere, tools return results but the response structure is not specified for the LLM to understand what fields to expect; (2) Error handling descriptions are missing, no guidance on what to do if a transaction ID is invalid or a PDF import fails; (3) Security-sensitive tools like import_from_email expose credentials as required parameters (email_address, email_password) instead of using server-side injection; (4) No idempotency or confirmation patterns for destructive operations (delete_transaction); (5) Some parameter descriptions lack format guidance (e.g., 'date' parameters state 'YYYY-MM-DD' but no validation rules about future dates, past date limits, etc.); (6) Batch tools (add_multiple_transactions, bulk_update_transactions) are good composition patterns but lack per-item error reporting, if 1 of 50 transactions fails, the response is unclear. Parameter schemas are well-formed with types and enums, which is a strength. Tool descriptions are actionable and non-trivial (avg ~120 chars), which is above the p10 baseline. The main weakness is the missing response schemas and error recovery guidance.
Credentials exposed as tool parameters. import_from_email requires email_address and email_password as required input parameters. These will be logged in agent traces, audit logs, and prompt history. Credentials must use server-side secret injection via environment variables or a vault.
import_from_email
Recommendations
Document output schemas for all tools. For each tool, add a 'Returns' section in the docstring or tool metadata describing the response structure. Example for list_transactions: 'Returns { transactions: [{ id: number, description: string, amount: number, direction: enum, category: string, datetime: ISO8601, ... }], total: number, limit: number }'. This lets LLMs understand chaining opportunities.
Move email credentials to server-side secret injection. Remove email_address and email_password from import_from_email parameters. Instead, read them from environment variables (BUDGET_EMAIL_ADDRESS, BUDGET_EMAIL_PASSWORD) or a secure vault at runtime. Document this in the server setup guide.
Add error recovery guidance to all tool descriptions. E.g., delete_transaction: 'Deletes the transaction with the given ID. Returns 404 if the ID does not exist; call list_transactions to find valid IDs. Deletion is permanent and non-reversible.' Similarly, for import tools: 'Returns an error if the file is not readable or in an invalid format; ensure the file path is correct and the format is supported.'
Implement confirmation pattern for delete_transaction and bulk_update_transactions. Add a dry_run boolean parameter (default false) that shows what would be deleted/updated without executing. When dry_run=true, return a preview of affected transactions. When dry_run=false, execute the operation. This prevents accidental data loss.
Add per-item error reporting to batch operations. For add_multiple_transactions and bulk_update_transactions, return a structured response like { successful: [{ index: 0, id: 123 }, ...], failed: [{ index: 5, error: 'Invalid category' }, ...] }. This lets the LLM know which items succeeded and which failed, enabling selective retry.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 32 points across a rubric change (v1 → v2)
67/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
67
<=2025-11-25
v2
2026-03-09
F
35
-
v1
50/100
Import transactions from email (Gmail, Outlook receipts)
import_from_imagewrite50/100
Import transactions from an image (receipt photo, screenshot)
import_from_pdfwrite50/100
Import transactions from a PDF file (bank statements, receipts)
list_transactionsread onlysource verified83/100
List recent transactions with optional filters
search_transactionsread onlysource verified83/100
Search transactions with advanced filters including date range and amount
No error handling guidance. Tool descriptions do not explain what happens on failure or what the LLM should do next. E.g., delete_transaction description does not say 'Returns error if transaction_id not found; try list_transactions to find valid IDs.' No recovery paths documented.
Destructive operations (delete_transaction, bulk_update_transactions) lack confirmation or dry-run patterns. An agent could accidentally delete the entire transaction history. No idempotency guarantees stated, unclear if retry with same transaction_id is safe or causes duplicate entries.
Batch tools do not specify per-item error reporting. add_multiple_transactions description does not state whether a single failed transaction causes the entire batch to fail, or if partial success is possible. No guidance on retry strategy for failed items within a batch.
Date parameters lack validation constraints. start_date and end_date are documented as 'YYYY-MM-DD' but descriptions do not state: are future dates allowed? What is the lookback limit? Can end_date be before start_date? These ambiguities let LLMs pass invalid values.
Numeric parameters unbounded. 'limit' in list_transactions has a default of 10 but no stated maximum. An LLM could pass limit=1000000, causing timeouts or OOM. Amount parameters (in transaction objects) have no min/max stated beyond 'must be positive'.
No permission/scope declarations. Tools do not state what permissions are required (e.g., 'read:transactions', 'write:transactions', 'delete:transactions'). This prevents least-privilege agent configuration and clear audit trails.
Parameter relationships undocumented. For import_from_email, the relationship between 'days' (lookback window) and 'keywords'/'senders' is unclear, are keywords AND senders required together, or optional separately? The description does not explain dependencies.
LLM providers specified as tool parameters. import_from_pdf, import_from_image, import_from_email, import_from_directory all accept 'provider' or 'llm_provider' as user input (enum: anthropic|local). This couples the tool interface to LLM availability and forces the LLM to decide which provider to use. Better: make provider a server config and remove from tool params.
Constrain numeric parameters. In list_transactions, add 'maximum 100' to limit description. In add_transaction and add_multiple_transactions, add 'amount must be between 0.01 and 999999.99' (or appropriate business limits). In search_transactions, constrain min_amount and max_amount similarly.
Add date range validation guidance. For start_date and end_date, document: 'Format YYYY-MM-DD. start_date must be on or before end_date. Maximum lookback is 10 years. Future dates are not allowed.' Prevent LLMs from passing invalid date ranges.
Add scope/permission declarations. For each tool, add a comment in the schema or description: 'Requires: read:transactions' (for list/search) or 'Requires: write:transactions' (for add/update) or 'Requires: delete:transactions' (for delete). This documents access control boundaries.
Clarify import_from_email parameter relationships. Update the description: 'Optional filters: keywords (search for any of these words in email subject/body) AND senders (filter by from address). If both are provided, emails must match both. If neither provided, returns all emails from the lookback window.'
Remove provider/llm_provider from tool parameters and make it a server configuration. Add a note to import_from_pdf/image/email/directory: 'The LLM provider is configured at server startup. This tool uses the configured provider for text extraction.' Remove the provider enum parameter. This decouples the tool interface from LLM selection and simplifies the agent's responsibility.
Add examples of chaining in tool descriptions. E.g., search_transactions: 'Returns transaction IDs and details. Use the returned transaction_id to call update_transaction or delete_transaction. Call get_spending_summary with the same date range to compare individual transactions against category totals.'
Document pagination limits. For list_transactions, state 'Default limit is 10. Results are paginated; if total > limit, the LLM should call list_transactions again with an offset parameter (if supported) or iterate with different filters.' If offset is not supported, state 'Not paginated; results are always the most recent N transactions.'
Add audit logging declarations. Document that all write operations (add, update, delete, import) are logged with user identity, timestamp, and parameters for compliance. This reassures users that their data changes are traceable.