Autonomous invoice processing via MCP + LangGraph for Accounts Payable (AP) automation. Processes pending invoices, makes approval/rejection decisions, and manages notifications.
Business AP Agent exposes 8 tools with reasonable naming conventions (verb-first: list_, get_, validate_, approve_, reject_, flag_, send_, generate_) and actionable descriptions. However, critical gaps emerge: (1) Most tool descriptions lack depth on WHEN to use them vs alternatives, dependencies, or consequences; (2) Parameter descriptions are minimal or absent in several tools; (3) No documented output schemas, LLMs cannot anticipate response structure for planning; (4) Error handling is basic (success/error JSON wrappers) without guidance on recovery or retry logic; (5) No input validation, type constraints (enums), or range checks visible. The tools are functionally distinct and cover a legitimate AP workflow, but fall short of production-grade LLM optimization. Average per-tool score: 54.
Approves a pending invoice and updates its status to approved. Logs the approval action with optional reason.
Flags an invoice for human review by updating its status to under_review. Logs the action with a reason.
Generates a summary of all invoices with counts and totals by status (pending, approved, rejected, under_review). Returns total invoice count, status breakdown, and total pending/approved amounts.
Retrieves detailed information for a specific invoice by invoice number.
Retrieves all pending invoices with their details including invoice number, vendor name, amount, currency, description, due date, priority, and status.
Rejects a pending invoice and updates its status to rejected. Logs the rejection action with a required reason.
No documented output schemas for any tool. LLMs cannot predict response structure, forcing them to guess field names when chaining tools. E.g., after approve_invoice, does the response include 'timestamp' (ISO 8601) or 'approved_at'? Missing schema documentation forces LLM exploration and wastes tokens.
Error handling is minimal and non-actionable. All errors return JSON with 'status': 'error' and a message, but provide no guidance on recovery. E.g., 'Invoice INV-123 not found' tells the LLM nothing about whether to search, retry, or ask the user. No error classification (retryable vs fatal).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Sends a notification message to Slack via webhook. Logs the notification attempt with status. Requires SLACK_WEBHOOK_URL environment variable to be configured.
Validates an invoice by checking for issues (missing vendor name, invalid amount) and warnings (high value invoices, exceeds approval threshold). Returns validation status, issues, warnings, and a recommendation (APPROVE, REVIEW_REQUIRED, REJECT, or ESCALATE).
No input validation visible. Tools accept strings without length, format, or pattern constraints. E.g., invoice_number is not validated against expected format (e.g., 'INV-\d{6}'). reason parameter in reject_invoice is required but no minimum length enforced, empty strings may be accepted.
Ambiguous tool relationships. List of tools (list_pending_invoices, generate_ap_summary) return summaries, but it's unclear when to call list vs generate. No 'discovery' guidance or differentiation in descriptions. LLM may waste calls choosing between them.
Parameter descriptions incomplete or missing field semantics. E.g., 'reason' parameter in approve_invoice is optional but described as 'Optional reason for approval', no guidance on what constitutes a valid reason (audit trail? business logic?). In reject_invoice, 'reason' is required but no minimum length or allowed character set specified.
No permission or authorization checks documented. Tools like approve_invoice, reject_invoice, flag_for_review perform write operations on business-critical invoices with no visible permission validation. Description does not state what user/role can call these tools or whether authority is checked.
Idempotence undefined. The approve_invoice tool code checks if invoice is already approved and returns a success message, but this is not documented. LLM does not know whether retrying a call is safe or if it triggers side effects (e.g., duplicate audit logs).