An MCP server providing invoice management capabilities with financial calculation, compliance checking, and template management
MCP Toolbox has moderate definition quality with significant gaps. Tools have descriptions, but most parameter schemas are incomplete. Of the 5 tools, 2 have well-structured input schemas (analyze_invoice, generate_invoice) while 3 have minimal parameter documentation (list-invoices, create-invoice, mark-invoice-paid). Tool naming is verb-based but some names use hyphens inconsistently (list-invoices vs create-invoice vs mark-invoice-paid), which is not aligned with standard snake_case conventions. Descriptions vary in quality: analyze_invoice and generate_invoice are detailed (100+ chars), but list-invoices, create-invoice, and mark-invoice-paid are generic (20-50 chars). Error handling is not visible in the source code provided. Output schemas are not documented.
Analyzes an invoice document to extract structured information including document structure, financial summary, compliance status, template matches, and suggested improvements
Creates a new invoice for a hotel with guest name and amount
Generates a new invoice document from invoice data with optional template selection, including validation, financial calculation, and template application
Lists all invoices in the system, optionally filtered by hotel_id
Marks an invoice as paid or unpaid by invoice_id
Inconsistent tool naming: three tools use hyphens (list-invoices, create-invoice, mark-invoice-paid) while standard MCP convention is snake_case. This confuses LLM tool selection and differs from the verb_noun pattern used by analyze_invoice and generate_invoice.
Three tools have minimal descriptions (20-50 characters): list-invoices ('Lists all invoices in the system, optionally filtered by hotel_id'), create-invoice ('Creates a new invoice for a hotel with guest name and amount'), and mark-invoice-paid ('Marks an invoice as paid or unpaid by invoice_id'). These do not meet the 50-200 character target for LLM optimization and lack context on when/why to use each tool.
Output schemas are not documented for any tool. LLMs cannot plan downstream calls or extract required fields without knowing what structure to expect. LLMs need to know what fields to expect.' This affects all 5 tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 39 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 13 | - | v1 |
Parameter descriptions are missing or minimal for 3 tools. list-invoices has 'Optional hotel ID to filter invoices (optional)', redundant and vague. create-invoice parameters are documented but lack format constraints (e.g., no mention of currency format, date format expectations, or valid amount ranges). mark-invoice-paid lacks detail on what 'paid' boolean actually controls or what the returned state is.
No tool annotations detected (readOnlyHint, destructiveHint, idempotentHint). Risk levels are declared (READ_ONLY, WRITE) but not in the standard MCP annotation format. This prevents proper protocol adherence and LLM awareness of tool safety characteristics.
No error handling guidance documented. Source code shows financial validation and compliance checking logic, but no explicit error messages returned.
Parameters across tools lack type specificity and constraints. generate_invoice accepts 'invoice_data' as a generic object with nested properties, but no enum validation for template_name (claimed to be 'standard, detailed, simple, or professional' in description only, not in schema). create-invoice amount parameter has no numeric constraints (min/max). list-invoices hotel_id is untyped.
Duplicate/overlapping functionality between tool sets. Both agentic_invoice.py (analyze_invoice, generate_invoice) and invoices.py (list-invoices, create-invoice, mark-invoice-paid) manage invoices, but use different data models (invoice_document vs hotel_id+guest_name+amount). LLMs will struggle to determine which tool set to use and risk data inconsistency.