MCP server for a cafeteria point-of-sale system providing sales, expense, and cash control reporting and analytics
Strong tool design with clear naming, comprehensive descriptions, and well-structured schemas. All 7 tools follow verb_noun convention (get_*, list_*). Descriptions are detailed (150-300 chars) and include WHEN to use each tool and WHY it's preferred over alternatives. Input schemas are complete with type definitions and parameter descriptions. However, output schemas are not formally documented in the code, only inferred from JSON responses. Error handling is minimal; no recovery guidance or actionable error messages visible. No tool annotations (readOnlyHint, etc.) despite all being read-only operations. Pagination is implemented (limit/offset) but truncation handling could be more explicit in all tools.
Cash-drawer health for a period: per-cashier shift counts, how many closes came up short or over, net and total shortfall, worst single difference — plus any days that took money with no cierre de caja recorded. Use this for 'were there shortfalls', 'which cashier has the most differences', 'did they close the till', or any question about missing cash. A single small difference is normal cash handling; the pattern across shifts is what matters. A day with sales and no close is itself a red flag.
One row per day: sales total and count, expenses total and count, and net. Use this for any question spanning more than a couple of days — trends, best/worst days, 'how did the month go', comparing weeks. It stays compact over long ranges where listing individual transactions would not fit. Drill into a specific day afterwards with list_sales if the user asks what was actually sold. Days are business-local (America/Bogota).
Total expenses and breakdown by category for a date range. Use this for 'what did I spend' questions rather than listing individual expenses.
Transactions and revenue grouped by hour of the day across a date range, in business local time. Use this for staffing, opening hours, prep timing, and 'when are we busiest' questions. Hours with no sales are omitted. At most 24 rows however long the range. This is the only view of time-of-day — no other tool exposes it.
Output schemas not formally documented. Tool descriptions explain what is returned (e.g., 'totals for a date range, transaction count, breakdown by payment method'), but the actual JSON structure is not declared in code. LLMs cannot plan downstream operations without knowing field names and types.
No tool annotations despite all tools being read-only. MCP spec supports readOnlyHint to signal safe tools. This helps agents reason about side effects and retry safety.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 87 | 2026-07-28+ | v2 |
Per-product units sold, revenue, share of revenue, and number of orders containing it, ranked by revenue. Use this for 'what sells best', 'what makes the most money', 'what should I drop from the menu', or questions about a specific dish. Prefer this over get_sales_summary's top_items, which ranks by quantity only and so overstates cheap items. Revenue is what the customer paid; it is not profit, since item costs are not recorded.
Totals for a date range: total sales, transaction count, breakdown by payment method, and top 5 items by quantity. Use this for 'how were sales' questions — it is far cheaper than listing rows. Dates are business-local days (America/Bogota).
Individual shifts for a date range: cashier, opening float, cash sales, cash movements, expected versus counted, the difference, and any note left at closing. Use this when the user wants the specific shifts rather than the pattern — 'show me yesterday's closes', 'what happened on María's shift'. For 'were there shortfalls' use get_cash_control. Returns at most 100 rows; check `truncated`.
Error handling not visible in tool implementations. No recovery guidance (e.g., 'if date range is invalid, try narrowing to a single day'). No categorization of errors as retryable vs. user-fixable. Agents cannot self-correct on failures.
Truncation handling inconsistent. get_daily_totals and get_product_performance return truncated flags, but get_sales_summary, get_expense_summary, get_hourly_pattern, get_cash_control do not. Agents cannot reliably detect incomplete result sets.
No batch operations. If an agent needs to fetch sales summaries for 10 different date ranges, it must call get_sales_summary 10 times sequentially. A batch variant would reduce latency and token overhead.