Serverless API and Model Context Protocol (MCP) integration to order pizzas from AI agents.
The Pizza MCP server defines 9 tools with reasonable naming conventions and mostly complete schemas. All tools use verb_noun naming (get_*, place_*, delete_*) which follows the agentic pattern. However, descriptions are brief and lack strategic guidance for LLM tool selection. Parameter documentation is present but inconsistent in depth. Output schemas are not explicitly documented. Error handling guidance is absent. The server demonstrates foundational quality but lacks the polish and completeness expected of production-grade tools.
Cancel an order if it has not yet been started (status must be "pending", requires userId)
Get a specific order by its ID
Get a list of orders in the system
Get a specific pizza by its ID
Get a list of all pizzas in the menu
Get a specific topping by its ID
Get a list of all topping categories
Output schemas are not documented. LLMs cannot infer what fields to expect from tool responses, forcing them to guess at response structure and increasing error rates in downstream reasoning.
Tool descriptions lack strategic context for LLM selection. Descriptions are 30 - 65 characters; baseline for A-grade tools is 50 - 200 chars. For example, 'Get a list of all pizzas in the menu' (38 chars) does not explain when to call this vs similar tools or what the response structure contains.
No error handling guidance. Error responses do not tell the LLM what to do next. For example, if 'get_pizza_by_id' fails with a 404, the LLM has no guidance to call get_pizzas() to discover valid IDs or understand whether to retry.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 76 | <=2025-11-25 | v2 |
| 2026-03-09 | B | 72 | - | v1 |
Get a list of all toppings in the menu
Place a new order with pizzas (requires userId)
Parameter descriptions for optional fields lack detail about expected behavior when omitted. For example, 'category' in get_toppings is marked optional with description 'can be empty', but it is unclear what 'empty' means (null vs empty string vs omitted) and what the response is when omitted (all toppings? or no toppings?).
place_order and delete_order_by_id describe side effects but lack confirmation or dry-run patterns. No mention of idempotency. If an agent retries a failed place_order call, will it create a duplicate order? This ambiguity invites accidental duplicate orders.
get_orders 'last' parameter accepts human-friendly formats ('60m', '2h') but no validation rule is documented. What happens if an LLM passes '2days' or 'invalid'? No guidance on error recovery.
place_order allows arbitrary quantity (minimum 1) but no maximum is specified. An LLM could order 1,000,000 pizzas. No guard against runaway agent loops.