MCP server for automating Zepto cafe orders using Playwright browser automation. Provides tools for logging in, adding items to cart, handling OTP verification, selecting addresses, and completing orders.
This MCP server exhibits multiple critical gaps across naming, schema completeness, and error handling. While 8 of 10 tools have descriptions, descriptions are often vague and under-optimized for LLM reasoning. Parameters lack type consistency, and output schemas are not documented. The server conflates multiple responsibilities in single tools (e.g., add_to_cart handles both lookup and OOS decisions), and several parameters use free-form strings where enums are required. No evidence of input validation, error recovery guidance, or permission scoping. The codebase shows promise (well-organized REST API wrapper with Playwright), but the tool interface itself falls well below production standards for agentic use.
Add a product to the cart by product name or URL. Handles out-of-stock items by returning options for user decision.
Get the list of available products in the Zepto Cafe catalog.
Get the current status of the order (idle, waiting_for_login, waiting_for_otp, item_added, out_of_stock, address_selected, processing_payment, completed).
Handle out-of-stock product decisions: cancel order, proceed with remaining items, or replace with alternatives.
Log in to Zepto with a phone number and wait for OTP verification. Returns a session token.
Proceed to payment page after address selection. Initiates payment flow.
Select or provide a delivery address for the order.
add_to_cart combines multiple responsibilities: product lookup, cart manipulation, and OOS handling. This violates the single-responsibility principle and forces the agent to make unclear decisions about when to call it vs. handle_out_of_stock. Should split into separate tools: add_product_to_cart (just adds) and discover_product (searches catalog with fuzzy matching).
No output schemas documented. get_order_status describes states but does not specify the response format (does it return {status: string}, {status, items, total}, etc.). get_catalog returns 'available products' but LLMs cannot reason about the structure without a documented schema. This forces agents to guess field names and causes downstream tool call failures.
Decision parameters (handle_out_of_stock.decision, select_address.address) use free-form strings instead of enums. 'decision' accepts 'cancel', 'proceed_with_remaining', or 'replace_items' but this is not enforced in schema, LLMs can hallucinate invalid values like 'ignore' or 'postpone', causing silent failures. Use enum constraint.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Stop the current order and reset the automation state.
Submit the OTP received during login to complete authentication.
Submit the OTP received for payment confirmation.
No error handling guidance. What happens if login fails (invalid phone, network timeout)? What happens if add_to_cart tries to add an out-of-stock item that has no alternatives? No error responses document recovery steps, retryability, or user-fixable issues. Agents will halt without guidance.
Descriptions are vague and do not explain WHEN to use tools or what to do NEXT. 'Log in to Zepto with a phone number and wait for OTP verification. Returns a session token.' does not say: is this required before add_to_cart? What does the caller do with the session token? Does it need to be passed to other tools? Descriptions should guide composition.
Parameters lack required descriptions or type specificity. add_to_cart.quantity has no min/max constraint (default 1 is good, but what if LLM passes quantity=1000?). select_address.address is a string but does not clarify: is this a saved address name like 'home', or a full address string like '123 Main St'? replacement_items is an array but does not describe the array element structure.
No permission scoping or audit trail. Tools perform sensitive operations (payment, OTP submission, address selection) but do not declare what permissions they require or log who called what. No evidence that credentials are server-injected vs. passed as parameters. Payment tools should gate behind user verification.
Tool composition broken across login flow. login returns a 'session token' but no other tools document accepting it as a parameter. If the token is required for add_to_cart and other operations, this dependency must be explicit in parameter descriptions. Current structure forces agent to guess whether to retry login or pass token.
get_catalog returns a hardcoded PRODUCT_CATALOG dict (visible in zepto_api_server.py) with ~40 products, but the tool description just says 'Get the list of available products in the Zepto Cafe catalog.' No mention of pagination, limits, or how the agent discovers products by name vs. URL. If catalog is static, say so; if dynamic, document how to filter/search.
No idempotency guarantees. add_to_cart, submit_payment_otp, and submit_login_otp modify state but do not declare whether they are idempotent. If an agent retries submit_payment_otp on a flaky network, does it double-charge? This must be explicit.