Contextual embedded insurance at deal redemption - MCP Server for GrabOn
GrabInsurance defines 4 tools with mixed quality. Two tools (classify_deal_intent, get_insurance_quote) have complete input schemas with typed parameters and descriptions, meeting baseline requirements. However, descriptions are underdeveloped (60-100 chars), falling below the production baseline of 194 chars average. No output schemas are documented in the visible source code. Tool naming follows verb_noun convention (classify_, get_), which is good. The `insurance_catalog` tool is listed as a resource, not a tool, creating inconsistency. Error handling is mentioned in docstrings (5-second timeouts, graceful fallback to Claude) but not reflected in actual error recovery guidance in tool descriptions. Parameter descriptions lack detail about constraints, formats, and valid ranges. The server shows foundational competence but lacks the depth required for production recommendation.
Classify deal intent and return top insurance products. Takes a deal object and returns the top 2 insurance products with confidence scores. Uses rule-based classification first, falls back to Claude API if category is unknown.
Generate personalized insurance copy for a deal. Creates a prompt that instructs Claude to generate one copy string under 120 characters. The copy mentions the deal amount and premium explicitly.
Calculate premium quote for an insurance product. Returns a premium quote based on the product's base rate, deal value, and user's risk tier. Premium is floored at Rs 19 and capped at Rs 499.
Return the full insurance product catalog. This resource provides Claude with complete product context including all 8 insurance products with their triggers, rates, and descriptions. Load this before performing any classification.
No documented output schemas. All four tools lack explicit response schema definitions. LLMs cannot know what fields to expect or how to chain results to downstream tools. This violates the baseline expectation that 100% of A+ tools have documented return types.
Tool descriptions are 50-100 characters, well below the baseline of 194 chars average for production tools. Descriptions lack WHEN/WHY context. For example, 'Classify deal intent and return top insurance products.' does not explain when to call this vs other tools, what prerequisites exist, or what failure modes to expect.
Parameter descriptions lack actionable constraints. 'category' accepts 'One of travel, electronics, food, health, fashion' but this is a prose statement, not a schema constraint. 'user_history' is described as 'Optional dict with risk_tier, total_purchases, categories_bought' but no schema validation is visible. LLMs cannot validate inputs without formal constraints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
insurance_catalog is defined as a resource (READ_ONLY), not a tool, but is listed in the TOOLS array. This creates confusion about whether it should be invoked as a tool or fetched as a resource. Resource definitions require explicit URI/description format, which is not visible in the source.
No error recovery guidance in tool descriptions. The source mentions 'graceful fallback' and '5-second timeout on Claude API calls' but tool descriptions do not state what happens on timeout, what errors are retryable, or what the LLM should do next. Per pattern:recovery-guide, error responses must tell the LLM what to do.
Parameter type constraints not enforced in schema. 'deal_value' is type 'number' with no minimum/maximum bounds. 'risk_tier' has description mentioning valid values (low, medium, high) but no enum constraint in schema. No validation rules are visible in the source code.
Missing chaining IDs in response design. If classify_deal_intent returns product recommendations, does it include product_id, product_name, and deal_value in the response? If get_insurance_quote returns a quote, does it include product_id and deal_value? Source code does not show response structures, making it impossible to verify if downstream calls can be chained.