Observable e-commerce order agent with RAG, memory, MCP tools and human approval. Multi-module Spring Boot application featuring order query, refund eligibility evaluation, and agent-based order operations via MCP protocol.
This MCP server exhibits significant gaps in definition quality. While two tools are present with basic descriptions and partial schemas, they lack the rigor expected for production use. Tool names are noun-heavy rather than verb-action focused (ORDER_QUERY, REFUND_ELIGIBILITY vs. query_order, check_refund_eligibility), making intent ambiguous. Descriptions exist but are minimal (89-131 chars) and lack clear WHEN/WHY guidance. Parameter descriptions are present but several critical constraints are missing: no enums for categorical inputs like reasonType, no explicit format/pattern constraints, no bounds on arrays. Output schemas are not documented anywhere in the provided source. The parameter relationships (e.g., customerOpened/customerUsed/conditionStatus forming a composite condition) are undocumented. No error handling guidance is visible. The schema for ORDER_QUERY appears incomplete, 'query' is free-form string, forcing LLM to construct natural language rather than structured queries. REFUND_ELIGIBILITY has 8 parameters, several redundant: both 'query' and 'orderId' suggest confusion about the primary input method.
订单查询工具:调用 mall-order 服务(/orders)并将结果写入 Graph 状态。Queries orders by order ID or user ID with authorization scope checking and persona-based masking.
判断退款资格:评估订单是否符合退款条件。Evaluates refund eligibility for an order based on reason type, product condition, and customer actions via MCP protocol.
Tool names violate verb-action convention. 'ORDER_QUERY' and 'REFUND_ELIGIBILITY' are noun-noun pairs. Should be verb-noun like 'query_order', 'check_refund_eligibility', or 'evaluate_refund_eligibility'. LLMs infer tool purpose from the action verb; noun-heavy names force them to read full descriptions, wasting reasoning cycles.
Output schemas are not documented. Neither tool specifies what fields it returns, their types, or structure. LLMs cannot plan downstream operations or extract required fields without documentation. This violates pattern:tool and pattern:response-shaper.
REFUND_ELIGIBILITY has redundant inputs: both 'query' (natural language) and 'orderId' (structured). Unclear which is primary or how they interact. Both are marked required without explanation. LLMs will become confused whether to pass one, both, or how to construct the natural language query.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 22 | - | v1 |
Parameter 'reasonType' in REFUND_ELIGIBILITY accepts free-form string with examples (NO_REASON, QUALITY_ISSUE, NOT_AS_DESCRIBED) embedded in description. Should be an enum constraint. Free-form strings invite hallucinated reason types; enums are self-documenting and machine-parseable.
No documentation of parameter relationships or dependencies. In REFUND_ELIGIBILITY, parameters 'customerOpened', 'customerUsed', 'conditionStatus', 'reasonDescription' form a logical group describing product state, but their interdependencies are unexplained. When are all required? When are some optional? LLMs will guess.
No error handling documentation. Tools lack guidance on what to do if order not found, refund ineligible, or service unavailable. Errors should be categorized as retryable, user-fixable, or fatal, with recovery guidance.
ORDER_QUERY parameter 'query' is free-form natural language with no format constraints. Description says it should contain order ID (ORD20250101120000) or user ID (USER1005), but examples are in description, not enforced as constraints. LLMs may pass malformed IDs or unrelated text.
No documentation of authorization or security scopes. ORDER_QUERY description mentions 'authorization scope checking' and 'persona-based masking', but no tool-level scope declarations (e.g., 'read:orders', 'read:customer-data'). Agents cannot understand permission requirements.
REFUND_ELIGIBILITY parameter 'evidenceUrls' is an array with no bounds. No minimum/maximum length, no validation rules for URL format. LLMs could pass 100+ URLs, overwhelming the service or context window.