Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
This server has significant structural and documentation gaps. Of 20 tools, only 3 have complete schemas visible in the provided code (ontology.explain_discount, ontology.normalize_product, ontology.validate_order). The remaining 17 tools show parameter definitions in the inventory but NO visible implementation code, input validation, or output schemas. Descriptions are present but often generic and under-optimized for LLM decision-making. The code snippet cuts off mid-implementation, making full assessment impossible. Critical issues: (1) No tool annotations (readOnlyHint/destructiveHint/idempotentHint) despite clear risk stratification in the inventory. (2) Error handling is minimal, _parse_order_id shows input validation but most tools lack equivalent guards. (3) Output schemas are not documented anywhere, the code shows payloads being assembled but no formal schema declarations. (4) Parameter descriptions in the inventory are minimal (e.g., 'The user ID' for user_id) and do not meet the 10 - 1024 character quality baseline. Naming is consistent (verb_noun pattern) but several tools conflate concerns: create_order + create_support_ticket + process_return are each complex multi-step operations that should be decomposed.
Missing implementation code for 17 of 20 tools. Only the first 3 ontology.* tools have visible implementation in tools.py; the remaining 17 commerce.* tools show only parameter definitions in the inventory. No actual tool handlers, validation logic, or output schema definitions are visible in the provided code. Cannot verify that these tools are actually registered or functional.
Add tool annotations to all 20 tools. Map the inventory's Risk column to MCP annotations: READ_ONLY → readOnlyHint=true, WRITE → destructiveHint=true, REVERSIBLE → (destructiveHint=true AND idempotentHint=true or map to a confirm-before-execute flow). This is REQUIRED for agent planning.
Document output schemas for every tool. Define what fields each response contains, their types, and their meaning. Example for ontology.explain_discount: {discount_applied: boolean, discount_rate: number (0-1), rule_source: string}. This is critical for agent reasoning and chaining.
Expand parameter descriptions from 1-sentence stubs to 50 - 150 character actionable text. Example: Instead of 'The user ID', write 'The internal user ID (integer > 0) or username (string). If unsure, call search_users() first.' This guides the LLM on formats, constraints, and fallback actions.
Add enums to categorical parameters. Examples: process_payment payment_method should be enum [credit_card, debit_card, paypal, alipay]; create_support_ticket priority should be enum [low, medium, high, urgent]; process_return return_type should be enum [refund, exchange, store_credit].
Implement input validation for all commerce.* tools, following the pattern in _parse_order_id. Validate user_id, product_id, quantities (>0), amounts (>=0), dates (ISO 8601), and categorical fields against allowed enums. Return actionable error messages, not stack traces.
Add pagination to list/search tools: commerce.search_products, commerce.get_product_reviews, commerce.get_product_recommendations should accept (limit: 1 - 100, offset: 0+) or (limit, cursor) and return {results: [...], total: int, next_cursor: string|null}. This prevents context bloat.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 10 points across a rubric change (v1 → v2)
45/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
45
<=2025-11-25
v2
2026-03-09
F
35
-
v1
read only
48/100
获取订单的物流状态
commerce.get_user_ordersread only50/100
获取用户订单列表,支持按状态筛选
commerce.process_paymentwrite50/100
处理订单支付
commerce.process_returnreversible50/100
处理退货/退款请求
commerce.remove_from_cartwrite50/100
从购物车移除商品
commerce.search_productsread only50/100
搜索商品,支持关键词、分类、品牌、价格范围过滤
commerce.track_shipmentread only48/100
追踪物流信息
commerce.view_cartread only47/100
查看购物车
ontology.explain_discountread only50/100
解释折扣规则,基于 VIP 状态和金额返回折扣率及规则来源
ontology.normalize_productread only50/100
将产品文本进行同义词归一化处理
ontology.validate_orderread only50/100
使用 SHACL 校验订单数据
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present in the code, despite clear risk stratification in the inventory (READ_ONLY, WRITE, REVERSIBLE). LLMs cannot determine which tools are safe to retry, which require confirmation, or which are read-safe without these hints. Required for proper agent planning.
Output schemas are not documented anywhere. The code shows tools assembling result dictionaries (e.g., ontology.explain_discount returns {discount_applied, discount_rate, rule_source}) but no formal schema definitions or documentation. Agents cannot reason about what fields to extract or plan downstream calls.
Parameter descriptions are minimal and generic. Examples: 'The user ID', 'The product ID', 'The order ID'. These fall well below the 72-character baseline for parameter annotations and do not include type constraints, valid ranges, or guidance on when/why to use the tool. LLMs cannot infer intent or validation rules from such sparse text.
No enum constraints on categorical parameters. Examples: commerce.create_order has no format enum for 'shipping_address' type; commerce.process_payment and commerce.create_support_ticket have no enums for 'payment_method', 'category', or 'priority'. Free-form strings invite hallucinated values. LLMs should be constrained to valid options.
Error handling is minimal. Only _parse_order_id includes validation logic; most other tools (17 out of 20) show no visible input validation, error recovery guidance, or actionable error messages. If a tool fails, the agent has no guidance on what to try next.
Tool descriptions are generic and do not guide LLM selection. Examples: 'Add product to cart', 'Get product detail', 'Process payment'. These are literal name expansions, not LLM-optimized descriptions. They lack context on when to use each tool, what dependencies exist (e.g., 'Call search_products first if you don't have a product_id'), or what the output will contain. Baseline expectation is 194 characters average; most commerce.* tools are under 80 characters.
Potential composition issues: commerce.create_order, commerce.create_support_ticket, and commerce.process_return appear to be multi-step operations that should be decomposed. For example, create_order likely needs to validate stock, apply discounts, and lock inventory, these should be separate tools so the agent can compose and retry them independently.
Parameter names lack type suffixes. Examples: 'product_id', 'user_id', 'order_id' are good, but generic 'status', 'category', 'type' could be disambiguated with context. Best practice: when a field could accept multiple types, use suffixes (e.g., category_name vs category_id, status_code vs status_text).
No pagination support visible. commerce.search_products accepts 'limit' but no 'offset', 'page', or 'next_cursor'. commerce.get_product_reviews accepts 'limit' but no pagination tokens. Without pagination, the agent cannot reliably retrieve large result sets, and all results are dumped into context at once, which can exhaust the context window.
Provide tool implementation code or at least visible registration for the 17 commerce.* tools. Currently, only the first 3 tools have visible call_tool handlers. The rest are undocumented. Either show the full tools.py or confirm they are implemented in a separate module.
Expand tool descriptions beyond name echoing. Provide LLM-optimized guidance: 'When to use', 'What it returns', 'Common next steps'. Example: 'Search for products by keyword, category, brand, and price range. Returns up to [limit] products with id, name, price, and availability. Use get_product_detail for full specs. Supports pagination via offset/limit.'
Decompose multi-step operations. If commerce.create_order internally calls check_stock, apply_discount, and lock_inventory, expose these as separate tools so agents can retry and compose them independently. This is especially important for WRITE operations that may partially fail.
Add recovery guidance to error responses. When a tool fails, respond with: 'X failed because [reason]. Try [alternative tool] or [parameter adjustment] next.' Examples: 'Product not found. Try search_products() with a keyword.' or 'Payment failed. Verify amount and payment_method, then retry.'
Define and enforce minimum/maximum for numeric parameters. Examples: limit (1 - 100), quantity (>0), amount (>=0), page (>=1), days (1 - 365). Document these constraints in parameter descriptions and validate on input.
Complete the tool implementation in the provided code snippet (tools.py cuts off at line 38). Show the full call_tool routing, including handlers for commerce.* tools, validation logic, error handling, and response formatting.
Ensure tool response fields include all IDs needed for chaining. If commerce.search_products returns {id, name, price}, and the next call is get_product_detail(product_id), ensure the 'id' field is present and named 'product_id' to match the next tool's parameter.
Adopt tool annotations NOW, not as a future enhancement. This is a blocker for safe agent planning, agents cannot safely use WRITE and REVERSIBLE tools without knowing which ones require confirmation or can be retried.