A multi-agent system using CrewAI for function calling with MCP (Model Context Protocol) integration. Handles product inquiries, inventory checks, and order placement through coordinated agents.
The server defines 5 tools with significant quality gaps. Tool definitions are present in mcp_server.py and multi_agents/tools/ but lack critical schema rigor, parameter validation, and error recovery guidance. Naming is acceptable but descriptions are generic. Parameter schemas exist but lack formal constraints (enums, ranges, patterns). Error handling is minimal, responses provide status codes but no actionable recovery steps. No tool annotations (readOnlyHint/destructiveHint/idempotentHint). Output schemas are documented in code comments but not formally declared. The server exposes write operations (create_order) without confirmation patterns or dry-run support. No pagination support for potentially large result sets. Composition is reasonable (each tool has one responsibility), but chaining requires the LLM to infer field mappings.
Retrieves inventory details from storage based on the product. Input is a JSON string or object with product name, and optionally storage and color.
Saves order data to a file in the 'orders' subdirectory with a standardized format. Input is a JSON string from SaveOrderInput model.
Saves the given order data (dictionary) to a file in the 'orders' subdirectory, with a standardized format. Returns a success message with the filename or an error message.
Retrieves the content of an order file by its order_id. Returns a dictionary with file_content or error message.
Retrieves inventory details from storage based on the product. Input is a JSON string or object with product name, and optionally storage and color.
Duplicate tools with inconsistent schemas. 'create_order' appears twice: once in mcp_server.py accepting dict, once in multi_agents/tools/create_order.py accepting JSON string. This violates single-responsibility and forces LLMs to reason about which variant to use.
No input schema validation constraints. Parameters like 'product', 'storage', 'color' in get_product_info are free-form strings. No enum, min/max length, regex pattern, or required/optional distinction. This invites hallucinated values (e.g., 'color' → user passes 'red' when system uses 'Titan tự nhiên').
Destructive write operation without confirmation. 'create_order' modifies file system (creates orders/) with no dry-run, confirmation step, or preview. Agents can accidentally create erroneous orders without review.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Error responses lack recovery guidance. get_order returns {"error": "Order file with ID ... not found", "status": 404} but does not suggest next steps: 'Try search_orders() to find similar order IDs' or 'Available order IDs: [...]'. This forces the LLM to backtrack without hints.
Tool descriptions are under 100 characters and lack WHEN/WHY context. 'Retrieves the content of an order file by its order_id.' does not explain when to call this instead of a batch retrieval, what format the returned file_content is in, or how to handle parsing. Descriptions should be 50-200 chars and answer: what, when, and format of output.
Output schema not formally declared in MCP metadata. Code comments mention return types (dict, str, JSON string) but these are not exposed as MCP tool output schemas. LLMs cannot plan downstream calls without knowing fields returned.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Agents cannot infer which tools are safe to retry. create_order should be marked destructiveHint=true; get_order should be marked readOnlyHint=true. Without annotations, agents must guess.
Parameter descriptions lack format/constraint details. 'order_id' in get_order is described as 'The order ID to retrieve' but does not specify format (UUID? alphanumeric? length?). 'storage' in get_product_info says '(optional)' but does not list valid values like '256GB', '512GB', '1TB'.
No pagination support. get_product_info and Check inventory detail may return large result sets (e.g., 100s of product SKUs). No limit, offset, or page_size parameters. Response includes raw product arrays that could blow context window.
MongoDB connection errors not gracefully degraded. If db_client is None in get_product_info, the tool returns error JSON but create_order and get_order do not use MongoDB, they operate on the file system. Inconsistency: some tools fail if MongoDB is down, others succeed. Should fail consistently or offer offline modes.
No schema validation in tool code. create_order checks for required fields at runtime but does not use Pydantic models in the MCP schema definition. Other tools (get_product_info, get_order) rely on the LLM to pass valid types. A malformed or missing parameter could cause a runtime error instead of a clear validation error.