A multi-stage agent system implementing Agent Oriented Architecture with MCP (Model Context Protocol) servers for product catalog management, inventory management, and multi-agent A2A (Agent-to-Agent) communication. Stage 1 provides MCP tools for product databases, Stage 2 adds SMOL Agent business intelligence tools, and Stage 3 implements A2A protocol for agent discovery and inter-agent communication.
This STDIO-only multi-stage MCP system exposes 15 tools across three product/inventory domains. While tool names follow verb-noun conventions well and descriptions are present for most tools, critical gaps emerge: parameter descriptions are frequently missing or trivial, input schemas lack full type definitions (nullable without clear alternatives), and output structures are not documented. Many tools are READ_ONLY but lack explicit tool annotations (readOnlyHint). Error handling is absent from tool definitions. The system is architectural sound (stages 1 - 3 demonstrate good separation of concerns) but individual tool definitions are below production baseline. Median tool score: 42/100.
Analyze price trends and statistics for products in a category
Check stock levels for a product across warehouses
Discover available agents in the system that you can communicate with
Find products similar to a given product based on category, price range, and rating
Get a list of all product categories with counts
Get a summary of inventory across all warehouses
Get the minimum and maximum prices in the catalog
Missing output schema documentation. No tool explicitly documents what fields are returned, their types, or structure. LLMs cannot plan downstream tool calls or extract required data without knowing output shape.
Parameter descriptions missing or trivial for optional parameters. 'category' in get_price_range and check_stock say only 'Filter by category (optional)' without explaining what happens if omitted, what values are valid, or how category names map to internal IDs.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Get detailed information about a specific product by ID
Get suggestions for products that need reordering
Get information about the product catalog database schema
Get warehouse information and capacity utilization
Predict which products will need restocking soon
Send a query to another agent and get their response
Execute a SQL query on the product catalog (read-only)
Search for products by name, category, or brand
No tool annotations despite all 15 tools being READ_ONLY. readOnlyHint should be set on every read-only tool to inform agents which tools are safe for retry and exploration, vs which require confirmation.
query_products accepts a raw 'query' parameter (SQL string) with no validation, constraints, or description of allowed SQL syntax. Exposed to SQL injection risk if LLM is tricked into malicious input. Description says 'read-only' but does not enforce it.
discover_agents and query_agent have minimal schema detail. query_agent accepts 'agent_name' as a free-form string with no enum, pattern, or discovery mechanism. If agent names are 'product-catalog', 'inventory-management', etc., these should be enums or linked to discover_agents output.
No error handling or recovery guidance in any tool definition. No description of what happens on failure, which errors are retryable, or what the LLM should do next. E.g., if get_product_by_id fails, should it retry, search, or abort?
find_similar_products (stage 2) wraps MCP tools and has incomplete schema definition in code. Visible schema in source shows product_id (int) and max_results (int, nullable) but lacks default value documentation and return type annotation.
warehouse_id parameter in check_stock is optional but no guidance on what 'all warehouses' returns if omitted. Does it sum stock across warehouses, return per-warehouse data, or fail?
Pagination and result limits not specified. search_products and analyze_price_trends provide no limit parameter, and descriptions don't state max result count. Returning hundreds of items risks context window exhaustion.