A professional hotel room service ordering assistant that manages menu items, checks availability, and maintains conversation history for room service orders.
This hotel room service agent exhibits significant quality gaps typical of early-stage community servers. While all 4 tools are explicitly registered with FastMCP and have basic descriptions and schemas, there are critical issues: (1) Parameters lack type declarations in schema definitions, with 'preferences' in get_menu() being an undeclared object and no 'type' field specified; (2) Descriptions are minimal (10-60 chars) and lack crucial context about WHEN to use each tool or HOW to invoke them correctly; (3) Output schemas are not documented, callers have no way to know what structure to expect from responses; (4) Parameter descriptions are generic ('The session identifier', 'Maximum number of recent messages') without constraints or examples; (5) The tool composition mixes stateful session management (save_conversation, get_conversation_history) with read-only domain queries, creating unclear separation of concerns; (6) Error handling is minimal, check_item returns {"error": "Item not found"} with no guidance on recovery; (7) No validation of dietary preference enums (vegetarian/vegan are boolean flags but no enum documentation). The per-tool analysis reveals all tools have visible schema and description, but depth is insufficient for production LLM integration.
Check item availability and preparation time.
Retrieve recent conversation history for context.
Get menu items filtered by dietary preferences (vegetarian, vegan).
Save conversation message for context retention.
Missing output schemas for all 4 tools. Tools return structured data but LLMs have no formal declaration of response structure, field names, or types. This forces LLMs to infer shape from code execution, risking hallucinated field names or type mismatches.
Parameter schemas lack type field declarations. get_menu 'preferences' parameter is described as object but the JSON Schema definition does not explicitly include 'type': 'object' at the top level. 'vegetarian' and 'vegan' lack 'type' declarations in their property definitions.
Descriptions are below the 50-200 character baseline for LLM-optimized tool docs. All 4 tools have descriptions under 65 characters, missing context on WHEN to call the tool, what it depends on, and what to expect. Per pattern guidance, descriptions should answer: What does it do? When should I use it instead of a similar tool? What does it return?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 40 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
No enum constraints on categorical parameters. 'role' in save_conversation should be declared as enum ['user', 'assistant'], not a free-form string. 'vegetarian' and 'vegan' should be documented as boolean enums or clarified in parameter descriptions.
Minimal error handling and no recovery guidance. check_item returns {"error": "Item not found"} but does not guide the LLM on next steps (suggest alternative items? retry with a different ID?). save_conversation and get_conversation_history offer no error cases at all in the tool description.
Parameter constraints and validation rules not documented. 'limit' in get_conversation_history has a default but no bounds (min=1, max=?). 'session_id' format is undefined, is it a UUID, integer, or free-form string? This invites invalid LLM calls.
Unclear tool composition and overlapping concerns. save_conversation and get_conversation_history manage session state, while get_menu and check_item query domain data. No clear guidance on which tools to call in sequence or how they interact. This forces LLMs to reason about call ordering without explicit dependencies.
Response structure mismatch in check_item. Error case returns {"error": "..."} while success case returns {"available": true, "name": "..."}. This inconsistency forces LLMs to handle two different response shapes, increasing hallucination risk.