An MCP server for MCP client benchmarking
The server defines 7 tools with explicit schema registration via the mcp-sdk. Tool names follow verb_noun conventions (start_benchmark, choose_food_category, select_menu, submit_reservation_details, get_confirmation_email, verify_confirmation_code, try_again). However, there are significant gaps: (1) descriptions are minimal to generic (most under 100 chars), lacking WHEN/WHY context for LLM selection; (2) parameter descriptions are sparse, many parameters lack domain guidance (e.g., 'confirmationCode' has only 'The confirmation code sent in your email', not explaining why 6 chars, alphanumeric vs mixed, what happens on mismatch); (3) no documented output schemas, LLMs cannot predict response structure for chaining; (4) no error handling guidance, no recovery hints if verification fails or category is invalid; (5) no resource context, the benchmark flow is opaque without understanding state transitions. The codebase shows sophisticated state management (AbstractBenchmarkState, StateMachine) and input validation (Zod schemas), but these are buried in implementation; tool descriptions do not surface this for the LLM. Average tool description length ~65 chars (well below the 194-char production baseline). Naming is acceptable but generic ('try_again' could be 'reset_benchmark_session'). Schema presence is good (all tools have Zod-defined input), but descriptions for parameters are inconsistent, some tools (submit_reservation_details) have detailed parameter constraints; others (start_benchmark, get_confirmation_email) have empty/trivial descriptions.
Select a food category for the reservation.
Generates and returns the text for a confirmation email for your reservation.
Select a specific restaurant menu.
Begin the MCP benchmark evaluation.
Submit the final details for the reservation.
Resets the session and starts a new benchmark run.
Submit the confirmation code from the email to finalize the benchmark.
Tool descriptions lack WHEN/WHY context for LLM selection. Descriptions are 30-60 chars; baseline is 194 chars. 'Begin the MCP benchmark evaluation' and 'Resets the session and starts a new benchmark run' do not explain prerequisites, state transitions, or when to invoke. LLMs cannot distinguish similar tools or reason about flow.
No documented output schemas. Tools return CallToolResult with unspecified content. LLMs cannot predict response fields for downstream chaining (e.g., does choose_food_category return available menus? a state acknowledgment?). Forcing inference.
Parameter descriptions are minimal or missing domain context. 'confirmationCode': 'The confirmation code sent in your email' does not specify format (alphanumeric? uppercase?), error handling (invalid code behavior?), or retry logic. 'menu_id': 'Restaurant menu identifier' lacks guidance on format or valid ranges.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
No error handling guidance. Tools lack recovery hints. If verify_confirmation_code fails, does the LLM retry? Request a new email? Abandon flow? Code shows validation (Zod safeParse) but tool description does not document failure modes or next steps.
Tool 'try_again' has a weak name. 'try_again' is vague, does it retry a failed step, reset the entire session, or abort? 'reset_benchmark_session' or 'restart_benchmark' would be clearer. Current name violates verb_noun convention.
State transitions are implicit. The flow expects agents to call tools in a specific order (start → category → menu → details → email → verification). Tool descriptions do not document prerequisites or valid successor tools, forcing the LLM to guess or rely on trial/error.