Single tool 'play-quiz' has a well-structured schema with Zod validation, comprehensive description (400+ chars), and proper parameter documentation. Tool uses tool annotations (readOnlyHint, destructiveHint, idempotentHint) and structured output via structuredContent. However, the tool combines quiz generation guidance with execution, and error handling is minimal (auto-fix silently modifies input rather than rejecting invalid data). Output schema is partially documented via structuredContent but lacks formal return type documentation. The description is lengthy and prescriptive (guidelines embedded) rather than concise.
Display an interactive quiz game. Generate well-crafted quiz questions, then call this tool. Guidelines: - Generate 5-10 questions per quiz (unless the user specifies a count) - Mix question types: mostly mcq (4 choices), with some true_false for variety - Each answer choice text: concise, under 60 characters - Write clear explanations for every choice (correct and incorrect) - Exactly one correct answer per question - Vary difficulty: start easier, get progressively harder - For document-based quizzes: ground questions in specific document facts - Avoid "all of the above" or "none of the above" answers The quiz renders as interactive mini-games (archery, puzzles, switches, etc.) — one unique game per question.
Tool description embeds 10+ lines of usage guidelines (question count, mix types, difficulty progression) instead of concisely stating what the tool does. Descriptions should be 10-1024 chars; this is ~450 chars but reads as a prompt template rather than a tool contract.
Error handling uses silent auto-fix (modifies input, logs to stderr) instead of rejecting invalid data with actionable error messages. When multiple correct answers exist, the tool silently keeps only the first instead of returning 'Multiple correct answers detected in question N. Exactly one correct answer required.'
Output schema (structuredContent) is not formally documented. The tool returns gameId, questions, title, templates, mobileTemplates but no schema definition explains these fields to the LLM. LLMs cannot infer downstream tool compatibility without documented return types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 72 | 2026-07-28+ | v2 |
Parameter 'questions' accepts 1-50 items but no guidance on optimal count. Description says 'Generate 5-10 questions' but schema allows up to 50, creating ambiguity about when to split into multiple quizzes.
Tool combines two concerns: quiz generation guidance (for the LLM) and quiz execution (server-side). The description should focus on execution contract; generation guidelines belong in system prompts or separate documentation.