MCP server for TutorClaw — an AI-powered tutoring system that teaches programming through PRIMM-Lite pedagogy, adaptive learning, and tiered content access.
TutorClaw demonstrates solid definition quality with well-structured tool names, comprehensive descriptions, and complete input schemas. All 9 tools follow verb_noun naming conventions (register_, get_, update_, submit_, generate_, assess_). Descriptions are detailed and include WHEN/NEVER guidance (50-250 chars, well above the 10-char minimum). All parameters have type definitions and descriptions. Tool annotations are properly declared. However, output schemas are not documented in the source, parameter constraints could be more explicit (e.g., stage enum values), and error handling guidance is minimal. The pedagogical domain is well-served, but production-grade error recovery patterns are underdeveloped.
Score a learner's answer against expected concepts and return a confidence adjustment, feedback, and next-step recommendation. WHEN to call: When a learner submits a natural-language answer to a predict, run, or investigate prompt and you need to evaluate their understanding. NEVER call for general conversation — only use when the learner is answering a PRIMM-Lite stage question. Use submit_code for code submissions. Related: submit_code (execute code, not text answers), update_progress (apply the returned confidence_delta and recommendation).
Prepare the chapter content excerpt and teaching instructions for the learner's current PRIMM-Lite stage. WHEN to call: After fetching chapter content with get_chapter_content, to get stage-appropriate code excerpts and a system prompt for the predict, run, or investigate stage. NEVER call for raw content — use get_chapter_content to fetch chapter text first, then pass it here. Related: get_chapter_content (fetch chapter_content input), assess_response (evaluate the learner's answer after guidance is delivered).
Fetch the markdown content for a chapter, optionally narrowed to a specific section. WHEN to call: When the learner needs to read or study chapter material, or when generate_guidance needs chapter_content as input. NEVER call for exercises or practice — use get_exercises to retrieve practice problems instead. Related: get_exercises (practice problems), generate_guidance (needs this tool's output as input).
Return practice exercises for a chapter, optionally filtered to specific topics. WHEN to call: When the learner is ready to practice, especially in the modify or make PRIMM-Lite stages, or when targeting weak_areas. NEVER call for reading material — use get_chapter_content to retrieve chapter text instead. Related: get_chapter_content (reading material), assess_response (evaluate exercise answers).
Output schemas not documented. Tool descriptions explain inputs but do not specify what fields/structure the LLM should expect in responses. This forces LLMs to guess at downstream field names and risks broken tool chains.
Stage parameter in update_progress and generate_guidance accepts free-form strings ('predict', 'run', 'investigate', 'modify', 'make') but is not declared as an enum. LLMs may hallucinate invalid stage values. Constraint is buried in store.py validation, not visible in tool schema.
Error handling lacks recovery guidance. Tool descriptions do not explain what errors are possible, how to classify them (retryable vs. fatal), or what the LLM should do next. E.g., 'learner not found' should suggest calling register_learner.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 83 | 2026-07-28+ | v2 |
Return the current tutoring progress for a learner, including chapter, stage, confidence, and tier. WHEN to call: Before any tutoring action to check the learner's current chapter, stage, confidence, tier, and remaining exchanges. NEVER call to change progress — this is read-only. Use update_progress to advance chapter or stage. Related: update_progress (write progress), register_learner (create new learner).
Return a Stripe checkout URL to upgrade a free-tier learner to the paid plan. WHEN to call: When a free-tier learner hits a content gate (chapter > 5) or exhausts daily exchanges, and needs an upgrade link. NEVER call for paid-tier learners — they already have full access. Use get_learner_state to check tier first. Related: get_learner_state (check current tier before calling).
Register a new learner and return their ID, API key, and welcome message. WHEN to call: At the very start of a conversation when a user wants to begin learning and has no learner_id yet. NEVER call for existing learners — use get_learner_state to look up a known learner_id instead. Related: get_learner_state (read existing learner), update_progress (change progress).
Execute a learner's Python code in a sandbox and return stdout, stderr, and execution status. WHEN to call: When a learner submits Python code to run, typically during the modify or make PRIMM-Lite stages. NEVER call for non-code messages — use assess_response to evaluate natural-language answers instead. Related: assess_response (evaluate text answers), get_exercises (get problems that may require code submissions).
Advance a learner to a new chapter and stage, and apply a confidence score adjustment. WHEN to call: After assess_response returns a recommendation to advance, or when the tutor decides to move the learner forward or back. NEVER call to check current state — this is write-only. Use get_learner_state to read progress first. Related: get_learner_state (read current progress), assess_response (get recommendation before updating).
submit_code tool accepts arbitrary Python code with only a runtime sandbox restriction (no os/subprocess/shutil/open). No mention of execution timeout, memory limits, or what happens if code runs forever. Agents need explicit bounds.
Tool composition: get_chapter_content and generate_guidance are tightly coupled (generate_guidance requires chapter_content as input). No tool chains documented; LLMs must infer the sequence. Consider adding a 'prepare_lesson' wrapper or explicit dependency hints.