The tutor-mcp server exhibits significant gaps in definition quality. While 46 tools are registered, the vast majority lack complete input schemas visible in source code. Most tool descriptions are present but terse (averaging ~30-50 chars, well below the 194-char production baseline). Critical parameters lack type definitions and detailed constraints. The codebase shows tool registration in tools_register_test.go and some input contract testing in tool_input_contract_test.go, but the actual tool definitions and schemas are not visible in the provided source. Two tools (record_learning_event, record_session_close) show partial schema evidence in test files, but the remaining 44 tools lack visible schema documentation. Only tool annotations (destructiveHint, idempotentHint) are noted as implemented, suggesting some schema awareness, but without access to the tool definition code, schema completeness cannot be verified.
44 of 46 tools lack visible input schema definitions in provided source code. Only record_learning_event and record_session_close show schema evidence (enum constraints and object properties in test assertions). Cannot verify that 95% of tools have complete, properly-typed input schemas.
Create comprehensive input schemas for all 46 tools with explicit type definitions (string, integer, boolean, object, array) for every parameter. Current code shows only 2 tools with partial schema evidence.
Expand all tool descriptions to 100-250 characters, following the production baseline of 194 chars. Include WHAT the tool does, WHEN to use it (vs similar tools), and WHAT it returns. Examples: current 'Initiate a new learning session' → 'Start a new learning session to begin teaching a learner. Call this before recording interactions or assessment attempts. Returns session_id and start_timestamp.'
Add descriptions to every input parameter. For each tool, document parameter meaning, valid values (enums where applicable), format constraints, and whether each is required or optional. Example: 'learner_id (string, required): The unique identifier of the learner. Accepts email or system ID.'
Document return schemas for all tools, specifying field names, types, and descriptions. Include fields that downstream tools need to chain calls. Example: 'Returns {session_id: string, learner_id: string, start_time: ISO8601, status: enum[active,paused,closed]}'.
Add dry-run or confirmation modes to destructive operations: delete_domain, publish_curriculum_revision, update_learner_profile, learning_negotiation, set_domain_priority, mark_domain_high_stakes, update_learner_memory, update_implementation_intention. Allow agents to preview changes before committing.
Implement structured error responses with actionable recovery guidance. Map common failure cases to next steps. Example: 'Domain not found (id=xyz). Available domains: domain1, domain2. Try list_domains() to discover more.'
Spec posture evidence
Inferred effective spec: 2025-06-18+.
Relies on Dynamic Client Registration (deprecated) - use Client ID Metadata Documents (CIMD)
Tool descriptions are consistently brief (25-40 chars). Rubric baseline is 194 chars. Descriptions lack context about WHEN to use each tool, prerequisites, and what the tool returns. Examples: 'Initiate a new learning session for a learner' (44 chars), 'Retrieve pending alerts for the current learner' (47 chars), 'Check learner mastery status for a domain or concept' (50 chars).
No parameter descriptions are visible in source code for any tool. Rubric requires every input parameter to have a non-empty description. Without param descriptions, LLMs cannot infer parameter meaning from names alone. Example: a 'status' param could mean HTTP status, project status, or boolean state.
No output schemas documented. Rubric requires documentation of return types and structure for every tool. LLMs need to know what fields to expect to plan downstream tool calls and extract relevant data. This is invisible in provided source.
Multiple tools perform destructive or write-heavy operations (delete_domain, publish_curriculum_revision, update_learner_profile, learning_negotiation, set_domain_priority, mark_domain_high_stakes, update_learner_memory) with no visible confirmation or dry-run support. Agents make mistakes, irreversible operations need safeguards.
No error handling guidance visible in source. Rubric requires error responses to tell LLMs what to do next with actionable recovery steps (e.g., 'User not found. Try search_users() with a partial name.'). Raw error codes or stack traces are not useful to agents.
Tool naming includes some ambiguity. Tools like 'get_olm_snapshot', 'get_pedagogical_snapshots', 'get_decision_replay_summary' use domain jargon (OLM, pedagogical) without clarification in descriptions. 'read_raw_session' vs 'get_memory_state' are unclear about their distinction.
No evidence of permission gating or scope declarations in visible source. Tools like delete_domain, update_learner_profile, and publish_curriculum_revision are high-risk but show no permission checks. Rubric requires destructive tools to declare required permissions and verify authority before execution.
Add permission checks and audit logging to high-risk tools. Declare required scopes (e.g., 'admin:delete', 'write:learner_profile') and verify before execution. Log who called what with parameters and outcome for compliance.
Clarify ambiguous tool names and distinguish overlapping tools. Rename or expand descriptions for: get_olm_snapshot (explain OLM), get_pedagogical_snapshots (pedagogical for what purpose?), read_raw_session vs get_memory_state (what's the difference?). Provide dependency hints in descriptions.
Implement pagination for tools that return lists (get_pedagogical_snapshots, get_misconceptions, list_implementation_intentions, get_dashboard_state). Add limit and offset/cursor parameters with a cap (20-50 items). Return total_count or next_cursor.
Add input validation and format constraints. Use enums for fixed choices (kind in record_learning_event already does this well; extend to other tools). Add regex patterns, length limits, and numeric ranges to parameter descriptions.