MCP server for HITL Kit. Exposes the 15 HITL primitive event kinds as MCP tools so Claude Code, Cursor, Claude Desktop, and any MCP-aware client can emit schema-validated human-in-the-loop events.
HITL Kit MCP exposes 16 UI-oriented tools with well-structured schemas and clear descriptions. All tools have input schemas with typed parameters and descriptions. However, naming conventions deviate from standard verb_noun patterns (e.g., 'hitl_interrupt_card' vs 'show_interrupt_card'), and descriptions lack actionable guidance on WHEN to use each tool or what downstream actions follow. Output schemas are not documented. Error handling and recovery guidance are absent. Tools are READ_ONLY UI primitives, not state-modifying operations, which limits composition patterns. No tool accepts natural-language identifiers or provides pagination. The server is STDIO-only, capping protocol readiness at 50.
Surface an ordinal scale indicating how much AI involvement was used for a given output (0 = fully human, 4 = fully AI).
Ask the human for a binary approve/reject decision on a specific item.
Present a batch of mixed agent actions to the human for sequential approve/reject.
Surface a single source-backed citation for a claim the agent is making. Use when the agent wants the human to verify the source supports the claim.
Show the human the removable context items (notes, files, URLs) attached to this agent run.
Show the human a proposed before/after diff for a text or code edit. Use whenever the agent wants to apply an in-place edit and the human should accept or reject before it lands.
Tool names do not follow verb_noun convention. All 16 tools use 'hitl_<noun>' pattern (e.g., 'hitl_interrupt_card') instead of action verbs like 'show_interrupt_card' or 'emit_interrupt_card'. LLMs infer intent from verb prefixes; this naming obscures what action each tool performs.
Descriptions lack WHEN-to-use guidance and downstream action hints. E.g., 'hitl_interrupt_card' says what it surfaces but not when an agent should call it vs. 'hitl_approve_reject', or what the agent does after the human responds. Descriptions average ~70 chars; production baseline is 194 chars with context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Surface a multi-step plan the human can edit, reorder, or delete steps from before the agent executes. Steps marked locked cannot be removed.
Point at WHERE a claim is grounded — a character span, a normalised bounding box, or a time segment in a named source — so the reviewer's attention lands on the disputed thing rather than the whole document. Record sources consulted but unused in notAssessed, so absence of evidence stays visible.
Surface an in-thread approval boundary for an agent action. Use for citations, quote verification, or any write step that needs explicit human approval before proceeding.
Emit a collapsible thought/action/result reasoning trace so the human can inspect how the agent arrived at a result.
Ask the human a multi-part question supporting single-choice, multi-select, and free-text follow-ups.
Surface a research-agent config panel so the human can inspect what search session is running (create / follow-up / read-URL).
Show the human a single ranked search result with metadata and relevance.
Report the execution status of a subagent task to the human (one of: idle, running, completed, error, skipped, cancelled).
Preview a tool call (name, args, optional rationale and signals) so the human can approve or reject before execution. Use for destructive or high-stakes tool calls.
Surface a draft-in-progress widget summarising what the writing agent is producing (title, target section, word range, evidence notes).
No output schemas documented. Tools are UI event emitters; responses are implicit (human approval/rejection). Without explicit output documentation, agents cannot plan multi-step flows or extract structured results for downstream tools.
No error handling or recovery guidance. Tools lack descriptions of failure modes (e.g., what if human rejects? what if timeout?). Agents have no guidance on retryability or next steps after tool invocation.
STDIO transport only. Server is not remotely accessible and cannot be used by hosted MCP clients. Limits deployment to local/embedded scenarios. No HTTP or SSE support.