Open Protocol for AI Agent Governance — webhook-driven GitHub App for AI agent orchestration with handoff FSM, engagement level governance, CODEOWNERS routing, and MCP tools
AgentCraftworks exposes 6 well-named tools with complete JSON schemas and non-empty descriptions. However, descriptions are generic (median ~80 chars) and lack LLM-optimization guidance. Parameter descriptions exist but are minimal. Output schemas are not documented, callers cannot see what fields to expect in responses. Error handling is basic (errorReporting flag present but no recovery guidance visible in code). Tools follow verb_noun naming convention and have clear single responsibilities. The server lacks tool annotations (readOnlyHint, destructiveHint, idempotentHint) which would help LLMs understand side effects. Overall, this is a functional but incomplete implementation, definitions are present and syntactically correct, but lack depth for optimal LLM agent reasoning.
Accept a handoff as the receiving agent and start working on it
Attach structured typed data to a handoff (e.g., security findings, code reviews, test results)
Mark a handoff as completed with results
Create a new agent handoff to delegate work to another agent
Retrieve structured context data attached to handoffs
Query the current state of handoffs and workflows
Output schemas not documented. Callers cannot see what fields create_handoff, accept_handoff, and complete_handoff return. This forces LLMs to guess at response structure and makes multi-step tool chaining error-prone.
Missing tool annotations. No destructiveHint on create_handoff, accept_handoff, complete_handoff, attach_context, tools that mutate state. No readOnlyHint on query_workflow_state and get_context. LLMs cannot distinguish safe discovery calls from irreversible writes without explicit annotations.
Parameter descriptions are minimal (10-20 chars average). E.g., 'Unique ID of the handoff to accept' is clear, but 'Name of the agent accepting the handoff' leaves ambiguous whether this is a display name, system identifier, or role. Descriptions under 50 chars risk LLM misinterpretation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 13 | - | v1 |
No recovery guidance in error paths. Code has errorReporting=true but no visible error classification (retryable vs user-fixable vs fatal) or next-step hints. When create_handoff fails (e.g., invalid agent name), LLM has no actionable recovery path.
Parameter 'to_agent' accepts free-form string with example '@code-reviewer'. LLMs may hallucinate agent names or pass the '@' symbol literally. Should enumerate valid agents or provide a discovery tool (list_available_agents).
Tool descriptions lack WHEN and WHY guidance. 'Create a new agent handoff' explains WHAT but not WHEN to use it (vs. direct invocation?) or WHEN query_workflow_state is required first. LLM-optimized descriptions should answer: What does it do? When should I call it? What comes next?