The server defines 20 tools with generally good naming (verb_noun pattern) and comprehensive descriptions. Most tools have well-structured input schemas using Zod with type definitions and descriptions. However, there are systematic gaps: (1) Output schemas are not documented, callers cannot see what fields to expect from responses; (2) Error handling guidance is minimal, tools do not describe what to do on failure; (3) Some parameter descriptions lack format/constraint details; (4) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite having tools with clear READ/WRITE/DESTRUCTIVE risk levels already labeled. The tool set is well-composed (minimal duplication, clear separation of concerns), and naming follows patterns. Most tools have non-trivial descriptions (>20 chars), which supports discoverability. However, lack of output schemas and error recovery guidance prevents this from reaching 75+.
Tools (20)
claim_taskwriteauthsource verified74/100
Claim a task for execution. Acquires a physical lock. Only "todo" tasks with met dependencies can be claimed. Tasks assigned to disconnected workers can also be claimed.
create_taskwriteauthsource verified78/100
Create a new task on the board. Main agent only.
create_workspacewriteauthsource verified75/100
Register a git repository as a managed workspace. Main agent only.
delete_taskdestructiveauthsource verified73/100
Delete a task. Cannot delete locked/in_progress tasks. Main agent only.
dispatch_taskwriteauthsource verified75/100
Dispatch a task to a specific worker agent. Claims the task and sends it to the worker. Main agent only.
force_releasewriteauthsource verified70/100
Force release a locked task. Use for unresponsive workers. Main agent only.
No output schemas documented for any tool. Callers cannot see what fields to expect from responses. LLMs must infer output structure from description text alone, increasing hallucination risk. Example: create_task returns a JSON object, but the fields (task_id, created_at, status, etc.) are undocumented. Per pattern:tool and pattern:response-shaper, all tools must document return types.
Add output schema documentation for every tool. Use JSDoc or Zod output types to define response fields. Example for create_task: {task_id: string, title: string, status: 'todo', created_at: ISO8601, created_by: string, ...}. This enables LLMs to extract and chain results.
Add error handling guidance to every tool description. Example: 'Throws task_not_found (404). Try list_tasks() to find the correct task_id.' For destructive tools like delete_task, add: 'Returns cannot_delete_in_progress (409) if task is locked. Try force_release() first, or ask the user to wait for the worker to finish.'
Implement tool annotations in MCP server registration. Add readOnlyHint, destructiveHint, and idempotentHint to the tool definition in the server.tool() calls. Example: server.tool('delete_task', ..., {hint: 'destructive'}). This informs LLM planning and safety policies.
Add explicit enum constraints to all parameters that have a fixed set of valid values. Example: update list_tasks's 'priority' param from string to z.enum(['critical', 'high', 'medium', 'low']). Replace generic descriptions like 'Filter by priority: critical, high, medium, low' with formal constraints.
Add pagination support to list tools. Add optional parameters: limit (default 20, max 100), offset or cursor. Update tool descriptions: 'Returns up to 20 tasks. Pass offset to retrieve the next batch. Include total_count in response for client-side pagination.' Example: server.tool('list_tasks', ..., {limit: z.number().min(1).max(100).optional(), offset: z.number().min(0).optional()}).
Score history
Overall score trend
↑ 6 points across a rubric change (v1 → v2)
66/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
66
2026-07-28+
v2
2026-03-09
C
60
-
v1
read onlyauthsource verified74/100
Get board overview: task counts by status, active agents, and recent events. Main agent only.
get_taskread onlysource verified75/100
Get detailed information about a specific task, including history, comments, and progress logs.
heartbeatread onlyauthsource verified73/100
Send heartbeat signal. For main agents, also returns pending events.
list_tasksread onlysource verified78/100
List tasks with optional filters.
list_workspacesread onlyauthsource verified74/100
List all managed workspaces with optional filters. Main agent only.
poll_eventsread onlyauthsource verified74/100
Poll for recent events. Main agent should call this periodically to monitor workers.
register_agentwritesource verified80/100
Register a new agent. role=main enforces uniqueness per project (only one active main per project allowed).
release_taskreversibleauthsource verified73/100
Release a claimed task back to "todo". Use when you cannot complete the task.
report_progresswriteauthsource verified76/100
Report progress on a claimed task. Also refreshes the lock expiry timer.
review_taskwriteauthsource verified76/100
Review a task in "review" status. Approve → done, Reject → todo. Main agent only.
set_dependencywriteauthsource verified73/100
Set task dependencies. Validates no circular dependencies. Main agent only.
sync_with_basewriteauthsource verified73/100
Sync your task worktree with the latest base branch changes (rebase). Use when the base branch has moved ahead.
Error handling guidance is absent. Tools do not describe what LLM should do on failure (e.g., 'If task not found, try list_tasks() to find the correct task_id'). Per pattern:recovery-guide and pattern:error-classification, every tool should guide recovery. Examples: delete_task says 'Cannot delete locked/in_progress tasks' but does not say how to unlock or what error to expect; dispatch_task does not explain what happens if agent_id is invalid.
Tool annotations missing. Tools are labeled with Risk levels (READ, WRITE, DESTRUCTIVE) but MCP schema does not expose these via readOnlyHint, destructiveHint, or idempotentHint. LLMs cannot read risk labels from the file structure, they must be in the MCP tool registration. Constrains agent planning and safety.
Parameter descriptions lack format/constraint specifics for some tools. Example: list_tasks has 'priority' param with description 'Filter by priority: critical, high, medium, low' but no enum constraint in the schema. create_task has 'priority' with enum (good), but some optional params like 'labels' and 'depends_on' have generic descriptions that do not specify format (comma-separated string vs array, or what task ID format is valid). Per pattern:constrained-input, all enums should be explicit in schema, and all string parameters should document format/length/pattern.
No pagination guidance for list tools. list_tasks and list_workspaces have no limit, offset, or page parameters documented. Per pattern:paginated-result and mxe:enforce-result-limits, list tools must support pagination and cap results at a reasonable limit (e.g., 20-50). Without pagination, large result sets blow context window and degrade LLM reasoning.
Idempotency not documented. Tasks like create_task, claim_task, and update_status do not state whether they are idempotent. Per pattern:idempotent-operation, agents retry on ambiguous failures, non-idempotent tools risk duplicate side effects (e.g., claiming the same task twice). If retrying create_task with same inputs should be safe, that must be explicit.
Some descriptions are vague or assume context. Example: set_dependency says 'Validates no circular dependencies' but does not explain what happens if a circular dependency is detected (error? rejection? which status is attempted?). force_release says 'Use for unresponsive workers' but does not explain the operational flow (does worker get notified? does task move to todo?). Per pattern:tool-description, descriptions must answer WHAT, WHEN, and WHAT IF.
Workspace mode parameter unclear. register_agent's 'workspace_mode' param has description 'Workspace mode: "required" (task-based agents) or "disabled" (TUI agents). Defaults to "disabled".'. The distinction between 'task-based' and 'TUI agents' is not self-evident to an LLM. What does each enable? When should an agent choose one over the other? Per pattern:tool-description, describe the consequence of each enum value in user-facing terms.
Chaining field availability unclear. After dispatch_task, will the response include the assigned agent's active status, or must the LLM make a separate call to poll_events or heartbeat? After create_task, does the response include task_id? Per mxe:include-chaining-ids, responses must include all IDs and references that downstream tools will need to avoid forcing extra discovery calls.
Document idempotency explicitly. For idempotent tools like create_task with a unique identifier, add: 'Idempotent: calling with the same title/description returns the existing task if found, does not create duplicates.' For non-idempotent tools, add: 'Not idempotent: each call dispatches the task. Retry only if the first call failed visibly (no response received); if the task was already claimed, retry will fail with task_locked.'
Clarify workspace_mode in register_agent. Expand description: 'Workspace mode controls whether the agent can operate on git repositories. Set to "required" for agents that need to execute tasks in workspace worktrees (e.g., code review agents). Set to "disabled" for TUI-only agents that interact only with the kanban board (e.g., orchestrators that do not execute code).'
Include chaining references in responses. Document that create_task response includes task_id (for update_task, dispatch_task), create_workspace response includes workspace_id (for sync_with_base), and claim_task response includes lock_token (for update_status, report_progress, release_task). Add this to tool descriptions: 'Returns task_id for use in follow-up calls like dispatch_task or update_task.'
Add format constraints for string parameters. Example: agent_token, task_id, agent_id, lock_token, specify if these are UUIDs, alphanumeric strings, or opaque tokens. Example: 'agent_token: opaque string returned by register_agent. Do not attempt to parse or modify.'
Document circular dependency behavior in set_dependency. Example: 'If setting depends_on would create a cycle, returns circular_dependency (422) with the problematic path. Example: task A depends on B, B depends on C, you try to make C depend on A. Resolve by removing one edge in the cycle, then retry.'
Expand force_release description. Example: 'Force-releases a locked task, moving it back to "todo" status. Use when a worker agent has crashed or become unresponsive (no heartbeat for >5 minutes). The worker will lose its lock_token and cannot further update this task. The task is available for re-claim by another agent. Worker is notified via poll_events() of the force release.'
Add natural-language identifier support where applicable. Example: dispatch_task currently requires agent_id. Consider supporting agent_name as an alias, documented as: 'agent_id: worker agent ID or display name. If a name is passed, it is resolved to the agent ID; if multiple agents have the same name, returns agent_ambiguous (409) with a list of matching IDs.'
Document the data model for dependencies and task lifecycle. Add to create_task and set_dependency descriptions: 'Dependent tasks cannot transition to in_progress until all tasks in depends_on have status done. Attempting to claim a task with unmet dependencies returns dependency_not_met (409) with a list of blocking task IDs.'
For poll_events, document the event structure. Example: 'Returns array of events like {type: "task_claimed", task_id, agent_id, timestamp} or {type: "task_completed", task_id, status, timestamp}. Available event types: task_created, task_claimed, task_completed, task_failed, task_reviewed, agent_registered, agent_disconnected.'
Add confirm/dry-run pattern to destructive tools. Example: delete_task could accept optional dry_run: boolean parameter. 'If dry_run=true, validates that the task can be deleted (not locked) and returns what would happen without executing. Default false. Recommended: always call with dry_run=true first when invoked by agents, to prevent accidental deletes.'
Document the session_id reuse pattern in dispatch_task and register_agent. Example: 'session_id: opaque identifier to group related MCP sessions. If provided, the worker agent should reuse this session ID when registering, enabling the main agent to track and reconnect to the same worker context. Useful for persistent agent contexts.'
Specify timeout behavior for long-running operations. Example: 'sync_with_base may take 5-30 seconds depending on repository size and rebase complexity. Client should set HTTP timeout to at least 60 seconds. If timeout occurs, the rebase may be partially applied, call get_task() to check current status and retry if needed.'
Add batch variants for tools called in loops. Example: dispatch_task dispatches one task to one agent. If the LLM needs to dispatch multiple tasks, it will call dispatch_task repeatedly. Consider adding dispatch_tasks (plural) accepting an array of {task_id, agent_id, prompt?} objects, returning per-item results. This reduces round-trips and tokens.