Agent coordination and step management server for distributed task execution with dependency tracking, heartbeat monitoring, and step recovery
The server provides 15 tools with explicit tool registration and clear verb-noun naming conventions. All tools have descriptions and documented parameters with types. However, there are systemic issues that significantly limit the quality: (1) Parameter descriptions are often minimal (5-15 chars) and lack actionable constraint guidance. (2) Output schemas are completely undocumented, no tool definition includes what fields are returned or their types. (3) Many parameters accept free-form strings (steps, files_modified, commit_hash) that should be constrained or formatted. (4) Error handling guidance is absent, no tools document recovery paths or categorize errors as retryable vs. fatal. (5) Complex parameters like 'steps' expect JSON strings but lack parsing examples or validation rules. The tool design follows a coherent agent coordination domain (projects, steps, claims, recovery), but falls short of production-grade agentic tool patterns. Most tools would score 45-60 individually.
Atomically claim the next available step
Mark step as completed
Create a new project with steps. Steps parameter should be a JSON array of objects with step_num, branch, scope, and depends_on fields.
Find steps with stale heartbeats (crashed agents)
Mark step as failed
Get agent activity log
Get all steps available to work on (no incomplete dependencies)
No documented output schemas for any tool. Tools define inputs but return types are completely unspecified, LLMs cannot plan downstream tool chains or extract required fields. Critical for composition.
Parameter descriptions are minimal and lack constraints. Example: 'steps' parameter says 'JSON string containing array of step definitions' but provides no parsing format, example, or validation rules. 'commit_hash' offers no guidance on format (40-char hex? full or abbreviated?). 'files_modified' accepts JSON string with no schema. LLMs will guess at correct format.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Get project and agent metrics
Get project details and all steps
Get step details
Update step heartbeat (call every 30-60 seconds while working)
List all projects
Recover a failed step by archiving it to history and resetting to not_started
Recover a stale step (reset to not_started)
Mark a claimed step as in progress
No error handling documentation. Tools like 'claim_step' (atomic operation) and 'complete_step' (state mutation) do not document failure modes, retryability, or recovery paths. An LLM has no guidance if a step claim fails or a heartbeat times out.
Missing pagination and result limiting guidance. 'list_projects' and 'get_agent_events' could return unbounded result sets. No documentation on max result count, pagination strategy (offset/cursor), or recommended limits. Large result sets will blow context windows.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite tools having clear risk profiles. WRITE and REVERSIBLE operations are declared in metadata but not in the tool schema itself, which limits LLM planning and safety checks.
'status' parameter in 'list_projects' lacks enum constraint. Description says 'Filter by status (active, completed, aborted)' but parameter is a free-form string. LLM may invent status values like 'pending', 'failed', or 'closed'.
'timeout_minutes' in 'detect_stale_work' lacks bounds. No minimum/maximum specified. LLM could pass -1, 0, or 999999, causing nonsensical queries or backend errors.
'limit' parameter in 'get_agent_events' has no bounds. Could default to 100 but lacks maximum constraint documentation. No guidance on recommended limits or pagination strategy.