Multi-skill agent orchestration platform with support for collaborating with Claude, Antigravity, Codex, and Grok CLIs, plus issue-driven workflow automation
Agent Designer provides 7 tools with mixed quality. Four bridge tools (agy_bridge, claude_bridge, codex_bridge, grok_bridge) have well-structured schemas with detailed parameter descriptions (8-12 params each, all typed). Three workflow tools (create_plan, create_issues, list_plans) have simpler but adequate schemas. However, tool descriptions are present but relatively brief (50-150 chars), lacking the depth and context LLMs need for confident tool selection. No output schemas are documented, callers cannot know what fields to expect from responses, violating the output-schema requirement. Error handling exists in code but is not surfaced in tool metadata. One critical issue: tool names use underscores (agy_bridge) rather than clear verb_noun patterns that convey action (run_agy, invoke_agy would be clearer). The server supports session persistence via SESSION_ID parameters across multiple tools, which is good for multi-turn continuity, but this is not explicitly documented in tool descriptions as a feature.
Wraps the Antigravity CLI (`agy --print`) to provide a JSON interface, live stderr progress, and multi-turn continuity via SESSION_ID
Wraps the Claude Code CLI (`claude --print`) to provide a JSON interface, live stderr progress, multi-turn sessions via SESSION_ID, and structured result telemetry (termination reason, cost, tokens, turns)
Wraps Codex CLI exec mode to provide a JSON interface and multi-turn continuity via SESSION_ID
Create the Issue CSV paired with a plan file. Derives issues/<timestamp>-<slug>.csv from the plan filename and validates row content before writing
Create a repo-local plan markdown file under ./plans with frontmatter and body
Wraps the Grok CLI (`grok -p`, headless mode) to provide a JSON interface, live stderr progress, and multi-turn continuity via SESSION_ID
List repo-local plan summaries by reading frontmatter only, with optional filtering and JSON output
No output schemas documented for any tool. Callers cannot know what fields to expect from responses, forcing LLMs to guess field names for downstream tool chaining and result extraction.
Tool names use underscore separators (agy_bridge, claude_bridge) without clear action verbs. Names like 'run_agy_prompt', 'invoke_claude', or 'execute_codex' would better convey the action and help LLMs infer intent from the name alone.
Tool descriptions are brief (50-150 chars) and lack context about WHEN to use each tool and WHY one might be chosen over another. E.g., agy_bridge describes wrapping Antigravity CLI but does not explain use cases vs claude_bridge or codex_bridge.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-04-07 | F | 26 | - | v1 |
SESSION_ID parameter for multi-turn session resumption is present across bridge tools but not documented as a composition feature. LLMs cannot infer that SESSION_ID enables conversational continuity without explicit description.
Parameter descriptions present, but format constraints are inconsistent. E.g., 'print_timeout' says 'Go format (e.g. 5m, 90s)' but this is example-based rather than a formal pattern or enum. Enums exist for sandbox mode (codex_bridge) and permission_mode (grok_bridge) but not all constrained parameters use this pattern.