Unified CLI wrapper for AI agents (Claude, Gemini, Codex, Cursor, Antigravity, ACP, Pi) with multi-agent consensus, comparison, and council features. Implements MCP server protocol for agent task management.
CAG's tool definitions exhibit significant quality gaps across naming, descriptions, and schemas. While 8 tools are explicitly registered with basic descriptions and parameter schemas, the implementation has critical issues: (1) Tool naming lacks clear verb-noun patterns, 'cag_agent_task_tool' is a multi-action dispatcher that violates single-responsibility principle; (2) Descriptions are generic and lack actionable context for LLM selection, e.g., 'Create a background or foreground task' does not explain WHEN to use foreground vs background or what the implications are; (3) Parameter schemas are present but incomplete, missing enums for 'mode' (foreground/background) and 'agent_name', no range/format constraints on strings; (4) Output schemas are not documented anywhere in the provided source, LLMs cannot infer what fields task creation returns; (5) Error handling is not evident in the source, no recovery guidance, categorization, or actionable error messages visible. The tool definitions are skeletal and would force LLMs to guess at behavior.
Create a background or foreground task for executing a CAG agent request with specified prompt, model, system prompt, and optional working directory
Cancel a running CAG agent task
Get details of a specific CAG agent task by task ID
List all visible (non-expired) CAG agent tasks with their current status
Get the result of a completed or failed CAG agent task
Multi-action tool for managing CAG agent tasks: list, get, result, wait, wait_any, cancel. Action is specified via 'action' parameter.
Wait for a specific CAG agent task to reach a terminal state (completed or failed)
Wait for any of the specified CAG agent tasks to reach a terminal state
Tool 8 (cag_agent_task_tool) violates single-responsibility principle, combines 6 distinct operations (list, get, result, wait, wait_any, cancel) behind an 'action' parameter. This forces LLMs to reason about which action to invoke and creates an ambiguous parameter schema (action is a string with no enum, task_id/task_ids are both optional, unclear interdependencies).
All 8 tools lack documented output schemas. LLMs cannot infer what fields are returned by create_task, list, get, result, wait, or wait_any. This blocks downstream planning and tool chaining.
Parameter 'agent_name' in cag_agent_create_task lists 7 valid values in description ('claude, gemini, codex, cursor, antigravity, acp, pi') but does not declare them as a formal enum. This invites hallucinated agent names outside the valid set.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 40 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter 'mode' in cag_agent_create_task lists valid values ('foreground' or 'background') in description but lacks formal enum declaration. Description does not explain WHEN to use each mode or what the implications are (e.g., foreground blocks until completion; background returns immediately?).
Descriptions are extremely brief (5 - 10 words for most tools), offering no context on WHEN to use each tool vs similar tools, no recovery guidance, and no hint about return values. Baseline for A+ tools is 50 - 200 chars per description; these are 30 - 70 chars.
Parameter 'resume' in cag_agent_create_task has vague description: 'Optional session ID to resume from previous execution.' No indication of format, how to obtain a valid session ID, or what happens if an invalid one is passed.
cag_agent_task_list provides no pagination mechanism (no limit, offset, or cursor parameters). Description says 'List all visible (non-expired) tasks' but does not explain how many tasks are returned, what happens with thousands of tasks, or how to handle large result sets.
cag_agent_task_cancel is marked as WRITE (destructive) but its description does not explicitly warn that cancellation is irreversible. No recovery guidance is provided. Per pattern:command-tool, destructive operations should state consequences clearly.
No error handling examples or recovery guidance visible in source. Tools do not indicate which errors are retryable vs user-fixable. Per pattern:recovery-guide, error responses must tell the LLM what to do next.