MCP server bridging the Antigravity (agy), OpenAI Codex, GitHub Copilot, Cursor, opencode, Grok Build, and Kimi Code CLIs so Claude Code can drive them as sub-agents
The server defines 10 tools with consistent verb_noun naming (antigravity_ask, antigravity_continue, etc.) and non-empty descriptions. However, definition quality is undermined by several systematic issues: (1) Descriptions are moderately detailed but often exceed the recommended 200 chars (e.g., antigravity_ask at ~180 chars is borderline; codex_ask exceeds 200), and critically lack actionable context about when to use each tool vs. alternatives. (2) Input schemas are visible and properly typed (string, enum), but parameter descriptions are present yet generic, many lack guidance on valid formats, constraints, or dependencies. (3) No documented output schema for any tool, responses are described informally in text ('returns the answer') without structured schema documentation. (4) The 5 sandbox-accepting tools (codex_ask, codex_continue, copilot_ask, copilot_continue, cursor_ask, cursor_continue, grok_ask, grok_continue) all expose a 'sandbox' parameter with enum values ('read-only', 'workspace-write', 'danger-full-access') but lack clarity on implications of each mode or when to use each. (5) Error handling is not evident in the tool definitions, no guidance on how tools report failures or what an agent should do on errors. (6) No idempotency guarantees documented for the _continue tools, which retain state across calls. (7) The model parameter validation differs across tools (agy validates via 'agy models', codex has 'not validated', cursor validates via 'cursor-agent models') but this variance is not explained to users. Average tool description length is ~140 chars (within baseline range 34-392), but clarity and actionability are moderate. All parameters have descriptions, meeting that baseline, but quality varies significantly.
Run a prompt against the Antigravity (agy) CLI and return the answer. Runs `agy -p` headless with optional model selection.
Resume the last Antigravity conversation in a workspace. Continues the pinned session or falls back to the newest conversation for the cwd.
Run a prompt against OpenAI's Codex CLI and return the answer. Runs `codex exec` headless with optional sandbox and model selection.
Resume the last Codex session in a workspace. Continues the pinned session or falls back to the newest rollout for the cwd.
Run a prompt against GitHub Copilot CLI and return the answer. Runs `copilot -p` headless with optional sandbox mode.
Resume the last GitHub Copilot session in a workspace. Continues the pinned session or falls back to the newest session for the cwd.
No documented output schemas for any of the 10 tools. Responses described only as informal strings ('return the answer') without structured field definitions. LLMs cannot parse responses or plan downstream tool chains without knowing output shape.
Sandbox parameter (present in 8 tools) has enum constraint but no guidance on implications or when to use each mode. 'danger-full-access' is particularly risky without explanation of what access is granted or when it's necessary.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 65 | 2026-07-28+ | v2 |
Run a prompt against Cursor CLI and return the answer. Runs `cursor-agent -p` headless with optional sandbox mode.
Resume the last Cursor chat in a workspace. Continues the pinned chat or falls back to the newest chat for the cwd.
Run a prompt against Grok Build CLI and return the answer. Runs `grok -p` headless with optional sandbox mode. EXPERIMENTAL — no authenticated round-trip verified yet.
Resume the last Grok session in a workspace. Continues the pinned session or uses grok's `-c` to resume the most recent session for the cwd.
Model parameter validation is inconsistent and undocumented. antigravity_ask validates against 'agy models', codex_ask explicitly states 'not validated', cursor_ask validates against 'cursor-agent models'. Agents cannot anticipate which tools validate input, risking invalid model slugs.
No error handling guidance. Tool descriptions do not explain how failures are reported, what error conditions are possible, or what agents should do on failure. Missing recovery guides for common cases (e.g., model not found, authentication failure, timeout).
Stateful _continue tools (antigravity_continue, codex_continue, copilot_continue, cursor_continue, grok_continue) retain session across calls but lack documented idempotency guarantees. Agents cannot know if retrying a failed _continue call will re-run the same prompt or append the new one.
Tool descriptions lack differentiation and actionable context. 'antigravity_ask' and 'antigravity_continue' differ only in continuation semantics, but descriptions don't explain when to choose one over the other or how they interact with workspace state.
Workspace parameter documented as 'Absolute workspace path, or the server's cwd when omitted' but no guidance on how cwd is determined, whether relative paths are accepted, or how missing/invalid paths are handled.
grok_ask marked as 'EXPERIMENTAL, no authenticated round-trip verified yet' but no guidance on whether agents should avoid it, what failures to expect, or when it may be stable. Experimental status should be clearer in metadata.