Standalone MCP server for Warp Oz agents — no VS Code required. Works with Claude Code, Cursor, and any MCP client.
OzBridge exposes 6 tools with generally clear intent and reasonable descriptions. Naming follows verb_noun conventions (oz_agent_run, oz_run_get, oz_list_models, etc.), making intent discoverable. Most tools have descriptions in the 100-300 character range, which is solid. However, there are significant gaps: (1) Input schemas are NOT visible in the provided source, only parameter names and types are inferred from package.json/README examples; (2) Output schemas are completely undocumented, no tool declares what fields it returns or how results are structured; (3) Parameter descriptions, while present in the extension manifest, are brief and sometimes lack critical constraints (e.g., no validation rules for model IDs, no guidance on skill IDs beyond examples); (4) No error handling guidance, tools do not explain how to handle failures, retries, or partial results. Two tools (oz_agent_run, oz_agent_run_cloud) properly flag their risk/side-effect profile (WRITE, credit consumption), which is excellent. However, the lack of visible input schemas in the MCP server code (vs. the VS Code extension manifest) suggests the MCP transport layer may not be fully specifying these constraints. The tool set is well-composed (each does one thing) and idempotent where needed (oz_run_get, oz_run_list are read-only and safe to retry). The main weakness is documentation completeness for downstream tool chaining, agents cannot easily infer what fields to extract from one tool's output to pass to another.
Run a Warp Oz AI agent locally in the host workspace and return its full output synchronously. Runs `oz agent run` with your prompt and blocks until the agent finishes. Use for local coding tasks — refactor, write/run tests, debug, explain code — that should NOT consume cloud credits; for cloud execution call `oz_agent_run_cloud` instead. NOT read-only: the agent may create or modify files in the workspace. Requires the `oz` CLI on PATH (install Warp).
Launch a Warp Oz AI agent in Warp's cloud (not on the local machine). ⚠️ CONSUMES WARP CREDITS — confirm with the user before calling. Returns the run id immediately WITHOUT waiting for completion; poll the terminal status and output with `oz_run_get`, or find the run later via `oz_run_list`. Requires the `oz` CLI on PATH and a Warp account with cloud credits.
List the AI model ids available to the connected Warp Oz account (from `oz model list`) and report the current default. Read-only; takes no arguments. Call this first to discover valid ids before passing `model` to `oz_agent_run` / `oz_agent_run_cloud`, or before `oz_set_default_model`.
Fetch the current status and output of a single Warp Oz run by its id. Read-only and idempotent — safe to call repeatedly while polling a cloud run to completion (SUCCEEDED / FAILED). Typically called after `oz_agent_run_cloud`, which returns the run id; ids also come from `oz_run_list`.
Output schemas completely undocumented. No tool declares what fields are returned (e.g., what fields does oz_run_get return? what is the structure of oz_run_list results?). This forces LLMs to guess or make exploratory calls, breaking downstream tool chaining.
Input schemas not visible in source code. Parameter definitions are inferred from the VS Code extension manifest (package.json), not the MCP server implementation (src/mcp/tools.ts). Without explicit JSON Schema in the MCP layer, LLMs may not enforce type/enum constraints and the server's type system is invisible to MCP clients.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
List recent Warp Oz runs (id, status, timing), newest first. Read-only. Use to discover run ids to pass to `oz_run_get`, or to review recent agent activity.
Set the default Oz model for every OzBridge surface by writing `defaultModel` into the workspace `.warp/warp-bridge.yaml` (the highest-precedence config source). Persistent side effect: edits that file on disk. The id is validated against `oz model list` when reachable. Requires a workspace root (an extension workspace, or `--cwd` for the standalone server).
No error recovery guidance. Tools do not explain what to do on failure. E.g., oz_agent_run says 'NOT read-only' but never describes failure modes: What if the oz CLI is not found? What if the agent times out? What if the workspace is invalid? No guidance for agent error handling.
Parameter descriptions lack validation constraints. E.g., 'model' accepts 'auto' or a model ID from oz_list_models, but no description states the format/pattern for model IDs (e.g., are they alphanumeric? case-sensitive? length limits?). 'skill' parameter example is '5-test-agent' but no description specifies if all skills follow this format or if custom names are allowed.
Pagination incomplete. oz_run_list accepts a 'limit' parameter but no description of 'offset' or 'cursor' for iterating results. Does it return a next_cursor or total_count? Without pagination metadata, agents cannot safely fetch large result sets.
Ambiguous status enum values in oz_run_list. The status parameter accepts both short forms ('all', 'active', 'completed') and full status codes ('QUEUED', 'INPROGRESS', 'SUCCEEDED', 'FAILED', 'CANCELLED', 'PAUSED', 'SKIPPED', 'UNKNOWN'). No description clarifies which to use when or if the full codes are returned by oz_run_get. LLMs may pass invalid combinations.
Tool chaining IDs not guaranteed. oz_agent_run_cloud returns a 'run id immediately WITHOUT waiting for completion' but it's unclear what other fields are returned (e.g., workspace path, agent name, etc.). oz_run_get expects a runId but the description does not state where this ID comes from or what format it takes. Agents need explicit guidance on which field to extract and pass downstream.
No idempotency guarantees for write operations. oz_agent_run and oz_agent_run_cloud do not declare whether they are idempotent (i.e., will retry with the same inputs always produce the same result?). For a local run (oz_agent_run) that modifies files, re-running with identical inputs could produce different results (files already exist, agent state changed). Agents need to know whether they can safely retry.