MCP server: launch and steer Cursor Cloud Agents from Claude Code, Cursor Desktop, or any MCP client
Strong tool naming (verb_noun pattern), comprehensive descriptions (100-300 chars each), and well-structured schemas with typed parameters. All 9 tools have explicit descriptions and input schemas. However, output schemas are not formally documented, error handling lacks recovery guidance in some tools, and parameter descriptions could be more prescriptive about constraints and formats. Tool composition is excellent, each tool has a single responsibility and outputs chain well (agent_id/run_id returned for polling). Security is sound (no secrets in params). Descriptions are LLM-optimized and include prerequisites and when-to-call guidance.
Cancel the active run. Cancellation is terminal — to continue the conversation, send a follow-up (which starts a new run on the same agent). Does not archive the agent. Returns the run's actual post-cancel state.
Send a follow-up prompt to an agent. Starts a NEW run — the returned run_id is what you poll with cursor_status (the old run_id is stale). Only call when the previous run is terminal; a follow-up during an active run fails with "agent is busy" — poll cursor_status first.
Launch a Cursor cloud agent and enqueue its first run. Pass a complete task prompt: goal, constraints, and how to verify done — vague prompts produce vague agents. Repo-less launches (no repo_url) are for research, reviews, and writing; pass repo_url (a GitHub https URL) when the agent should write code. starting_ref is a branch or SHA; omit it to use the repo default (never assume "main"). Call cursor_models first and pass a model id verbatim — never guess ids, and never pass the literal string "Auto" (omit model instead). model_params is a list of {key, value} dicts for per-model options (only ids/params from cursor_models are accepted). mode is "agent" or "plan". Returns agent_id AND run_id: poll cursor_status with both. Launch can take minutes; if the tool reports status "unknown", the agent may still have been created — reconcile with cursor_list, or retry with the same idempotency_key (replays are safe and never duplicate).
List your agents, newest first. Use this to reconcile after a launch reported status "unknown" (match by name/creation time), or to find an agent you lost track of.
Output schemas not formally documented. Tool descriptions mention what is returned (e.g., 'Returns agent_id AND run_id') but JSON Schema output types are absent. LLMs cannot reliably parse response structures without formal schema.
Error handling lacks recovery guidance. cursor_launch returns {ok: False, error: str(exc)} but does not tell the LLM what to do next (e.g., 'Retry with idempotency_key' or 'Call cursor_list to reconcile'). Errors should guide the agent's next action.
Parameter constraints not fully specified in descriptions. 'timeout_s' is clamped to 5 min in code but description does not state the range (10 - 300 seconds). 'limit' in cursor_list has no min/max. LLMs cannot infer numeric bounds from code.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 74 | 2026-07-28+ | v2 |
List available Cursor models with their ids, parameters, and variants. Call this before cursor_launch and pass the id verbatim — ids are not guessable. Each model may support params (e.g. reasoning effort); only valid id/params combinations from this list will be accepted.
Wait for a run to finish (polling server-side) and return its result. Convenience wrapper over cursor_status for when you just want the final answer. Waits at most timeout_s seconds (clamped to 5 min) — on timeout, keep polling with cursor_status using the returned run_id. Large results are truncated with truncated=true. Transient API blips (429/5xx) are retried with backoff.
Check a run's status. Run status is the source of truth — the agent's lifecycle status (ACTIVE/IDLE) is NOT "is it still working". If run_id is omitted, the agent's latest run is used. Poll this every 10-30 seconds; never sleep inside a tool call. When terminal=true, result holds the agent's final reply.
Token usage and cost for an agent (or one run). Costs real money — check this when a run felt expensive.
Verify your API key works and see which identity it belongs to. Call this first if anything else fails with auth errors.
cursor_models output structure not described. Description says 'List available Cursor models with their ids, parameters, and variants' but does not specify the response schema (e.g., is it an array of {id, params, variants}? What fields does each model object contain?).
Pagination not implemented for cursor_list. Description mentions 'limit' parameter but no offset/cursor or total_count returned. If an agent has >10 agents, it cannot iterate through all of them.