One brain, many arms — spawn multiple specialized Codex agents as MCP servers
Codex Octopus has three tools with clear, detailed descriptions and well-structured schemas. The naming follows verb_noun conventions (codex, codex_reply, create_codex_mcp). Descriptions are detailed (150-500+ chars) and explain WHAT, WHEN, and WHY. Input schemas are comprehensive with enums, optional parameters, and good type coverage. However, there are meaningful gaps: (1) no explicit output schemas documented in the source code, while the tools likely return structured data, the response format is not specified anywhere in the visible code; (2) no error handling guidance, tools don't show what errors are possible or how to recover; (3) parameter descriptions lack some detail on constraints and formats (e.g., what constitutes a valid thread_id format, what happens if a model override is invalid); (4) the factory tool (create_codex_mcp) has a description with example triggers that reads more like a prompt injection list than a tool definition. These gaps are not critical but prevent the server from reaching 'excellent' territory.
Send a task to an autonomous Codex agent. It reads/writes files, runs shell commands, searches codebases, and handles complex software engineering tasks end-to-end. Returns the result text plus a thread_id for follow-ups via codex_reply.
Continue a previous codex conversation by thread ID. Use this for follow-up questions, iterative refinement, or multi-step workflows that build on prior context.
Generate a .mcp.json config entry for a new Codex Octopus MCP server instance. WHEN TO USE: user says 'codex octopus agent', 'codex octopus mcp', 'new codex octopus', 'add codex octopus', 'create codex octopus', 'codex octopus instance', 'codex octopus config', 'codex octopus server', or any phrase combining 'codex octopus' with agent/mcp/new/add/create/config/server/setup. This is a wizard: only a description is required. Returns a ready-to-use .mcp.json config and lists all customization options. Call again with more parameters to refine.
No output schemas documented for any tool. While input schemas are detailed, response structures are not specified in the source code. LLMs cannot predict what fields to expect or plan downstream tool calls.
No error handling guidance. Tools do not document what errors are possible (e.g., invalid API key, network failure, model not found) or how the LLM should recover. Per the rubric, error responses should tell the LLM 'what to do next'.
Parameter constraints lack specificity. For example, 'model' parameter accepts any string, no guidance on valid model names or consequences of invalid overrides. 'effort' enum is good, but no explanation of what each level means (e.g., 'xhigh = deepest reasoning, ~10x latency'). Thread_id format is undocumented.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
create_codex_mcp description includes trigger phrases ('when to use: user says...') that read like a prompt injection list. This is unusual for a tool definition and may confuse LLMs. The description should focus on WHAT the tool does, not on prompt patterns that trigger it.
No documentation of idempotency. The 'codex' and 'codex_reply' tools perform complex operations (running shell commands, reading/writing files). Per the rubric, destructive tools should support dry-run or confirmation. Current descriptions do not indicate whether operations are safe to retry or if they have irreversible side effects.