MCP server for OpenAI Codex CLI integration
This server has a moderate foundation with some positive elements (explicit Zod schemas, tool descriptions present) but significant gaps in LLM-optimized design, parameter guidance, and error handling. Tool names lack consistency in verb-first patterns. Several descriptions are present but generic. Schema coverage is partial, while input parameters are typed, output schemas are undocumented. Error messages in code show some recovery guidance, but not consistently applied. The 'ask-codex' and 'exec-codex' tools expose complex configuration parameters without sufficient constraint documentation or defaults. The 'ping' and 'version' tools are included but not typically essential. Overall, this reads as a technically functional but not agent-optimized toolkit.
Apply the latest diff produced by Codex agent to the local git repository
Execute OpenAI Codex with comprehensive parameter support for code analysis, generation, and assistance
Non-interactive Codex execution for automation and scripting
Get information about available Codex MCP tools and usage
Test MCP connection and server responsiveness
Get version information for Codex CLI and MCP server
Tool names lack consistent verb-first pattern. 'ask-codex' and 'exec-codex' use domain jargon; they should be 'execute-codex' or 'query-model'. 'ping' and 'help' are generic utility names that don't scale with domain-specific tools.
Output schemas are completely undocumented. Tools return formatted strings (e.g., apply-diff returns markdown-formatted text) but agents cannot parse structured fields from responses. Downstream tools have no way to extract IDs, status, or metadata.
Parameters with constrained values (e.g., 'model', 'sandbox', 'approval') are documented as free-form strings in descriptions (e.g., 'Options: gpt-5') rather than enforced as enums in the schema. This invites invalid LLM inputs and increases error rates.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | - | v1 |
'ask-codex' and 'exec-codex' accept complex 'config' parameter as either string or object (z.union([z.string(), z.record(z.any())])) with no validation or description of valid keys. LLMs will struggle to construct valid config without examples or a schema.
Many parameters lack minimum/maximum constraints. 'timeout' is unbounded (could be 0 or 1000000). 'image' accepts array with no size limit. 'prompt' has min=1 but no max, inviting arbitrarily long strings.
'help' and 'version' tools have empty or near-empty input schemas ({}), making them low-utility placeholders. They do not contribute to agent capability and consume tool registration slots.
Error handling is scattered in tool implementations (try-catch blocks in execute()) but not consistently structured. apply-diff shows good recovery hints ('No Diff Available', 'Git Repository Required'), but ask-codex and exec-codex do not return recovery guidance on failure.