A safe Streamable HTTP MCP bridge from ChatGPT/GPTs to local Codex.
Strong foundation with clear naming, comprehensive descriptions, and proper schema definitions. All 6 tools have verb-noun names (bridge_status, codex_read, codex_run, codex_reply, codex_job_status, ask_chatgpt) and detailed descriptions (100-200 chars). Input schemas use Zod with type constraints and descriptions. Tool annotations present (readOnlyHint, destructiveHint, idempotentHint, openWorldHint). Key gaps: output schemas not documented in code; error handling guidance minimal; no pagination/limit enforcement visible for list operations; ask_chatgpt lacks enum constraints for model parameter.
Ask an OpenAI API model a prompt and return the text response. This is not the ChatGPT web UI; it uses the official Responses API.
Read-only status check. Returns bridge safety policy, allowed local roots, tracked session count, and upstream Codex MCP tool availability. Does not start Codex and does not read project files.
Check a long-running codex_read, codex_run, or codex_reply job that previously returned a jobId. Poll this tool until status is completed or failed.
Run a read-only local Codex inspection in an allowed working directory. The bridge forces Codex read-only and does not permit file modifications. If Codex does not finish before the bridge fast-return deadline, this tool returns a jobId; call codex_job_status with that jobId until it completes.
Continue a Codex session that was first created through this bridge. If Codex does not finish before the bridge fast-return deadline, this tool returns a jobId; call codex_job_status with that jobId until it completes.
Output schemas not documented. Tools return ToolResult but LLMs cannot see what fields to expect. Breaks downstream tool chaining and forces agents to guess response structure.
ask_chatgpt 'model' parameter lacks enum constraint. Accepts free-form string, inviting hallucinated model names. Should enumerate valid options (gpt-4, gpt-4-turbo, gpt-5.5, etc.).
Error handling lacks recovery guidance. Tools return jobId on timeout but no explicit guidance on polling strategy, backoff, or failure modes. Agents cannot self-correct.
codex_job_status polling lacks documented timeout/max-retries. Agents could loop indefinitely on stuck jobs. Should document expected polling interval and failure detection.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 82 | 2025-06-18+ | v2 |
Start a local Codex session in an allowed working directory using the bridge sandbox policy. Prefer codex_read for read-only project inspections. If Codex does not finish before the bridge fast-return deadline, this tool returns a jobId; call codex_job_status with that jobId until it completes.
ask_chatgpt 'reasoningEffort' parameter description vague ('Optional reasoning effort for models that support it'). Should enumerate valid values and which models support it.