MCP server that wraps AI CLI tools and proxies Chrome DevTools MCP to remote browser leases
This MCP server exposes 16 tools with significant quality gaps across naming, descriptions, and schemas. Most tools have basic descriptions (30-80 chars) but lack detailed parameter documentation, proper error handling guidance, and comprehensive output schema definitions. Tool names follow verb_noun conventions but many are compound operations (codex-reply, codex-start, codex-status) that could be clearer. Parameter descriptions exist but are often minimal (5-20 chars: 'ID of the job', 'Offset in the result'). Critical issue: No visible input schema validation or output schema documentation in the source code, schemas are inferred from the provided tool spec but not verifiable in server.js implementation. Security concern: The server wraps external CLI tools (Codex, Gemini, Claude) and passes prompts directly without visible sanitization or injection prevention.
Run Claude Code interpreter
Cancel a running Claude job
Get the result of a Claude job
Start a Claude job asynchronously
Get the status of a Claude job
Run a prompt in a Codex interactive workspace
Cancel a running Codex job
Missing output schemas for all 16 tools. No response structure documentation visible in source code. LLMs cannot infer what fields will be returned or plan downstream tool calls.
Parameter descriptions are minimal (5 - 20 characters on average). Most parameters lack validation rules, ranges, or format specifications. E.g., 'offset' in codex-result and codex-commentary lacks pagination semantics; 'model' in codex lacks enum of valid values; 'cwd' in codex lacks path validation guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 17 | - | v1 |
Get commentary from a Codex job
Peek at a Codex thread's in-flight status
Reply to a previous Codex thread
Start a Codex reply job asynchronously
Get the result of a Codex job
Start a Codex job asynchronously
Get the status of a Codex job
Run Google Gemini
Execute browser commands on a remote Chrome DevTools instance
Tool naming includes compound/hyphenated verbs (codex-reply, codex-start, codex-reply-start, remote-browser) that are less explicit than verb_noun patterns. Hyphenation makes parsing harder for LLMs and differs from the standard convention (get_*, list_*, create_*, etc.).
Insufficient tool descriptions. Most descriptions are 30 - 80 characters and lack WHAT/WHEN/WHY context. Examples: 'Run a prompt in a Codex interactive workspace' (44 chars) does not explain when to use codex vs. codex-start; 'Run Google Gemini' (18 chars, under 20-char minimum) offers no context; 'Execute browser commands on a remote Chrome DevTools instance' does not document allowed commands or error recovery.
No visible error handling guidance. Tools that invoke external services (Codex, Gemini, Claude, Chrome DevTools) lack documented error messages, retry strategies, or recovery instructions. Source code shows timeout constants (DEFAULT_CODEX_TIMEOUT_MS, DEFAULT_CLAUDE_TIMEOUT_MS) but no tool description explains what happens on timeout or how to recover.
Async job tools (codex-start, codex-reply-start, claude_job_start) lack clear documentation of job lifecycle. No description explains: How long does a job persist? What happens if you never call *_status or *_result? Is there a cleanup mechanism? How do cursor and wait_ms interact in *_status?
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in source. Tools marked with Risk=WRITE or Risk=REVERSIBLE lack semantic hints that would help agents reason about safety. Pattern: tool-annotations is not adopted.
Tool names 'gemini' and 'codex' are bare nouns, not action verbs. Convention is verb_noun (e.g., 'run_gemini', 'run_codex_interactive'). Bare nouns do not clearly signal to LLMs what action will be taken.
Parameter 'sandbox' in codex and codex-start is an enum (read-only, workspace-write, danger-full-access) but the description does not explain the security implications or when to use each mode. An LLM might default to 'danger-full-access' without understanding the risk.
Tool chaining is not documented. If codex returns a jobId, which tool does the agent call next? Are there dependency hints in descriptions? The tool set includes multiple codex tools (codex, codex-reply, codex-start, codex-reply-start, codex-status, codex-commentary, codex-result, codex-cancel, codex-peek) with overlapping semantics and no clear guidance on the intended call sequence.