Workspace multiplexer for AI agents — run fleets of Claude Code, Codex & Gemini in parallel on Windows & macOS, with MCP browser automation built in
wmux presents a high-risk browser automation MCP with severe definition quality gaps. While tool names follow verb conventions (browser_repl, browser_replay, repl_run, repl_reset, repl_sessions, Write), the descriptions are dangerously verbose and lack structured guidance for LLM selection. The browser_repl tool description (514 words) reads as a specification dump rather than an agent-optimized prompt, it enumerates allowed browser methods inline instead of guiding the LLM on when to use the tool. Parameter schemas are present but descriptions are inconsistent. Critical security concerns: no mention of permission gates, no audit trail documentation, and sensitive operations (file write, browser automation) lack confirmation patterns. Error handling is minimal, no recovery guidance, no categorization of retryable vs fatal errors. The Write tool has a critical security flaw: it exposes file_path as a parameter with only weak protection ('no escape to parent directories'), trivial to exploit with directory traversal. No tool defines output schemas explicitly. STDIO transport caps protocol readiness at 50 maximum.
Write markdown memory files to the orchestrator's memory partitions (_global or workspace-scoped). Restricted to .md files only, no escape to parent directories or siblings.
Run a JavaScript snippet that drives the browser through many steps in ONE call. Each allowed browser_X tool is `await browser.X(args)` with the same args, resolving to {text, events} (+ refs:[{ref,param,role,name}] for snapshot/smart_snapshot, diff text, all refs; pass refs[i].ref as the arg named refs[i].param). A failed step throws (catchable). screenshot adds image:"img-N" (attached below). Allowed: browser_click, browser_wait, browser_hover, browser_drag, browser_select, browser_scroll_into_view, browser_highlight, browser_dialog, browser_navigate_back, browser_snapshot, browser_smart_snapshot, browser_diff, browser_type, browser_paste, browser_press, browser_navigate, browser_dismiss, browser_focus, browser_get_url. Args for the steps whose standalone tools are unlisted: navigate_back() hover(ref) drag(sourceRef,targetRef|path) select(ref,values) scroll_into_view(ref) highlight(ref) dialog(accept,text). Other browser calls are not allowed — the steps listed above cover the broad cases, and a general eval would need a sandbox wmux does not have. A step that times out leaves the page exactly where it stopped; assign the timeout (default 30s, max 5m) per call. Runs inside an automation lease: only one repl per surface at a time, queueing the rest. The lease holds the browser context open across steps, so variables assigned to globalThis persist (e.g. globalThis.rows = await browser.smart_snapshot(); then browser.click({ref: rows[0].ref})). Playwright frames are handled: a ref from inside a frame names that frame, and a click on it frames-aware — same for every step. Steps that typed into a password field are never stored in the action ring and make a flow unrunnable, so repl_run cannot play them back. A script that errors mid-run stops at that step; the page is left exactly where it stopped, so the cheapest recovery is to finish from here and save a new flow. If a script times out, the runtime is torn down and the next call starts fresh — variables are gone, but the browser context survives.
browser_repl description is 514-word specification dump, not LLM-optimized prompt. Enumerates 20+ browser methods inline instead of high-level guidance. Baseline: 50-200 chars optimal; this is 5x the recommended length and buries intent under implementation details.
Write tool exposes file_path parameter with inadequate path traversal protection. Description claims 'no escape to parent directories or siblings' but provides no actual validation, relies on description rather than schema constraints. LLMs can be prompt-injected to pass '../' paths.
No permission gates documented for any tool. browser_repl, browser_replay, repl_run, and Write all perform destructive/sensitive operations (browser automation, arbitrary code execution, file writes) with no scope declarations or permission checks visible in definitions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 44 | 2026-07-28+ | v2 |
Record and replay a browser flow. After a flow works, save it by name; a later run repeats it without reading a single snapshot, which is where the saving is. A run that cannot find an element stops at that step and reports why, leaving the page there for you to finish live. Steps that typed into a password field are never stored and make a flow unrunnable. Recording is a ring: the last 100 actions (or 1048576 bytes of them, whichever comes first) are still available to save, and one trace holds 500 steps — so save a long session in parts, naming each tail with steps:<n>, rather than once at the end. Needs the chrome browser backend: the builtin webview can fall back to a DOM snapshot that mints no accessibility refs, and a flow recorded then saves but can never run.
Throw away a REPL session and its state; the next repl_run starts a fresh runtime.
Run JavaScript in a persistent Node runtime and get the return value back. State survives between calls: variables (including top-level let/const), required modules, and open handles are still there next call. Top-level await works, but declarations inside an awaiting snippet do not persist — assign to a global (x = await f()). Full fs/net/require access, NO sandbox. Lives only as long as your MCP connection: no wmux restart, no sharing with other panes or workspaces. await browser.X(args) drives the browser like browser_repl (full profile only).
List this connection's REPL sessions: cwd, pid, age, and current state.
browser_repl and browser_replay lack output schema documentation. Both describe return values in prose ('{ text, events }' for repl; 'cannot find element' for replay) but provide no structured JSON schema. LLMs cannot plan downstream tool usage without knowing output shape.
Error handling is absent across all tools. No recovery guidance, no categorization of retryable errors. browser_replay states 'A run that cannot find an element stops at that step and reports why' but provides no machine-readable error structure or next-step guidance.
repl_run description is vague on dangerous capabilities. States 'Full fs/net/require access, NO sandbox' in lowercase parenthetical, critical security posture buried in prose. No mention of audit logging, no guidance on when LLMs should use this vs safer browser tools.
browser_replay 'variables' parameter description is minimal: 'run: values for the {{placeholders}} the flow was saved with.' Does not explain format, constraints, or validation. How does LLM know what placeholders exist? No link to saved flow metadata.
repl_run accepts 'session' parameter but no description of session lifecycle, visibility, or isolation between concurrent calls. Can concurrent agents see each other's sessions? Are session names globally scoped or per-user? Undocumented dependency.