MCP server for tapflow — lets LLM agents control iOS/Android simulators via Model Context Protocol
Tapflow MCP server has 8 tools with clear action-verb naming (list_, connect_, boot_, shutdown_, query_, run_). All tools have descriptions (avg ~150 chars, within baseline 34-392 range). Input schemas are present with parameter descriptions. However, several tools lack output schema documentation, and some descriptions could be more explicit about prerequisites and error conditions. Tool composition is sound, each tool has a single responsibility. Parameter naming is consistent (sessionId, deviceId, buildId). No tool annotations (readOnlyHint/destructiveHint) despite clear risk levels documented in metadata.
Boot a simulator/emulator. Requires connect_device first. Waits up to 30 seconds for the device to be ready.
Join a device session so you can control it. Required before boot_device, install_app, and launch_app.
End a device session and release the connection.
List all apps and their builds available on the relay. Use this to find buildId before calling install_app or launch_app.
List all available simulators and emulators registered on the tapflow relay.
Query the accessibility tree of the current screen: interactive and text-bearing elements as { role, label, identifier, frame, enabled, rawRole }. Frames are normalized 0-1 relative to the screen. Prefer this over guessing coordinates from a screenshot: to tap an element, multiply the frame center by the screenshot pixel size — x = (frame.x + frame.width / 2) * screenshotWidth, y = (frame.y + frame.height / 2) * screenshotHeight — and pass those pixel coordinates to the tap tool.
Output schemas not documented. Tools like list_builds, list_devices, query_ui_tree, and run_flow lack explicit return type documentation. LLMs cannot plan downstream calls or extract required fields without knowing response structure.
Tool annotations missing. Risk levels (READ_ONLY, WRITE, DESTRUCTIVE) are documented in metadata but not exposed as tool annotations (readOnlyHint, destructiveHint, idempotentHint). Agents cannot distinguish safe from dangerous operations without reading descriptions.
Error handling guidance incomplete. Tools like boot_device and shutdown_device mention 30-second timeouts but do not document what happens on timeout, whether it is retryable, or what the LLM should do next.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 73 | 2026-07-28+ | v2 |
Replay a tapflow flow (YAML) deterministically — no LLM in the loop. Use this for verified scenarios instead of tapping step by step: author the flow once, then replay it idempotently. Pass the YAML inline via "flow", or a file path via "path" (resolved from the MCP server process cwd). Steps: clearState / launchApp / tapOn / inputText / pressKey / swipe / scroll / openUrl / assertVisible / assertNotVisible. launchApp launches the buildId argument. When buildId is set, the build is installed before replaying (like `tapflow flow run --build`) so clearState/launchApp have the app present — pass install:false to skip. Returns per-step results; on failure a screenshot is saved to a temp file.
Shut the session's booted simulator/emulator down — powers the device off to free resources or force a cold boot next time. Unlike disconnect_device (which only leaves the session, leaving the device running), this actually stops the device. Requires connect_device first. Waits up to 30 seconds.
Parameter constraints not formalized. run_flow accepts 'flow' (YAML inline) or 'path' (file path) but does not declare them as mutually exclusive. LLMs may pass both, causing ambiguous behavior.
Discovery tool output limits not stated. list_builds and list_devices do not specify max result count or pagination support. Large device/build inventories could blow context windows.