Mobile + canvas automation MCP for AI agents — one stdio server, 51 tools: iOS (simulator + real) & Android control, native UI automation, Maestro E2E flows, evidenced oracle-ladder assertions, React Native/Metro debugging, WebView DOM + network capture, and a no-vision canvas/WebGL brain (Pixi/Konva/Fabric/Phaser/Three/Babylon) that drives game UIs like DOM elements — ~5x fewer tokens than screenshot/vision loops.
Podium MCP exhibits strong naming conventions and functional schemas across 45 tools, but suffers from inconsistent parameter descriptions, missing error handling guidance, and output schema documentation. Tool names are generally verb-forward (device_boot, app_launch, tap_on, screenshot) and appropriately scoped. However, parameter descriptions are minimal, most are 1-2 line summaries ('Device UDID', 'Bundle ID of the app') without context on valid ranges, formats, or dependencies. Output schemas are not documented in the visible source; tool responses are inferred from e2e test assertions rather than declared. Error classification and recovery guidance are absent. Descriptions range 20-95 characters, mostly at the lower end of the baseline (avg 194 chars baseline). The server implements 45 tools covering mobile automation, debugging, and canvas inspection, a substantial and well-organized tool set, but definition quality lags behind orchestration sophistication.
Install an application on a device
Launch an application on a device
List installed applications on a device
Get the current state of an application
Terminate a running application on a device
Uninstall an application from a device
Assert that text is NOT visible on the screen
Parameter descriptions are minimal (mostly 1-2 line summaries) and lack context on valid ranges, formats, and dependencies. E.g., 'Device UDID' for udid param does not explain what a UDID is, format constraints, or how to obtain one.
Output schemas are not documented in tool definitions. Responses are inferred from e2e test assertions rather than declared in tool schemas. LLMs cannot reliably determine what fields to expect or how to chain tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 57 | 2026-07-28+ | v2 |
Assert that text is visible on the screen
Inspect canvas/WebGL element using pixel analysis
Tap a canvas/WebGL element
Get a quick reference cheat sheet for common Maestro flow actions
Retrieve a specific crash report
List crash reports on a device
Boot an iOS simulator or Android device
List available iOS simulators and Android devices
Inspect game engine objects via AltTester
Tap a game engine object via AltTester
Export automation steps to Maestro flow format
Inspect the accessibility hierarchy of the device screen
List React Native Metro debugging sessions
Capture console logs from a Metro-connected React Native app
Capture network traffic from a Metro-connected React Native app
Get the state of a Metro-connected React Native app
Open a URL on a device
Get the current screen orientation of a device
Set the screen orientation of a device
Health check for the Podium MCP server
Press a key or button on a device
Start recording a video on a device
Stop recording a video on a device
Run a Maestro E2E flow on a device
Run a sequence of automation steps on a device
Get the screen size of a device
Take a screenshot of a device
Set the simulated location on a device
Perform a swipe gesture on a device
Tap at specific coordinates on a device
Tap with fallback oracle ladder retry logic
Generate a token usage report for the latest interaction
Validate a flow with assertions and automated checks
Wait for a text element to appear on the screen
Evaluate JavaScript in a WebView
Inspect a WebView DOM element
Navigate a WebView (reload, back, forward)
Capture network traffic from a WebView
No error handling guidance. Tools lack descriptions of retryable vs fatal errors, actionable recovery steps, or what the agent should do when a tool fails. E.g., app_launch could fail due to device offline, app not installed, or app already running, each requires different recovery.
Complex parameter semantics undocumented. run_steps expects 'Array of step objects with action and parameters', what actions? what parameters? No schema, no examples, no linked documentation.
Some parameter names ambiguous without additional context. 'key' in press_key does not specify valid values (home, back, enter, etc., all string enums should be declared). 'direction' in swipe should be an enum constraint, not free-form.
Inconsistent tool composition. Multiple tools operate on the same resource (e.g., app_state, app_launch, app_terminate all manipulate app lifecycle) but no documented relationship or recommended calling order. Agent must infer intent.
Visibility of return structures varies. Some tools (e.g., screenshot) accept saveTo parameter but output schema is unclear, does it return file path? image data? both?
Description lengths are at the short end of the baseline distribution (avg 194 chars baseline; most Podium tools 50-95 chars). Descriptions lack 'WHEN to use' and 'WHY' context, forcing LLMs to guess tool intent.