This Android automation MCP server demonstrates basic tool structure but has significant gaps in definition quality. All 9 tools are explicitly registered with @mcp.tool() decorators in main.py, and schemas are visible for 8 of 9 tools. However, descriptions are present but often generic and lack LLM-optimized guidance. Parameter descriptions exist but are minimal. No output schemas are documented. Error handling is present but returns plain text rather than structured recovery guidance. The server follows a reasonable naming convention (mobile_* prefix with verb-noun pattern) but lacks composition patterns and actionability. Average tool score: 52/100.
Click on a specific coordinate on the Android screen. Args: x: X coordinate to click y: Y coordinate to click
Get UI elements from Android screen as JSON with hierarchical structure. Returns a JSON structure where elements contain their child elements, showing parent-child relationships. Only includes focusable elements or elements with text/content_desc/hint attributes.
Initialize the Android device connection. Must be called before using any other mobile tools.
Press a physical or virtual button on the Android device. Args: button: Button name (BACK, HOME, RECENT, ENTER)
Launch an application by its package name. Args: package_name: The package name of the app to launch (e.g., 'com.android.chrome')
List all installed applications on the Android device. Returns a JSON array with package names and application labels.
No output schemas documented. Tools return unstructured strings or images without declaring expected field names, types, or structures. LLMs cannot plan downstream tool calls or extract specific fields reliably.
Descriptions are generic and lack LLM-optimized guidance. Most are under 100 characters and lack WHEN/WHY context. Example: 'List all installed applications on the Android device.' does not explain when to call it (discovery phase?) or what structure is returned.
mobile_dump_ui and mobile_click have hidden interdependency not documented. mobile_click validates coordinates against ui_coords set populated only by mobile_dump_ui, but this prerequisite is not stated in mobile_click's description. LLM may call click without first calling dump_ui, causing 'Invalid elements coordinates' error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Perform a swipe gesture on the Android screen. Args: start_x: Starting X coordinate start_y: Starting Y coordinate end_x: Ending X coordinate end_y: Ending Y coordinate duration: Duration of swipe in seconds (default: 0.5)
Take a screenshot of the current Android screen. Returns an image object that can be viewed by the LLM.
Input text into the currently focused text field on Android. Args: text: The text to input submit: Whether to submit text (press Enter key) after typing
Error messages are plain text strings without recovery guidance or actionable next steps. Example: 'Error: Device not initialized. Please call mobile_init() first...' is returned as a string in the response, not as a structured error with category or retry guidance. No pattern distinguishes retryable vs fatal errors.
Enum constraints missing. mobile_key_press accepts 'button: str' with documentation listing (BACK, HOME, RECENT, ENTER) but no enum constraint in schema. LLM may pass invalid button names like 'POWER' or 'VOLUME_UP', causing runtime failures. Schema should use enum: ['BACK', 'HOME', 'RECENT', 'ENTER'].
No pagination or result limits documented. mobile_dump_ui and mobile_list_apps return potentially large structures (entire UI hierarchy or full app list) without limits or pagination. Large results can blow context window. Descriptions should state result caps (e.g., 'Returns first 50 apps') or include limit parameters.
Parameter descriptions are minimal. Examples: 'X coordinate to click' (12 chars) and 'Y coordinate to click' (21 chars) lack range/format guidance. Should state expected bounds (e.g., '0-1440 for standard Android screen width') or reference output of mobile_dump_ui for valid coordinates.
No confirmation step for destructive operations. mobile_swipe, mobile_click, mobile_key_press, and mobile_launch_app modify device state but lack dry-run or confirmation-request patterns. An LLM mistake (e.g., clicking wrong coordinate, launching wrong app) has irreversible consequences but no guard.
mobile_dump_ui returns 'str' (JSON as a string) instead of structured object. Returning JSON as a string requires LLM to parse it, wasting tokens and introducing parse errors. Should return parsed dict/object with typed fields so LLM can directly access coordinates and text.
No tool composition or chaining support. mobile_list_apps returns app labels and package names, but no direct link to mobile_launch_app. After discovering an app via list_apps, LLM must manually extract package_name and pass to launch_app. Could simplify by offering a single 'search_and_launch_app' tool or ensuring mobile_list_apps output includes IDs downstream tools need.