Fast browser automation MCP server for LLMs — persistent Chromium, ref-based interaction, cookie migration
Pilot MCP demonstrates good foundational tool design with strong naming conventions (verb_noun pattern: pilot_snapshot, pilot_snapshot_diff, pilot_find) and comprehensive parameter descriptions. All three tools are READ_ONLY with well-structured input schemas using Zod validation. However, there are notable gaps: (1) Output schemas are not explicitly documented, the tool descriptions mention return types (e.g., 'Text representation of the accessibility tree') but provide no structured schema definition for LLMs to parse; (2) Error handling descriptions exist but lack actionable recovery guidance and categorization (retryable vs fatal); (3) Parameter descriptions are good but some lack explicit format constraints (e.g., selector accepts 'CSS selector' but no pattern validation is visible). The server excels at practical UX (refs like @eN for element selection, diff capabilities, token-saving options) but falls short on LLM-optimized schema completeness.
Find an element by visible text, label, placeholder, or role — without running a full snapshot. Use when you know what you want to click or fill but don't need to see the entire page tree. Returns a @eN ref immediately usable by pilot_click, pilot_fill, pilot_hover, and other interaction tools. Saves tokens compared to pilot_snapshot when you only need one element.
Capture an accessibility tree snapshot of the page with @eN refs for element selection. Use when the user wants to see the page structure, find elements to interact with, or get refs for click/fill/hover. This is the primary way to understand what is on the page. Refs from this snapshot are used by pilot_click, pilot_fill, pilot_hover, pilot_select_option, and most other interaction tools.
Compare the current page state against the previously captured snapshot, showing a unified diff of what changed. Use when the user wants to verify the effect of an action (click, fill, navigation), check if dynamic content loaded, or see what changed on the page without re-reading the entire snapshot. The first call stores a baseline; subsequent calls diff against it.
Output schemas not documented. Tools return text or file paths, but LLMs lack formal definitions of return value structure (field types, required vs optional fields). pilot_snapshot_diff mentions 'unified diff text' but doesn't specify if it's plain text, structured diff format, or field-by-field mapping.
Error handling lacks actionable recovery guidance. pilot_snapshot describes 'Timeout: The page is too complex or unresponsive' but doesn't suggest next steps (e.g., 'Try again with max_elements=50 or selector to narrow scope'). pilot_snapshot_diff says 'No baseline snapshot' but doesn't explain how the baseline is stored or reset.
Parameter 'selector' accepts CSS selectors but lacks format constraints or examples. Descriptions say 'CSS selector to scope' but don't indicate whether '#main-content', '[data-qa=sidebar]', or complex selectors like 'div.nav > ul:first-child' are supported.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 64 | 2026-07-28+ | v2 |
pilot_find parameter relationships undocumented. Parameters text, label, placeholder, and role are all optional, but it's unclear if at least one is required, if they're mutually exclusive, or if passing multiple refines the search. LLMs cannot infer this logic.
Tool composition gap: pilot_find returns refs (@eN) usable by unshown tools (pilot_click, pilot_fill, etc.), but those tools are not included in this server definition. This breaks the 'tool-chain' pattern, refs are only useful if downstream tools exist and are discoverable.