Execution intelligence for browser automation agents. Exposes AIR execution intelligence as tools via stdio transport for discovering capabilities, executing actions, and reporting outcomes on websites.
AIR SDK provides 5 tools with complete input schemas and descriptions. Naming is action-oriented (extract_*, browse_*, execute_*, report_*). Descriptions are detailed (194 - 400+ chars) and explain WHEN to use each tool. However, output schemas are not documented in the visible code, LLMs cannot predict response structure. Error handling guidance is minimal. Parameter descriptions are thorough but some lack explicit constraints (e.g., 'domain' format, 'params' object structure). The report_outcome tool is complex with many optional nested fields, increasing cognitive load. No tool annotations (readOnlyHint/destructiveHint) are present despite clear read/write semantics.
Step 1 of 3: Discover what actions can be automated on a website domain. Returns capabilities (search, login, add to cart, etc.) with confidence scores, execution tiers, selectors, and macro availability. Even if a domain has no indexed capabilities, universal patterns from similar sites may provide execution guidance via execute_capability. Always follow up with execute_capability (step 2) and report_outcome (step 3).
Step 2 of 3: Get execution guidance for a specific capability. Returns the optimal path: API fast-path, pre-verified macro steps, selector hints, or context for unverified capabilities. Call this after browse_capabilities. After executing with your browser tools, you MUST call report_outcome (step 3) with the CSS selectors you used.
Extract text and structured sections from a file (PDF, DOCX, PPTX, XLSX, CSV, TXT, MD, HTML, EML). Returns the same ExtractionResult shape as extract_url: title, sections (as content items), full text, and diagnostics. Any agent can use this — no browser needed. Accepts a file URL (HTTP/HTTPS) which will be fetched and uploaded, or a local file path for desktop agents. Use this for document import pipelines: extract_content → research_import.
Extract structured data from any URL. Returns title, description, content items, metadata, and diagnostics. Handles SPAs (YouTube, Next.js, Nuxt), JSON-LD, RSS/Atom feeds, JSON APIs, and standard meta tags. Best results on: sites with JSON-LD, RSS/Atom feeds, JSON APIs, and SPAs with standard state globals. Limited results on: retail sites (Amazon, Walmart), CAPTCHA-protected sites, and heavy SPAs — for these, use browser tools (get_page_text, screenshot) and report what you find via report_outcome with browserObservations. Costs 1.0 credit (fresh) or 0.1 credit (cached). Meta-only results cost 0 credits.
Output schemas not documented. LLMs cannot predict response structure (title, description, content items, metadata, diagnostics). Agents must infer field names and types from examples or trial-and-error.
report_outcome has 10+ optional nested fields (steps array with 6 sub-fields, browserObservations object with 3 sub-fields). No guidance on which combinations are valid or required. Increases agent confusion and error rates.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). extract_url and extract_content are clearly read-only; report_outcome is destructive (creates macros). Agents cannot infer safety properties without annotations.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 65 | 2026-07-28+ | v2 |
REQUIRED: Report the outcome after executing ANY browser action via AIR. This is the final step of every browse → execute → report workflow. Without your report, the action you just performed is lost — no macro gets created, no future agent benefits. You MUST include the exact CSS selectors you used in each step. Steps with empty selectors are discarded. Example steps: [{ action: "click", selector: "input#twotabsearchtextbox", success: true }, { action: "fill", selector: "input#twotabsearchtextbox", value: "query", success: true }, { action: "click", selector: "input#nav-search-submit-button", success: true }]. Look at the CSS selectors from your browser tool calls and copy them exactly into each step. If you used browser tools (get_page_text, screenshot) because extract_url was insufficient, include browserObservations to help improve future extractions on this domain.
Parameter 'params' in execute_capability is additionalProperties: {type: 'string'} with no schema for expected keys. LLMs cannot validate capability-specific parameters without knowing what each capability accepts.
Error handling guidance missing. No recovery hints (e.g., 'If extraction fails, try extract_content or browser tools'). Agents cannot self-correct on failures.