Cross-platform Electron desktop chat application with agent tool execution, browser automation, bash command execution, and background task management.
King Louie provides 7 browser automation and task management tools with mostly complete input schemas and descriptive text. However, several critical issues limit score: (1) Output schemas are undocumented, no tool declares what it returns, breaking the pattern:response-shaper requirement and forcing LLMs to guess downstream field names. (2) Tool descriptions vary widely in quality, some like BrowserSession are comprehensive (500+ chars) while AskUser is minimal (95 chars). (3) Parameters lack consistent type annotations and constraints, e.g., BrowserPage.action has an enum but many selector parameters accept free-form strings with no format guidance. (4) Error handling guidance is absent, no tool describes what errors could occur or how to recover. (5) Composition issues: BackgroundTask and TaskStatus should be more tightly integrated to prevent agent confusion about task lifecycle. Evidence: BrowserSession description includes detailed action sequences and vault explanations (excellent); AskUser and TaskStatus descriptions are generic and underdescriptive; BrowserPage and BrowserExtract have 15+ action types each with overlapping semantics (e.g., both can interact with frames), risking LLM confusion. No tool explicitly states idempotency, retry safety, or side effects. Baselines: Average tool description is 280 chars (good), but range is 95 - 550 (high variance). All tools have input schemas with enums where appropriate, but none document output structure.
Ask the user a question and wait for their response. Use when you need clarification or input.
Spawn a task that runs in the background while you continue the conversation. The task runs an agent independently and stores its output for later review. Use TaskStatus to check progress, read output, or stop background tasks. NOTE: the bg agent gets its own tool runtime — it CAN call gated tools (Browser, Vault, Bash, Edit, Write, Git) and approval prompts will surface in this same chat, but it does NOT share the foreground browser instance. For long-running browser sessions, prefer to drive them inline rather than handing them to a bg task.
Execute shell commands for local development tasks.
Extract data from the active page, drive iframes, and intercept network requests. Read actions: content (full HTML), title, get_text, get_attribute, get_value, is_visible, count, bounding_box. Frame actions: list with "frames" then use *_in_frame variants targeting frame_name / frame_url / frame_index. Network: route_block (drop matching requests), route_fulfill (mock responses), unroute. Console: returns recent browser console output.
Navigation + element interaction + waits on the active browser page. Shares one singleton with BrowserSession and BrowserExtract — start a session first. Selector strategies (pass as "selector"): - CSS: "div.class" or "#id" (default) - Text: "text=Login" - Role: "role=button[name='Submit']" - Label: "label=Email" - Placeholder: "placeholder=Search..." - Test ID: "testid=submit-btn" - XPath: "xpath=//div[@id='app']" - Alt: "alt=Company logo" - Title: "title=Close dialog" All element actions auto-wait for the element to be actionable. Avoid wait_for_load_state state="networkidle" on chatty sites (LinkedIn, modern dashboards) — they stream telemetry forever and you'll burn the timeout. Prefer wait_for on a specific element or wait_for_url.
No tool documents its output schema or return structure. LLMs cannot infer what fields are available for downstream chaining (e.g., what does BrowserExtract.content return? Is it a string, object, or structured DOM tree?). This violates pattern:response-shaper and forces agents to guess or make discovery calls.
Error handling is undocumented. Tools do not describe what errors they may raise, which are retryable, or how to recover. For example, BrowserPage and BrowserExtract can timeout, but timeout guidance is missing. This violates pattern:recovery-guide and pattern:error-classification.
AskUser description is generic (95 chars) and does not explain when or why to use it vs. other input methods. Follows pattern:tool-description minimally. Severity: medium because the tool is simple, but the description leaves LLMs uncertain about appropriate use cases.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 58 | 2026-07-28+ | v2 |
Browser lifecycle, persistent profiles, credential vault, auth, tabs, cookies, storage state, viewport, and high-level credentialed flows. Shares one singleton browser instance with BrowserPage and BrowserExtract. Typical sequence: profile_list → start (with profile) → fill_credentials or login → use BrowserPage to interact. PROFILES persist cookies/localStorage between runs: - profile_list / profile_create / profile_delete / profile_current - start with params.profile="<name>" attaches that profile. VAULT CREDENTIALS (King-Louie's encrypted vault, not Chromium's): - save_credentials (profile?, host?, username, password) — stores creds. - fill_credentials (profile?, host?, submit?) — types stored creds into the current page's auto-detected username/password fields. Form-based logins only. - set_http_auth (username+password OR profile+host) — for HTTP Basic auth (the OS-style popup, not an HTML form). Call BEFORE navigating, or after a chrome-error://chromewebdata/ 401. PREFER HIGH-LEVEL: login / signup / fill_payment auto-detect fields across sites — use these instead of hand-rolling fill+click sequences.
Check the status of background tasks, read their output, list all tasks, or stop a running task.
BrowserPage and BrowserExtract provide overlapping selectors and frame interactions with 15+ action enums each. No clear guidance on which to use when. E.g., both can fill fields, click elements, and interact with iframes. This risks LLM confusion and violates pattern:tool composition principle (each tool should do one thing).
TaskStatus description is minimal and does not explain the task lifecycle (created → running → completed) or what status values are returned. Agents cannot predict task state transitions without more guidance.
Bash tool lacks timeout and safety guidance. The description says 'for local development tasks' but does not state: (1) what timeouts apply, (2) whether commands are sandboxed, (3) what commands are blocked, (4) how stderr/stdout are captured. This is critical for a destructive tool (risk: DESTRUCTIVE).
No tool explicitly states idempotency or retry safety. For a browser automation suite, this is important, e.g., does clicking a button twice double the action, or is it idempotent? This information should be in tool descriptions to guide agent retry logic.