An open-source generalist agent for the enterprise, supporting complex task execution on web and APIs, OpenAPI/MCP integrations, composable architecture, reasoning modes, and policy-aware features.
CUGA Agent presents a mixed portfolio. Two tools (spawn_agent, get_agent_result) have excellent descriptions and schemas; most browser/action tools have minimal descriptions and incomplete parameter documentation. Six of twelve tools lack meaningful descriptions (under 50 chars). Parameter schemas are inconsistent: some tools (click, select_option, type) have typed inputs with descriptions, but others (go_back, scroll, memorize, human_in_the_loop) have either no parameters or trivial descriptions. Output schemas are entirely undocumented across all tools. Error handling and recovery guidance are absent. Security concerns exist around spawn_agent's task parameter (4000 char limit, unclear sanitization against prompt injection). Composition is fragmented: browser navigation tools (go_back, scroll, click, type, select_option, open_app) are tightly coupled but lack a discovery tool to reveal clickable elements or page structure. Email tools (read_emails, send_email) have generic descriptions and no parameter validation or rate-limiting hints. Overall, this reads as an experimental/research system rather than a production-grade tool suite.
Click an element.
Wait for and retrieve the result of an async spawn_agent call.
Go back to previous page.
Facilitates communication between the agent and the user, allowing the agent to seek input or permission based on environment policies or complex decision-making scenarios.
Memorize key information for later!
Open an application
Read emails from Gmail inbox with dummy data
Six tools (go_back, memorize, human_in_the_loop, scroll, read_emails, send_email) have descriptions under 50 characters, violating pattern:tool-description baseline of 194 chars average. Descriptions like 'Go back to previous page.' and 'Scroll the page' lack context for LLM selection or disambiguation from similar tools.
Browser action tools (go_back, click, select_option, type, scroll) lack output schema documentation. No documentation of what these tools return, whether they return success booleans, new page content, element lists, or error status. This forces LLMs to guess downstream behavior.
Four tools (go_back, memorize, human_in_the_loop, scroll) have zero input schema visibility. 'go_back' and 'scroll' have empty input objects but no parameter definitions. 'memorize' and 'human_in_the_loop' have single 'information'/'message' params with descriptions, but the schema structure is unclear from source.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Scroll the page
Select one or more options in a dropdown – delegate to implementation.
Send an email using Gmail with dummy data
Spawn a SubCuga subagent with fresh context to handle a task independently. The subagent inherits all your tools and runs without any prior conversation history. Pass the complete task description — everything the subagent needs to succeed. When including prior results in task, build it with string concat + json.dumps (never nested f-strings/triple quotes around those values). mode='sync' (default): blocks until the subagent finishes — use only for a single sequential subtask. mode='async': returns a future_id immediately so you can spawn multiple subagents in parallel before collecting results with get_agent_result. Use mode='async' whenever you have two or more independent subtasks that could run simultaneously. share_workspace=False (default): isolated empty workspace. share_workspace=True: same workspace both ways — child reads parent uploads/files and anything the child writes (reports, .md, outputs) is visible in the parent session (avoid for parallel async writers on the same files).
Fill out a form field. It focuses the element and triggers an input event with the entered text. It works for <input>, <textarea> and [contenteditable] elements. use press_enter true when the search input field requires pressing enter after filling the element.
spawn_agent's 'task' parameter accepts 4000-char free-form text with no prompt injection sanitization guidance. Description warns against nested f-strings and triple-quote formatting, but does not address SQL injection, command injection, or template injection risks when the task is passed to a child agent. No sanitization guidance.
Browser navigation tool suite is tightly coupled but lacks discovery. No tool to query page structure, clickable elements, or available options before calling click, select_option, or type. LLMs must blindly pass element IDs ('bid') without knowing what elements exist. This violates pattern:tool-chain composition.
Email tools (read_emails, send_email) have no error handling guidance. No documentation of what happens if recipient is invalid, server is down, or rate limit is hit. Description of send_email says 'with dummy data', suggesting this is a mock/test tool, but no indication of whether real Gmail is behind it or if this is safe to call.
No documentation of idempotency or side effects. spawn_agent with mode='async' returns future_id but no guidance on retry behavior if get_agent_result fails. Click, type, select_option describe what they do but not whether they are idempotent if retried (e.g., clicking the same button twice). This violates pattern:idempotent-operation.
Parameter naming inconsistency: 'bid' (element identifier) is cryptic and not self-documenting. click, select_option, type all use 'bid' with description 'Element identifier', but LLMs cannot infer what a 'bid' is without domain knowledge. Should be 'element_id' or 'browser_element_id' to improve clarity.