Open source AI companion with terminal, voice, and SMS interfaces. Provides MCP tools for integrations, actions, browser control, display management, notifications, and AI agent automation.
Nero OSS exhibits solid tool definitions with clear, action-oriented naming and consistent parameter documentation. However, schema definitions lack formal JSON Schema structures with explicit type constraints, and several critical patterns are missing: no tool annotations (readOnlyHint/destructiveHint/idempotentHint), no output schema documentation, minimal error guidance, and no pagination support despite tools that could return unbounded results. The server shows good intent in descriptions (averaging ~120 chars, within baseline) but falls short of production-grade polish. 17 of 27 tools (63%) have adequate descriptions; 10 tools lack comprehensive parameter constraint documentation. No input validation errors or recovery guidance visible in tool definitions.
Ask the user to make a decision (or a few) and WAIT for their answer. Use this for genuine forks only they can settle (which approach, which option, confirm before something irreversible) instead of guessing or burying it in prose. Pass one or several questions; a focused card appears and this blocks until they pick, then returns their choices so you keep going in the same turn. The user can always type "something else", so keep your options to the likely paths. Do not ask what you can reasonably decide yourself.
Move an existing action to a different dial slot, or unbind it with slot -1.
Get an OAuth link to connect a built-in integration (e.g. google). Give the link to the user to open in their browser; once they approve, it connects automatically. The integration's secrets must be set first (see list_integrations).
Click an element by its ref from the latest read_page. Returns the new snapshot.
Fill a field with a stored secret WITHOUT ever seeing its value. Give the field ref and the secret's name (e.g. HULU_PASSWORD); the server types the real value into the page. Use this for every password/credential. If the secret isn't set, you'll be told to request it. Returns the new snapshot.
No JSON Schema type definitions visible for input parameters. Tools declare parameters with descriptions but lack formal 'type' (string|number|boolean|array|object) and constraint attributes (enum, minLength, maxLength, minimum, maximum, pattern). This prevents LLMs from validating inputs and makes parameter semantics ambiguous.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. Tools are labeled with Risk levels in metadata (READ_ONLY, WRITE, DESTRUCTIVE) but these are not exposed to the LLM via MCP tool annotations. LLMs cannot determine which tools are safe to retry or which require confirmation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Navigate an open browser panel to a URL. Returns the new page snapshot.
Scroll the page down (positive) or up (negative). Returns the new snapshot.
Type text into an input element by ref (from the latest read_page). For passwords or any credential, DO NOT use this - use browser_fill_secret. Returns the new snapshot.
Connect an external integration (MCP server) so you gain its tools. For OAuth servers this returns a link the user must click to authorize; share that link with them. Known: Lux (https://api.luxdb.dev/mcp), GitHub, Google. Pass api_key only if the user provides one.
Author a one-press action and bind it to a slot on the orb's radial dial. Use when the user wants a shortcut for something they do repeatedly (a script, a deploy, a status check, a canned request to you). Prefer `script` for anything you can do in shell; use `prompt` when the press should just start a conversation turn with you.
Throw an interactive panel onto a screen, a dashboard, controls, a graph, media, status, anything you want to show. Defaults to the screen you're on. Supports various component types including text, buttons, images, charts, metrics, and more.
Delete a dial action for good.
Disconnect and forget an integration.
List the user's dial actions and which slot each occupies. Use before binding a slot so you know what you'd displace.
List the screens on the network you can move to or throw panels onto: their names, ids, sizes, online status, and which one you are currently on.
List Nero's built-in integrations and their status: needs-secret (missing API credentials the user must set), needs-auth (secrets set, the user just needs to authorize), or connected. Use this to know what you can do and what setup a request needs.
List configured integrations and which are currently connected, with their tool counts.
Move yourself (the orb) to a different screen. You can only be in one place at a time. Use list_devices first if unsure of names.
Reach the user off-screen with a push notification (delivered to their Nero app). Use only when something genuinely needs them while they're away (a long job you finished, a deadline approaching, a reply they asked you to watch for) - keep it rare and worth the buzz, never for chit-chat. If it reports no device, the user hasn't opened the app / allowed notifications yet.
Open a live, interactive web page inside a panel on the Field - a real browser you (and the user) can see and click. Use to SHOW or operate the web: a dashboard, docs, a site, search results, a logged-in page. The user can click/scroll/type in it. Open the DIRECT url for what they want. NOTE: DRM video (Hulu/Netflix) will not play here (blank) - for actually watching, use open_url instead.
Open a URL in the user's own browser (fire-and-forget). Use when they want to actually WATCH or use something that can't be embedded - streaming (Hulu, Netflix, YouTube fullscreen), logged-in or paywalled sites, or anything DRM. Be precise: open the DIRECT url for the specific thing (the exact episode, article, PR, doc), never just a homepage. If you don't know the direct url, search for it first, then open it.
Read the current page of an open browser panel as a numbered list of interactive elements (links, buttons, inputs) plus the page text. Call this before acting, and note that refs are only valid for THIS snapshot - they change after any click/navigation, so re-read.
Change what an existing Dial button does, in place. Use this when the user wants one of their buttons fixed or altered - it keeps the same slot and survives a failed attempt, unlike deleting and rebuilding.
Fire one of the user's Dial buttons yourself. If they already have a button for something, press it rather than reaching for an MCP server or writing a one-off command: it's what they built and it's faster.
Commit the working draft to the dial slot. Only call this after test_action actually did the right thing.
Save the configured goal for an agent button, once you've asked the user what it should do. The goal is what you'll run every time they press it, so write it as a complete instruction to yourself.
Run a draft action right now and see what comes back. Do this before saving, every time - for a light or a speaker this is also how you confirm it physically worked. Reference credentials as ${NAME}; they resolve at run time and you never see the value.
No output schema documentation. Tools lack explicit documentation of what fields they return, their types, and structure. This forces LLMs to infer output shape and breaks downstream tool chaining when reference IDs are needed.
List tools (list_integrations, list_actions, list_devices) lack pagination parameters (limit, offset, page_size, cursor). No mention of result caps or how to handle large datasets. If any integration returns 100+ items, the entire result floods the context.
No error handling or recovery guidance visible in tool definitions. Error responses are not documented; LLMs have no instruction on what to do if a call fails (retry vs ask user vs skip). Example: if authorize_integration fails, should the LLM guide the user to set secrets first?
Duplicate tool names: 'list_integrations' appears twice (api/src/mcp/tool.ts and api/src/services/mcp/tool.ts). This is ambiguous for tool resolution and suggests inconsistent registration or namespace collision.
Parameter 'action' in run_action and revise_action accepts either an ID or a label (case-insensitive). No enum constraint or pattern documented. LLMs may pass invalid values; tool should validate and return clear error if no match found.
Icon parameter in save_action and create_action references an undocumented 'icon list'. LLMs cannot know valid values without seeing the list. Either embed enum values in the schema or provide a discovery tool (get_icon_list).
Questions parameter in ask tool expects a JSON array string but the schema shows type 'string' with no validation rule (minLength, maxLength, pattern). LLMs may pass malformed JSON; should validate and return clear parse-error guidance.
slot parameter in create_action and assign_action_slot accepts 0-7 or -1, but no explicit numeric constraints (minimum, maximum) documented. LLMs may pass out-of-range values (8, -2, 100).