A parallel AI operator with physical capabilities (webcam, voice, macOS control, memory) exposed as MCP tools for Claude Code / Claude Desktop integration.
HAL9000 has significant definition quality gaps. While 26 tools are defined with names and descriptions, the quality is uneven. Most tools lack comprehensive parameter documentation, lack output schemas, and lack structured error handling guidance. Tool descriptions are present but often lack actionable guidance on WHEN to use them vs similar tools. Parameters frequently lack type constraints (enums, ranges) and validation rules. Most critically, tools lack documented output schemas, forcing LLMs to guess at response structure. Security-sensitive tools (screenshot, webcam, clipboard, app control) lack permission scoping or audit trail documentation. The median tool description is adequate (~150-180 chars) but several fall below the 50-char minimum for clarity. No tool declares destructive hints or idempotent properties despite having irreversible operations (delete memory, quit app, volume/brightness changes). The composition is reasonable (26 single-purpose tools), but chain-ability is hindered by undocumented output fields.
Send an automation command to a running application. On macOS this uses AppleScript; on Windows this uses PowerShell COM. Use this to control apps AFTER they are opened — create documents, insert text, save, close, navigate tabs, etc.
Send a message to HAL and get his spoken response back. HAL will think about your message, speak the reply aloud, and return the text. Use this for bidirectional conversation between Claude Code and HAL. Example: hal_chat("I've finished reviewing the code. Found 3 issues.") HAL will process this as if the user said it, respond with his personality, and speak the response aloud through the browser.
Fetch a webpage's text content. Returns the page title and full text (with HTML stripped).
Remove a memory by ID. Permanent deletion.
Load context at session start. If a prior session was saved, returns its chat history, artifacts, and task state. Returns empty dict if no prior session.
List all stored memories, with optional type filter. Returns paginated results.
Missing output schemas: 13 tools lack documented response structures (hal_see, hal_screenshot, hal_listen, hal_chat, hal_recall, hal_save_session, hal_get_context, macos_wifi, macos_battery, hal_fetch_url, list_running_apps). LLMs cannot plan downstream calls or extract chaining IDs without knowing response fields.
No tool annotations (destructiveHint/readOnlyHint/idempotentHint). Destructive tools like hal_forget, quit_application, and macos_volume/brightness lack hints. hal_remember and hal_save_session lack idempotentHint. Agents cannot reason about safety or retry semantics.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | 2026-07-28+ | v2 |
Listen through the microphone for speech and return the transcription. Waits for the user to speak, records until silence, then transcribes. Returns the transcribed text, or an error if no speech was detected. Timeout is approximately 6 seconds of silence before giving up.
Search HAL's persistent memory for facts matching a query. Returns matching entries with type, content, source, and timestamp.
Store a fact in HAL's persistent memory. This survives restarts. Use this to remember user preferences, project context, names, decisions, etc. type: fact (default), decision, preference, task. source: who is storing this — claude_code (default), hal, user.
Save the current session context (recent chat, artifacts, task state) to persistent storage. Returns a session ID for later recovery.
Take a screenshot of the entire macOS screen and return the file path. Use this to see what's on the user's screen — code editors, browsers, design tools, terminal output, etc.
Capture a frame from the webcam and return it as a base64-encoded JPEG. Use this to see what's on the user's desk, who is in front of the computer, or understand the physical environment. The image is returned as base64 data that can be analyzed directly. NOTE: Requires HAL server to be running (python server.py) with Vision enabled.
Speak the given text aloud using HAL9000's cloned voice. Use this to give verbal feedback, read code aloud, announce results, or communicate with the user via speech. The voice is a cloned HAL9000 voice. Keep text concise — spoken output should be short declarative sentences.
Search the web via DuckDuckGo. Returns a list of results with title, URL, and snippet.
List installed applications. Optionally filter by a search query. Use this when the user asks 'what apps do I have', 'show my apps', 'find an app', or when you need to discover the correct app name before opening it. Returns app names and their install locations. Always prefer this over guessing app names.
List all currently running applications.
List, open, or quit macOS applications. Mode: list (default) to list running apps or search installed apps; open to launch an app; quit to close it.
Get battery status: percentage, is_charging, and time remaining (if available).
Get or set the display brightness. Specify a level 0-100 to set; omit to get current level.
Get or set clipboard contents. Pass text to copy to clipboard; omit to read current clipboard.
Send a macOS notification. Appears in Notification Center and (if Dock is configured) bounces the app.
Get or set the system volume level. Specify a level 0-100 to set; omit to get current level.
Get the current WiFi network name (SSID).
Open a desktop application by name (e.g. 'Safari', 'Chrome', 'Terminal', 'Claude').
Open a URL in the default web browser.
Quit a running application by name.
Inadequate parameter descriptions. Tools like hal_chat, hal_save_session, hal_get_context, macos_apps (mode/name/query) lack clear descriptions of when to use each mode or how modes interact. hal_remember type parameter ('fact, decision, preference, task') lacks usage guidance.
No input validation rules or constraints. Parameters like hal_remember (type enum), hal_list_memories (limit/offset), hal_web_search (max_results) lack range constraints. LLMs may pass invalid values (limit=999999, max_results=1000).
Security tools lack permission scopes and audit trail documentation. Sensitive tools (hal_see, hal_screenshot, macos_clipboard, app_action, quit_application) do not declare required permissions (e.g. read:camera, write:clipboard). No audit-trail pattern documented.
No error recovery guidance. Tools return generic errors without actionable next steps. E.g., hal_fetch_url likely fails on 404, timeout, or CORS, no guidance on retry, alternative sources, or fallback.
Missing mutual-exclusivity and dependency documentation. Tools like macos_apps (mode=list/open/quit with conditional name/query params), hal_remember (optional type/source), hal_list_memories (optional type filter) lack clear parameter interdependencies.
Tool naming ambiguity: macos_apps covers list/open/quit, three separate concerns. Consider splitting into list_running_apps, list_installed_apps (already exist separately), open_application, quit_application. Current overlap creates agent confusion on which to call.
No pagination or result limits documented. hal_web_search claims 'max 20' but tools like hal_list_memories and app responses lack documented result caps. Large responses (e.g. all installed apps, all memories) could exhaust context.
Documentation states 'Requires HAL server to be running' (hal_see) but no tool declares upstream dependency checks or clear error messages if HAL is offline. Agent cannot disambiguate 'permission denied' from 'server offline'.