Build and run autonomous AI agents with OpenClaw, Hermes, multiple model providers, orchestration, delegation, memory, skills, schedules, and chat connectors.
Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
SwarmClaw presents 41 tools with critical definition quality issues. While tool names follow verb_noun conventions (execute_command, read_file, write_file), the actual schema definitions are not visible in the provided source code. The tool-definitions.ts file is referenced but its content is not shown, preventing verification of input schemas, parameter descriptions, output structures, and error handling. Without access to the actual schema definitions, we cannot confirm the presence of required JSON Schema types, parameter descriptions, or output documentation. The sample source shows only Dockerfile, package.json, and test file references, no tool schema implementations. Tools spanning file I/O, web operations, delegation, and system management (41 total) lack demonstrable schema quality. The high-level tool list suggests broad functionality but offers no evidence of: (1) documented input parameters with type constraints and descriptions, (2) output schemas with field documentation, (3) error recovery guidance, (4) idempotency declarations, (5) permission/scope declarations. Critical tools like execute_command (WRITE), delete_file (DESTRUCTIVE), and delegate_to_agent (WRITE) have no visible schema documentation, parameter validation rules, or safety constraints. The descriptions provided ('Run shell commands in the working directory', 'Delete files or directories (when explicitly enabled)') are functional but minimal (34-53 characters), below the optimal 50-200 character range for LLM-optimized descriptions. Without seeing src/lib/tool-definitions.ts, we cannot verify whether parameters accept enums for constrained values, whether descriptions explain WHEN to use each tool vs. similar alternatives, or whether output schemas include chaining IDs needed for downstream tool calls.
Tools (41)
browserwritesource verified32/100
Browse the web, take screenshots, and interact with pages
check_delegation_statusread only38/100
Check the status of a delegated task
connector_message_toolwrite33/100
Send proactive outbound messages via running connectors
Inspect and document src/lib/tool-definitions.ts to expose actual JSON Schema definitions for all 41 tools. For each parameter, add: type (string|number|boolean|object|array), required flag, description (50-200 chars explaining WHAT, WHEN, and WHY), and constraints (enum values, min/max, regex patterns, length limits).
Rewrite tool descriptions following the LLM-optimized format: (1) What action does it perform? (2) When should the LLM call it instead of alternatives? (3) What does it return? (4) Any prerequisites or side effects? Target 100-200 characters. Example: 'Execute shell commands in the working directory. Use for running builds, tests, or system tasks. Returns stdout, stderr, and exit code. Irreversible, LLM cannot undo. Requires shell syntax validation.'
Consolidate 8 delegate_to_X_cli tools into a single delegate_to_agent tool with an explicit cli_provider enum parameter (claude_code|codex|opencode|gemini|copilot|droid|cursor|qwen). This reduces duplication, simplifies LLM decision-making, and follows single-responsibility principle.
Rename generic tools: 'memory' → 'store_memory' or 'recall_memory' (split concerns), 'browser' → 'web_browser_action', 'manage_*' → 'create_task'|'update_task'|'delete_task', 'whoami_tool' → 'get_agent_context', 'connector_message_tool' → 'send_connector_message'. Verbs must come first and clearly indicate the action.
Document output schemas for all tools. For each tool response, define: (1) root object type (object|array), (2) field names and types, (3) whether pagination is supported (include total_count, next_cursor, or limit parameters in responses). Example: execute_command returns {stdout: string, stderr: string, exit_code: number, command_truncated: boolean}.
Score history
Overall score trend
First recorded score · v2 rubric
43/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-23
F
43
2026-07-28+
v2
writeauth37/100
Delegate complex coding tasks to Codex CLI
delegate_to_copilot_cliwriteauth37/100
Delegate complex coding tasks to GitHub Copilot CLI
delegate_to_cursor_cliwriteauth37/100
Delegate complex coding tasks to Cursor Agent CLI
delegate_to_droid_cliwriteauth37/100
Delegate complex coding tasks to Factory Droid CLI
delegate_to_gemini_cliwriteauth37/100
Delegate complex coding tasks to Gemini CLI
delegate_to_opencode_cliwriteauth37/100
Delegate complex coding tasks to OpenCode CLI
delegate_to_qwen_code_cliwriteauth37/100
Delegate complex coding tasks to Qwen Code CLI
delete_filedestructive40/100
Delete files or directories (when explicitly enabled)
edit_filewrite43/100
Edit existing files with find-and-replace
execute_commandwrite40/100
Run shell commands in the working directory
google_workspacewriteauth35/100
Run Google Workspace CLI commands for Drive, Docs, Sheets, Gmail, Calendar, Chat, and more
Descriptions are minimal (30-53 characters), below the optimal 50-200 character range. They lack context on WHEN to use each tool, WHAT it returns, and WHY an LLM should select it over similar tools. Example: 'Run shell commands in the working directory' does not explain prerequisites, safety constraints, or error recovery.
7 tools follow 'delegate_to_X_cli' naming pattern, combining delegation concern with CLI provider specification. Should split into single delegate_to_agent + provider parameter, or consolidate with explicit enums for CLI provider selection. Current names lack clarity on parameter differences.
Tool names using generic nouns (memory, browser, google_workspace, manage_tasks, manage_sessions) are ambiguous. LLM cannot distinguish between tools without reading descriptions. Example: 'memory' could mean recall, store, or query, unclear what action it performs.
No evidence of output schema documentation. Cannot verify whether tools return structured objects with typed fields, pagination support, chaining IDs for downstream calls, or error categorization (retryable vs. user-fixable vs. fatal).
all
Add error handling and recovery guidance to descriptions. For destructive tools (delete_file, execute_command, delegate_to_agent): document failure modes, whether operations are retryable, how to check status, and what compensation tools are available. Example: 'On permission error, request higher-level access from the user. On timeout, retry with shorter timeout or break into smaller batches.'
For tools accepting lists (manage_tasks, manage_agents), document pagination: accept limit (1-100, default 20) and offset (0+) parameters; return total_count and items array. Prevents context window exhaustion and enables multi-page result handling.
For file operation tools (read_file, write_file, delete_file, list_files), add explicit path validation rules in descriptions: disallow '..' for path traversal, require absolute paths or paths relative to working directory, document max file size. Example: 'path must be < 4096 chars, no '..' components, not outside working directory'.
Add security-relevant parameter descriptions. For web_* tools and browser: accept URL constraints (https only, domain whitelist), return size limits. For delegate_to_* tools: document timeout, max retries, permission requirements. For execute_command: document shell syntax validation and command injection prevention.
Create a master tool composition matrix documenting tool dependencies: which tools should be called in sequence? What IDs does each return that downstream tools require? Example: spawn_subagent returns subagent_id → check_delegation_status requires subagent_id. This enables the LLM to plan multi-step operations without discovery detours.