Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
RustyHand presents 30 tools with consistent but minimal documentation. All tools have names, descriptions, and JSON Schema inputs. However, descriptions average ~80-120 characters (well below the 194-char baseline for A+ tools), lack context about when to use each tool, and omit guidance on error recovery or side effects. Parameter descriptions exist but are frequently generic ('Agent ID to terminate' vs 'The UUID of the agent to kill, required for cleanup of spawned subprocesses'). No output schemas are documented anywhere in the visible code, critical for tools like agent_list, memory_recall, task_list, and image_analyze. Tool composition is reasonable (file operations are properly separated), but critical security and correctness concerns exist: shell_exec accepts free-form command strings with only 'taint checking for injection patterns' (unverified in code), file_delete and agent_kill are destructive with no confirmation pattern, and secrets management is not visible. The 30-tool set is large and sprawling, many could be composed into fewer, more focused tools (e.g., agent_send + agent_spawn + agent_kill could reduce cognitive load). Error handling is absent from visible schemas, no guidance on what errors are retryable, what users should fix, or how to recover from partial failures.
Tools (30)
agent_findread onlysource verified73/100
Find an agent by name or ID
agent_killdestructivesource verified65/100
Terminate an agent
agent_listread onlysource verified67/100
List all agents
agent_sendwritesource verified72/100
Send a message to another agent
agent_spawnwritesource verified73/100
Spawn a new subagent with optional configuration and depth restrictions
No output schemas documented for any tool. LLMs cannot know what fields to expect from responses, breaking downstream tool chaining and multi-step planning. Tools like agent_list, memory_recall, task_list, image_analyze, and web_fetch return unspecified data structures.
Destructive operations (file_delete, agent_kill, schedule_delete) lack confirmation or dry-run patterns. Agents make mistakes, no protection against accidental deletion of critical agents or files.
shell_exec tool accepts arbitrary shell commands with only undocumented 'taint checking for injection patterns' as defense. Code for taint checking not visible. This is a critical RCE vector, LLMs can be tricked into executing malicious commands. No allowlist, no sandboxing, no parameter validation.
Recommendations
Document output schemas for ALL tools. Specify the JSON structure returned: fields, types, required vs. optional, and what data is safe to pass to downstream tools. Example for agent_list: {type: 'object', properties: {agents: {type: 'array', items: {type: 'object', properties: {id: {type: 'string'}, name: {type: 'string'}, status: {enum: ['running', 'idle', 'failed']}, depth: {type: 'integer'}}}}}}.
Expand descriptions to 150-250 characters and include: (1) What the tool does, (2) When to use it instead of similar tools, (3) Key consequences (if destructive), (4) Common next steps. Example for agent_kill: 'Terminate a running agent by UUID. Kills all child tasks and spawned subagents. Agents cannot be resumed after termination. Use before spawning a replacement to clean up failed subagents or to free memory in long-running systems.'
Add confirmation or dry-run support to destructive tools. For file_delete and agent_kill, require a boolean 'confirm=true' parameter or add a 'dry_run' mode that returns what would be deleted without executing. This prevents accidental destruction.
Secure shell_exec by: (1) Documenting the taint-checking logic (what patterns are blocked?), (2) Adding an explicit allowlist of permitted commands or command prefixes, (3) Running executed commands in a sandboxed subprocess with resource limits (CPU, memory, wall-clock time), (4) Logging all shell invocations with arguments for audit. Better: split shell_exec into smaller domain-specific tools (e.g., compile_code, run_tests, grep_files) so agents don't have raw shell access.
Score history
Overall score trend
First recorded score · v2 rubric
68/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-23
C
68
2026-07-28+
v2
Publish an event
file_copywritesource verified75/100
Copy a file to a new location
file_deletedestructivesource verified70/100
Delete a file
file_listread onlysource verified75/100
List files and directories in a given path
file_mkdirwritesource verified72/100
Create a directory
file_movereversiblesource verified73/100
Move or rename a file
file_readread onlysource verified80/100
Read the contents of a file
file_statread onlysource verified75/100
Get file metadata (size, permissions, timestamps)
file_writewritesource verified75/100
Write or create a file with the given content
image_analyzeread onlysource verified70/100
Analyze an image with computer vision
location_getread onlysource verified65/100
Get current location information
memory_recallread onlysource verified67/100
Search and retrieve memories from agent memory
memory_storewritesource verified70/100
Store information in agent memory
schedule_createwritesource verified78/100
Create a scheduled task (restricted to depth 0)
schedule_deletedestructivesource verified65/100
Delete a scheduled task (restricted to depth 0)
schedule_listread onlysource verified67/100
List all scheduled tasks
shell_execirreversiblesource verified63/100
Execute a shell command with taint checking for injection patterns
task_claimwritesource verified70/100
Claim a task for execution
task_completewritesource verified73/100
Mark a task as complete
task_listread onlysource verified72/100
List all tasks
task_postwritesource verified70/100
Post a new task
web_fetchread onlysource verified77/100
Fetch content from a URL with SSRF protection and taint checking
Descriptions are too generic and brief (average ~80 chars vs. 194-char baseline). They state WHAT tools do but not WHEN to use them or consequences. E.g., 'Terminate an agent' vs. 'Terminate an agent, freeing memory and stopping any running tasks. Killed agents cannot be resumed. Use before spawning replacements or to clean up failed subagents.'
No error handling guidance in any tool schema. LLMs don't know: is this retryable? Should I ask the user to fix it? Is it fatal? Examples: file_read fails on permission denied, should the agent try sudo, ask the user, or retry? web_fetch fails on timeout, retry or give up?
Parameter descriptions are minimal. Many parameters lack context: 'query' for web_search could mean what? Boolean syntax? Natural language? regex? 'status' for task_list, what are valid enum values? Agents guess and fail.
Many tools accept natural-language identifiers (agent name, schedule name) but schema only shows string type with no description of format or resolution behavior. If agent_find('my agent') returns agent_id='a1b2c3', does agent_kill accept 'my agent' or must the user pass 'a1b2c3'? Ambiguity forces extra lookups.
No pagination or result limits documented. Tools like agent_list, task_list, schedule_list, memory_recall could return unbounded results. If a system has 10,000 scheduled tasks, task_list returns all of them, exploding context. No mention of limit/offset/cursor parameters.
Agent management tools (agent_spawn, agent_send, agent_kill) and task tools (task_post, task_claim, task_complete) are underdocumented regarding lifecycle and state. What happens if agent_kill is called while the agent is running a critical task? Does task_complete trigger event_publish? Do events race with task status updates?
No visible audit or permission declaration in tool definitions. Tools like file_write, shell_exec, agent_spawn carry high risk but no indication of what permissions they require or who can call them. No audit trail metadata.
Add 'limit' and 'offset' parameters to all list/search tools (agent_list, task_list, schedule_list, memory_recall, web_search). Document default limits (e.g., 'Default limit=20, max=100; returns total_count for pagination'). Cap unbounded results.
Enhance parameter descriptions with format and constraint hints. Examples: (1) query (web_search): 'A natural-language search query; DuckDuckGo respects boolean operators (AND, OR, NOT). Example: python AND tutorial NOT Java', (2) status (task_list): 'Filter by task state. Valid values: pending (not yet claimed), claimed (assigned but not done), completed (finished). Omit to return all statuses.', (3) cron_expression (schedule_create): 'A standard cron expression (minute hour day month weekday). Example: 0 9 * * 1 runs at 9am every Monday. Invalid expressions return an error.'
Add error handling guidance for every tool. For each, specify: (1) Common failure modes, (2) Whether to retry, (3) How to recover. Example for web_fetch: 'Timeout (after 10s): Retryable. Try again with a simpler request or different URL. | 403 Forbidden: User-fixable. Check API credentials or permissions. | 404 Not Found: Fatal. URL or endpoint does not exist. Verify the URL and try a different search.'
For agent-lookup tools (agent_find, agent_send, agent_kill), clarify whether parameters accept both UUIDs and human-readable names, and document the resolution order. Example: 'The agent_id accepts both: (1) A UUID (e.g., 'a1b2c3d4-e5f6-4a8b-9c0d-e1f2a3b4c5d6'), (2) An agent name (e.g., 'research_bot'). If multiple agents share a name, returns the most recently spawned. To target a specific UUID, pass the full UUID.'
Add permission declarations to tool schemas. Include a new 'permissions' field: {permissions: ['read:files', 'write:files']} for file tools; {permissions: ['spawn:agent', 'kill:agent']} for agent tools; {permissions: ['execute:shell']} for shell_exec. This enables least-privilege agent configurations.
Document idempotency and side effects. Specify for each tool: 'Idempotent: yes/no. If called twice with the same parameters, the result is guaranteed to be identical (e.g., file_read, agent_find). Non-idempotent: Calling twice has cumulative effects (e.g., file_write appends, agent_spawn creates a duplicate). Retrying may produce duplicates, use exponential backoff with jitter.'
For the agent management subsystem (agent_spawn, agent_send, agent_kill, agent_list), document the relationship between agents, tasks, and events. When does an agent enter 'failed' state? What happens to tasks if an agent dies? Does event_publish guarantee delivery? Add a state diagram or sequence diagram in tool descriptions.
Cap unbounded string parameters. Add maxLength constraints: path fields (max 2048), command (max 4096), query (max 1000), message content (max 16384). Document these in parameter descriptions: 'The shell command to execute (max 4096 characters). Very long commands may timeout, break into multiple calls if needed.'
For tools returning structured data (agent_find, task_list, memory_recall, browser_navigate), explicitly document whether the response is an object, an array, or paginated. Example: 'Returns a paginated object: {total_count: integer, limit: integer, offset: integer, items: [...]}'.
Add 'dry_run' mode to high-risk tools. Examples: file_write, file_delete, agent_spawn, schedule_create. dry_run=true shows what would happen without committing changes: 'Would write 1024 bytes to file.txt. File does not exist, so it would be created.'
Document tool composition patterns. For example: 'To send a message to an agent by name, call agent_find(name) to get the UUID, then call agent_send(agent_id, message). To list and filter tasks, call task_list() then filter in-agent rather than re-calling with different status parameters.'
For browser_navigate and image_analyze, document timeout behavior and output format. What does browser_navigate return, raw HTML, markdown, accessibility tree? What format does image_analyze use, structured JSON tags or free text? Make this explicit.