Security auditor for AI agent configurations. Scans Claude Code setups for vulnerabilities, misconfigs, and injection risks.
AgentShield exposes 9 tools with basic definitions and input schemas, but lacks depth across multiple quality dimensions. All tools have descriptions (20+ chars), and all have input schemas with typed parameters, which prevents a failed score. However, descriptions are generic and lack LLM-optimization guidance; parameter descriptions are minimal; output schemas are undocumented; error handling is not visible in tool definitions; and no security annotations (readOnlyHint/destructiveHint) are present despite clear risk classifications. The tool set itself is problematic: 'bash', 'network', and 'external_api' are dangerous tools that should not be exposed as-is without substantial guardrails, multi-step confirmation, and dry-run patterns. The whitelist model is architecturally sound, but the tool interface itself needs refinement.
Execute shell commands. DANGER: Can access host system. Only enable with additional containment.
Edit existing files within the sandbox. Validates file exists before modification.
Call external APIs. DANGER: Can make authenticated requests to third-party services.
Pattern-match files within the sandbox. Scoped to sandbox directory only.
List directory contents within the sandbox. Cannot traverse above sandbox root.
Make HTTP requests. DANGER: Can exfiltrate data. Only enable with network policy allowlist.
Read file contents within the sandbox. Cannot access files outside the sandbox boundary.
Output schemas are completely undocumented. No indication of what fields are returned, structure of results, pagination support, or chaining IDs. LLMs cannot plan downstream tool calls without knowing response structure.
Dangerous tools (bash, network, external_api) lack confirmation/dry-run patterns. These tools can cause irreversible damage (command execution, data exfiltration, unauthorized API calls), yet there is no mention of a dry-run mode, confirmation step, or multi-round-trip request (MRTR) pattern to prevent accidental misuse.
Parameter descriptions are minimal and lack constraint guidance. E.g. 'path' param has no format hint, length limit, or example format. 'pattern' for search and glob lacks regex vs glob clarification. 'command' for bash has no safety constraints or allowed command hints. LLMs cannot validate inputs against undocumented constraints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Search file contents within the sandbox using text patterns. Scoped to sandbox directory.
Write file contents within the sandbox. Requires explicit session configuration.
Tool descriptions lack error recovery guidance. No indication of what errors are possible, whether failures are retryable, or what the LLM should do if a path is invalid, a file doesn't exist, or a command fails. Agents will be left without recovery paths.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk classifications. The 'read', 'search', 'list' are safe reads; 'write', 'edit', 'glob' are writes; 'bash', 'network', 'external_api' are destructive. These annotations should be present in tool definitions to guide LLM behavior.
'bash', 'network', and 'external_api' tools should never be exposed without substantial containment, allowlists, and granular permission gates. The descriptions warn about danger but do not enumerate required precautions. These are too risky to expose as-is.
Sandbox/path traversal validation is mentioned in tool descriptions but the actual validation logic in 'src/miniclaw/sandbox.js' is not shown. Cannot verify that path traversal attack mitigation is correctly implemented.