AI agent security scanner and npm audit for MCP servers, Claude Code, Cursor, and Windsurf. Find prompt injection, hallucinated packages, secrets, unsafe tools, and vulnerable code.
Server provides 22 security-focused tools with generally clear naming conventions and detailed descriptions. Most tools follow verb_noun patterns (scan_*, check_*, list_*, get_*) appropriately. Descriptions are comprehensive (100-300 chars typical), explaining WHAT the tool does and WHEN to use it. However, critical gaps exist: (1) Input schemas visible in specification lack complete type information for all parameters, several enum parameters are correctly declared but some descriptions lack clarity on constraints; (2) Output schemas are NOT documented anywhere in the provided source, LLMs cannot predict what fields to expect from calls; (3) No error handling patterns visible (no examples of BLOCK/WARN/LOG/ALLOW being used with guidance); (4) No pagination parameters despite tools like scan_project and scan_mcp_server likely returning large result sets; (5) Tool descriptions mention verbosity levels effectively, but no guidance on output field mapping for downstream composition.
Check if a package name is legitimate or potentially hallucinated (AI-invented)
Alias for scanner_health (deprecated, use scanner_health instead)
Evaluate a project against compliance frameworks (SOC2-technical, GDPR-technical, AIUC-1). Collects evidence from code scans, SBOM, vulnerability checks, and hallucination detection, then evaluates controls. Optionally saves timestamped evidence bundle.
Scan a file and return fixes. Use verbosity='minimal' for summary only, 'compact' (default) for fix list, 'full' for complete fixed file content.
Look up compliance controls with evaluation criteria. Supports multiple frameworks: aiuc-1 (default), soc2-technical, gdpr-technical. Filter by domain, control IDs, or OWASP LLM tags.
List statistics about loaded package lists for hallucination detection
No output schemas documented. LLMs cannot predict what fields tools return or plan downstream chaining. For example, scan_security returns findings in some format, but no schema exists to show if results include 'severity', 'line_number', 'fix', 'cwe_id', etc. This violates the pattern:response-shaper requirement.
No pagination support visible in tools that return lists (list_security_rules, list_package_stats, scan_project results, sbom_scan_vulnerabilities). Large result sets will exceed token budgets. Tools should accept limit/offset/cursor parameters and return total_count or next_cursor.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 71 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
List all available security fix templates and their descriptions
Scan a CycloneDX SBOM for hallucinated (AI-invented) packages across all components. Checks against real package registries.
Compare two SBOM files and report added, removed, and modified dependencies. Useful for tracking supply chain changes across builds.
Generate a human-readable security report from a CycloneDX SBOM. Includes executive summary, vulnerability tables, hallucination warnings, and compliance evidence.
Generate a CycloneDX v1.5 SBOM for a project. Discovers all dependencies (direct + transitive) across npm, Python, Go, Java, Rust, PHP, Ruby ecosystems. Includes license and vulnerability data.
Scan a CycloneDX SBOM for known vulnerabilities across all components. Returns vulnerability list with severity, remediation details, and affected versions.
Pre-execution security check for agent actions (bash, file_write, file_read, http_request, file_delete, cron, process_spawn, git, docker). Returns ALLOW/WARN/BLOCK.
Scan a prompt for malicious intent. Returns BLOCK/WARN/LOG/ALLOW. Use verbosity='minimal' for action only, 'compact' (default) for findings, 'full' for audit details.
Scan git diff for new security vulnerabilities. Only reports issues on changed lines. Use for PR reviews.
Scan an MCP server's source code for security vulnerabilities: overly broad permissions, missing input validation, data exfiltration, insecure patterns. Returns grade (A-F) and recommendations.
Scan code for package imports and check for hallucinated (AI-invented) packages. Use verbosity='minimal' for counts, 'compact' (default) for flagged packages, 'full' for all details.
Scan an entire directory for security vulnerabilities with .gitignore support and security grading. Use verbosity='minimal' for grade + counts, 'compact' (default) for top issues, 'full' for all details.
Scan a file for security vulnerabilities. Use verbosity='minimal' for counts only (~50 tokens), 'compact' (default) for actionable info (~200 tokens), 'full' for complete metadata.
Deep security scan of an OpenClaw skill. Multi-layer analysis: prompt injection detection, code analysis (AST+taint), ClawHavoc malware signatures, package supply chain verification, rug pull detection. Returns security grade A-F with detailed findings.
Check plugin health: engine status, daemon status, package data availability
Score findings using OWASP AIVSS v2. Accepts any scanner output or raw findings JSON. Returns per-finding AIVSS scores (0-10) and aggregate posture. Use verbosity='minimal' for posture only, 'compact' (default) for scores, 'full' for all metrics.
Error handling not documented. Tools like scan_agent_action return 'ALLOW/WARN/BLOCK' but no guidance on what LLMs should do on BLOCK or WARN. Should include recovery suggestions: 'Action BLOCKED due to unsafe shell expansion. Alternatives: ask_user_confirm() or rewrite_command()'. Current spec lacks actionable error classification.
clawproof_health is deprecated (description states 'deprecated, use scanner_health instead'). Deprecated aliases should not exist as separate tools, they confuse LLM tool selection and dilute signal. Remove clawproof_health entirely or implement it as a redirect within scanner_health.
Parameter descriptions lack explicit constraints. For example, check_package accepts 'ecosystem' as a string but descriptions like 'Package ecosystem (npm, pypi, etc.)' don't say if the list is exhaustive. Should declare: ecosystem must be one of: npm, pypi, maven, etc. (enum or explicit list).
No confirmation/dry-run pattern for potentially destructive operations. Tools like evaluate_compliance with save_evidence=true and scan_mcp_server with update_baseline=true modify server state (.mcp-security-baseline.json). Should support a dry-run parameter or require explicit confirmation to prevent accidental baseline overwrite.
Tool names do not clearly distinguish between discovery (e.g., list_security_rules) and action tools (e.g., fix_security). However, composition is somewhat unclear: after scan_security finds issues, fix_security should clearly return a fixed file or patch. No documentation shows how results chain together.