Autonomous red team security scanner with adaptive attack planning, kill chain execution, and AI-driven vulnerability analysis
Mergen is a sophisticated security toolkit with 26 tools covering reconnaissance, exploitation, and attack planning. However, it exhibits pervasive deficiencies in definition quality. While most tools have descriptions (rarely under 20 chars), parameter descriptions are frequently missing or minimal. Input schemas are present but often lack complete type information and descriptions for parameters. Critical security-focused tools like `run_command`, `write_and_exec`, `adaptive_attack`, and `elite_hunt` lack proper constraint documentation and error recovery guidance. The tool composition is reasonable, distinct tools for different attack phases, but parameter documentation is inconsistent across the toolkit. No tool declares permissions, which is a significant oversight for a security-sensitive application. Output schemas are undocumented in tool definitions, forcing LLMs to reason about response structures without explicit guidance.
Fully autonomous attack mode. Combines planning and execution in a loop.
Build a structured Logic Map of the target app: endpoints, params, tech stack, interesting targets. Run this FIRST before any scanner.
HTTP parameter discovery: finds hidden GET/POST parameters in web apps.
Firmware/binary analysis: extracts embedded files, filesystems, and signatures.
Binary security checker: NX, PIE, RELRO, Stack Canary, ASLR, Fortify.
Passive subdomain enumeration via Certificate Transparency logs (crt.sh). No API key required.
Advanced XSS scanner. Pass 'params' list from arjun/katana for targeted scanning instead of blind discovery.
CRITICAL: Destructive tools (`run_command`, `write_and_exec`, `adaptive_attack`, `elite_hunt`) lack confirmation/dry-run patterns. LLMs can execute arbitrary shell commands and exploits without user confirmation, risking catastrophic damage.
CRITICAL: No permission gates or scope declarations on any tool. Tools should declare required permissions (e.g., 'execute:shell', 'exploit:network', 'access:sensitive_data'). Currently, agents have unrestricted access to all capabilities.
HIGH: Output schemas are undocumented. Tools like `get_session_report`, `plan_attack`, `elite_hunt`, and `execute_plan` return complex findings/reports but provide no schema for LLMs to parse results. LLMs must guess response structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Detect auth bypass/IDOR by comparing responses under different auth tokens. Pass token_a, token_b (or no_auth vs token_a).
Fast web path scanner with extension support and smart filtering.
MERGEN 3.0 Elite Hunt — 6-layer autonomous red team pipeline. 41 kill chains, AI playbook selection, correlation engine, dual reporting.
Execute a previously generated attack plan (or generate one on the fly).
Find exploits for a specific service version using SearchSploit. Returns ready-to-run commands.
Get the live status and output of a specific job ID.
Get known vulns, bypasses, and best tools for a detected tech stack. Call before hypothesis generation. tech_stack can be a comma-separated string like 'Laravel,PHP' — multi-stack lookups are supported.
Get a summary of findings, risk scores, and tool outputs for a session.
Kill a currently running background job by ID.
List all available plugins, their descriptions, and installation status.
Generate a comprehensive attack plan for a given target without executing it.
CTF exploit development: generates exploit templates and runs pwntools scripts.
Execute a raw shell command on the server. Returns the output and status code. Use with caution.
Run a single tool directly (e.g. 'nmap', 'whois', 'sqlmap'). Returns a job_id immediately — tool runs in background to avoid MCP timeouts. Use get_job_output(job_id) to poll until status='done'. Options: JSON string '{"level":5}' or dict {"level": 5}.
Record what worked or didn't for future operations. Call after each confirmed finding.
Run an intelligent reconnaissance workflow (ports, services, HTTP).
Extract printable strings from binaries: URLs, credentials, IPs, hashes.
Run vulnerability analysis using Nuclei, Nikto, and SearchSploit. Best run AFTER recon.
Write exploit script to Kali disk and execute a command. Use for custom exploitation: write Python/bash exploit then run it.
HIGH: Parameter descriptions are frequently sparse or missing for optional parameters. Tools like `run_tool` (options param accepts 'JSON string or dict', unclear format), `app_map` (no description for headers param), `dalfox` (workers param lacks range/guidance).
HIGH: No audit trail or logging of tool invocations. Security-sensitive tools (exploitation, reconnaissance, command execution) should log who called what, with which parameters, and results, currently no evidence of this.
MEDIUM: Input validation and error guidance is absent. `run_command` accepts any shell command with no validation. `plan_attack` accepts free-form 'objective' without enum constraint (should be closed set: full_compromise, data_exfiltration, lateral_movement, etc.). Error responses do not guide LLM recovery.
MEDIUM: Parameter type ambiguity. `run_tool` options param documented as 'JSON string or dict', JSON Schema should enforce strict type (string or object, not union). Similarly, `diff_check` id_range is array but no description of element type (int?) or format ([1, 100]?).
MEDIUM: Missing pagination and result limits. Tools like `get_memory` may return unbounded learnings. `app_map` (spider depth 3 = unbounded result set), `dalfox` (worker count has no max). Description should state result limit caps and pagination strategy.
MEDIUM: Tool composition issue: `execute_plan` and `adaptive_attack` both generate AND execute plans. These should be split into separate tools (one generate, one execute) so agents can control the workflow. Currently, LLMs cannot inspect a plan before executing it.
LOW: `write_and_exec` description is minimal ('Write exploit script to Kali disk and execute a command'). Does not explain when to use vs `run_tool` or `run_command`, what happens to the script afterward, or how to debug if execution fails.
LOW: Tool naming ambiguity: `save_learning` vs `get_memory`, unclear relationship. Also, `list_tools` is a meta tool (lists plugins) but appears alongside action tools. Consider renaming to `list_plugins` or moving to a separate namespace.