An MCP server for AI agents to use during red teaming exercises, providing tools for network reconnaissance, vulnerability scanning, exploitation, and post-exploitation activities
This red team MCP server exhibits significant quality gaps across naming, descriptions, and schema completeness. While tool names are reasonably verb-based (resolve_, port_, enumerate_, get_, search_, domain_), most lack precise descriptions of when to use them versus similar tools. Parameter descriptions are inconsistently present: some tools (port_scan, enumerate_vulnerabilities) have well-documented parameters with examples and constraints, while others (get_banner, resolve_hostname_to_ip) lack parameter-level descriptions. Output schemas are not documented anywhere, no tool declares what fields the response contains, forcing LLMs to guess at response structure. Error handling is minimal; tools return JSON with 'success' flags but provide no recovery guidance. The server mixes responsibilities (scan_nuclei_stream appears redundant with enumerate_vulnerabilities) and lacks clear composition patterns. Security concerns are present but not addressed in tool definitions (no permission declarations, no audit trails). Per-tool scores range from 35 - 58, averaging 42.
DOMAIN DISCOVERY: Enumerate subdomains from a top-level domain using subfinder. This tool discovers subdomains for a given domain and resolves them to IP addresses.
ENUMERATE VULNERABILITIES: Find security issues and CVEs using nuclei scanner with streaming output. This tool ENUMERATES VULNERABILITIES, MISCONFIGURATIONS, and SECURITY ISSUES. It does NOT find open ports - use port_scan for port discovery. Provides real-time progress updates during scanning.
Use host and port and retrieve banner information.
Get all finished scan results from database.
NETWORK PORT SCANNING: Discover open ports on hosts using masscan. This tool performs PORT DISCOVERY to find open TCP/UDP ports on target hosts. It does NOT scan for vulnerabilities - use vulnerability_scan for that.
Resolve a hostname to an IP address with timeout.
No output schemas documented. Tools return JSON responses (resolve_hostname_to_ip returns {success, hostname, ip_address, message}; port_scan returns List[str]; enumerate_vulnerabilities returns []) but LLMs cannot anticipate field names, types, or structure. This forces agents to make assumptions about response shape and breaks downstream tool composition.
Redundant tools with unclear distinctions. enumerate_vulnerabilities and scan_nuclei_stream both scan hosts for vulnerabilities using nuclei. The LLM cannot determine which to call or whether they perform the same action. Violates composition principle of one tool = one responsibility.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Scan a host for vulnerabilities using nuclei scanner with streaming output. Runs Nuclei with -j and -vv, ensuring at least one status update is yielded.
Search scan results by hostname, IP address, port, or service. This allows flexible searching across all scan results to find: - All ports open on a specific host - All hosts running a specific service - All instances of a specific port across hosts
Search vulnerability scan results by various criteria. This allows searching across all vulnerability scan results to find: - All vulnerabilities on a specific host - All high/critical severity issues - All instances of a specific vulnerability type
Parameter descriptions missing or minimal. get_banner has 3 parameters (ip, port, timeout) but only ip and port have non-trivial descriptions (timeout lacks explanation of behavior). resolve_hostname_to_ip has only a hostname parameter with a brief description. Many parameters lack examples, ranges, or format constraints that would guide LLM input.
Tool selection ambiguity. Multiple search_* and get_* tools (search_scan_results, search_vulnerability_results, get_finished_scan_results) lack clear guidance on when to use each. Descriptions do not explain: what is the difference? When should I call search_scan_results vs search_vulnerability_results? This mirrors the pattern:name-clarity issue.
Error handling does not guide recovery. resolve_hostname_to_ip returns JSON with 'success: false' and a generic 'message' field. LLMs cannot determine: should I retry? Is this a bad hostname or a transient DNS failure? What should I try next? Lacks actionable error classification.
No permission declarations or security scope documentation. Tools perform sensitive red team operations (port scanning, vulnerability enumeration, banner grabbing) but lack declared scopes (read:network, execute:scan, etc.). No audit trail pattern implemented. Violates least-privilege and compliance requirements.
Response field naming inconsistency. resolve_hostname_to_ip returns {ip_address, ...} but other tools return bare IP strings or lists. When composing tools, an agent cannot extract a common 'ip' or 'ip_address' field reliably. Breaks tool-chain patterns.
Pagination and result limits not enforced. search_scan_results and search_vulnerability_results accept limit (default 20) but no explicit maximum is enforced in descriptions. If an agent requests limit=10000, will it blow the context window? No guidance on total_count or next_cursor for pagination.
Tool descriptions contain examples in dangerous contexts. port_scan description says 'Find open ports: target=192.168.1.1', LLMs reuse example IPs literally. enumerate_vulnerabilities shows example 'google.com', if agent is constrained to internal networks, it may try to scan Google's external IP.
No idempotency guarantees. Tools like port_scan, enumerate_vulnerabilities, domain_discovery perform network operations. If an agent retries due to a transient timeout, will duplicate scans be initiated? No idempotence declaration or request deduplication mentioned.