Professional pentest MCP server with stdio + Streamable HTTP transports, bundled MCP Inspector launcher, bearer auth, SoW-aware reporting, and modern tooling across sniffing/finding/bruteforce/cracking/priv-esc/extraction workflows.
Pentest MCP has 18 tools with inconsistent quality. Most tools have basic descriptions and schemas, but critical gaps exist: (1) Many tool names lack clear action verbs (e.g., 'nmap_scan' is acceptable but 'createClientReport' violates verb_noun convention); (2) Parameter descriptions exist but are minimal (10-40 chars average, well below the 72-char baseline for A+ tools); (3) Output schemas are not documented anywhere in the provided source, the response structure is completely opaque to callers; (4) Error handling is absent, no guidance on retryability, user-fixable errors, or recovery paths; (5) High-risk tools (john_crack, hashcat_crack, hydra_bruteforce) lack explicit acknowledgment of destructive intent or confirmation patterns; (6) The tool 'createClientReport' with scopeMode='ask' hints at elicitation, but the mechanism is not visible in the code provided, inferred rather than explicit. Security tools are exposed without rate limiting, timeout guidance, or audit trail documentation. Overall, this reads like a functional prototype with penetration test semantics, but lacks the production-grade polish of A-/B-tier tools.
Tools (18)
createClientReportwritesource verified55/100
Generate a professional engagement report with Statement of Work elicitation support
ffuf_scanread only50/100
Fuzz HTTP requests using FFUF
getEngagementRecordread onlysource verified60/100
Retrieve a specific engagement record by ID
gobuster_scanread only50/100
Brute force directory and DNS names using Gobuster
hashcat_crackread only50/100
Crack password hashes using Hashcat
httpx_scanread only50/100
Probe HTTP services using Httpx (ProjectDiscovery)
hydra_bruteforcewrite50/100
Perform brute force attacks on network services using Hydra
Output schemas are completely undocumented. Callers cannot know what fields to expect from any tool response. This forces LLMs to guess response structures, breaks tool chaining, and risks context loss between turns.
Parameter descriptions are minimal (10-50 chars) and lack actionable detail. E.g., 'args' for nmap_scan is 'Array of nmap command-line arguments', no guidance on valid options, format, safety constraints, or examples of what a user might pass.
Document output schemas for every tool. For nmap_scan, specify: {scan_id: string, status: 'initializing'|'scanning'|'complete'|'failed', ports: [{port: number, state: string, service: string}], host: string}. Use this format in the tool's definition or a separate schema file.
Expand parameter descriptions to 50-100 chars with actionable constraints. E.g., 'args: Array of nmap options (e.g. -sS, -p 1-1000, -O). Avoid -oX as output is handled internally. Max 10 arguments.'
Rename camelCase tools to snake_case for consistency: 'createClientReport' → 'create_client_report', 'listEngagementRecords' → 'list_engagement_records', 'getEngagementRecord' → 'get_engagement_record', 'setUserMode' → 'set_user_mode'.
Add destructiveHint: true annotation to high-risk tools (john_crack, hashcat_crack, hydra_bruteforce, tcpdump_capture) in the tool definition. This signals to LLMs that these tools modify state or have side effects.
Add explicit error guidance to tool descriptions. E.g., 'If nmap_scan returns timeout error, try reducing target scope or increasing timeout parameter. If denied-by-firewall, check network connectivity.'
Add min/max constraints to numeric parameters. E.g., 'timeout: number (1-3600 seconds, default 300)', 'threads: number (1-64, default 4)', 'port: number (1-65535)'.
Sanitize the 'args' parameter in nmap_scan. Instead of free-form strings, accept structured parameters (scanType: 'syn'|'ack'|'udp', ports: string like '1-1000' or '80,443', timing: 'paranoid'|'sneaky'|'polite'|'normal'|'aggressive'). This prevents command injection.
Score history
Overall score trend
↑ 49 points across a rubric change (v1 → v2)
49/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
49
2026-07-28+
v2
2026-03-09
F
0
-
v1
read onlysource verified60/100
List all engagement records created during the session
nikto_scanread only50/100
Scan web servers for vulnerabilities using Nikto
nmap_cancel_scanwrite50/100
Cancel an active nmap scan
nmap_scanread only50/100
Execute nmap network scan with customizable options and targets
nmap_scan_progressread only50/100
Get current progress of an active nmap scan
nuclei_scanread only50/100
Scan for vulnerabilities using Nuclei (ProjectDiscovery)
setUserModewrite50/100
Set the user mode (student or professional) for the session
sqlmap_scanread only50/100
Scan for SQL injection vulnerabilities using SQLMap
subfinder_scanread only50/100
Enumerate subdomains using Subfinder (ProjectDiscovery)
createClientReport violates verb_noun naming convention. Should be 'create_client_report'. Current name uses camelCase which is inconsistent with the rest of the tool suite (all other tools use snake_case).
Destructive/high-risk tools (john_crack, hashcat_crack, hydra_bruteforce, tcpdump_capture) are not marked with destructiveHint or similar annotation. LLMs cannot distinguish safe read-only tools from potentially dangerous ones. No confirmation pattern or dry-run option documented.
No error handling documentation. Tools lack guidance on retryability, user-fixable errors, or recovery paths. E.g., if john_crack times out, is it retryable? If sqlmap_scan fails, should the user try a different URL format?
Numeric parameters (timeout, packetCount, threads, port) lack min/max bounds in descriptions. An LLM could pass packetCount=999999 or timeout=1000000, causing resource exhaustion or hangs.
The 'args' parameter in nmap_scan (and similar command-line tools) is a free-form array of strings. This is a command-injection vector, LLMs can be tricked into passing malicious nmap options. No input validation or sanitization documented.
createClientReport has scopeMode='ask' but the elicitation mechanism is not visible in the code snippet. If this is a placeholder or incomplete implementation, it should be documented or removed. Callers cannot know how to interact with the 'ask' mode.
No pagination, result limits, or cursor documentation for list-like tools (listEngagementRecords, httpx_scan, subfinder_scan). An LLM could request 100k results, blowing the context window.
listEngagementRecordshttpx_scansubfinder_scan
Implement and document the elicitation flow for createClientReport's scopeMode='ask'. Either remove the feature if incomplete, or add a clear description of how the tool interacts with the LLM to gather scope details.
Add pagination to listEngagementRecords: support limit (1-100, default 20) and offset (0-based). Return {records: [...], total: number, hasMore: boolean}.
Add rate-limit guidance. E.g., 'hydra_bruteforce respects service rate limits but may trigger IDS alerts. Recommended: max 1 concurrent hydra_bruteforce call per team. Implement client-side backoff on 429 responses if the server implements rate limiting.'
Document timeout and resource constraints for long-running tools (john_crack, hashcat_crack). E.g., 'Default timeout 300s; if exceeded, returns partial results with best_cracked field. Server may kill process if CPU > 80% or memory > 2GB.'
For tools returning multiple results (httpx_scan, subfinder_scan, nuclei_scan), include chaining IDs in responses. E.g., if nuclei_scan returns vulnerabilities, include target_url in each result so agent can drill down without re-lookup.
Add security note to all tools: 'This MCP server is intended for authorized security testing only. Unauthorized scanning, cracking, or fuzzing is illegal. Ensure you have explicit written permission from the target owner before running any tool.'
Implement idempotency for tools where possible. E.g., nmap_scan with the same target and args should return the same scan_id if already in-flight, not spawn a duplicate scan. Document this behavior.
Add example invocations in descriptions (without concrete hostnames). E.g., 'sqlmap_scan: url: http://target-app.local/login, method: POST, data: username=admin&password=FUZZ, dbs: true' (with FUZZ or similar placeholder).