A comprehensive MCP server providing tools for penetration testing and vulnerability assessment, integrating multiple security tools including Commix, Nmap, Metasploit, Shodan, SQLMap, and others.
The server defines 29 tools with highly inconsistent quality. While all tools have names and descriptions present, the descriptions are often generic (under 20-50 chars), parameter descriptions frequently omit type constraints and valid ranges, input schemas lack proper JSON Schema structure with type definitions, and output schemas are entirely undocumented. Error handling is minimal, most tools return generic try/catch messages without guidance for recovery or classification. The tool set shows significant composition issues: multiple tools do similar things (metasploit_search vs searchsploit_search vs shodan_search), and many lack clear verb-noun naming conventions. Security concerns are severe: filesystem write/delete operations expose destructive capabilities without confirmation patterns, subprocess execution (commix, curl, gobuster, hashcat, nmap, nuclei, sqlmap) invokes external tools with minimal input validation, and no visible permission gates or audit trails. The code shows proper FastMCP framework usage but lacks production-grade tooling patterns for error classification, schema documentation, and permission modeling.
Tools (29)
commixread onlysource verified43/100
Perform a Commix request with the specified command.
commix_helpread onlysource verified40/100
Get the help information for Commix.
curl_requestread onlysource verified45/100
Perform a CURL request with the specified command. Shell operators like `&&` and `||` are not supported and will result in an error.
delete_filedestructivesource verified52/100
Delete a file from the filesystem.
get_cveread onlyauthsource verified53/100
Provides information on the provided CVE ID. Returns JSON.
gobuster_dirread onlysource verified50/100
Perform a Gobuster directory scan on the specified URL using the provided wordlist. Wordlists are the default Kali Linux wordlists in /usr/share/wordlists.
hashcat_dictionaryread onlysource verified52/100
Crack a hash using Hashcat with the specified wordlist.
Output schemas completely undocumented. No tools document what fields they return, their types, or structure. LLMs cannot parse results reliably or plan downstream tool calls.
Input schemas lack proper JSON Schema type definitions. Tools define parameters as simple string/object without explicit 'type' field. Schema score cannot exceed 30 when type definitions are missing.
Subprocess execution tools (commix, gobuster_dir, hashcat_dictionary, nmap_scan, nuclei_scan, sqlmap_*) use os.system() or subprocess.run() with minimal input sanitization. Vulnerable to command injection if attacker controls parameter values. No validation of --batch flags, wordlist paths, or option strings.
Recommendations
Document output schemas for all 29 tools. For each, specify the structure of returned data (JSON object fields, array items, primitive types). Example: 'Returns a JSON object with keys: success (bool), command_output (string), exit_code (int).' This enables LLMs to parse results and chain tools.
Add proper JSON Schema type definitions to all input parameters. Replace {"options": {"type": "string"}} with explicit schema including minLength, maxLength, pattern, or enum constraints. Example: {"options": {"type": "string", "minLength": 1, "maxLength": 1024, "description": "CURL flags (e.g. -X GET). Do not use shell operators like && or ||."}}.
Implement input validation and sanitization for all subprocess execution tools (commix, curl_request, gobuster_dir, hashcat_dictionary, nmap_scan, nuclei_scan, sqlmap_*). Validate wordlist paths exist and are readable, reject shell metacharacters (&&, ||, ;, |), whitelist allowed flags. Return clear validation errors: 'Invalid flag: --recursive. Allowed flags: --url, --data, --batch.'
Add confirmation/dry-run pattern to write_file and delete_file. Example: add a dry_run boolean parameter; return 'Would write X bytes to /path/file.txt (dry-run). Call again with dry_run=false to execute.' Prevents accidental data loss.
Upgrade error handling to distinguish retryable vs fatal errors and provide recovery guidance. Instead of 'Error: Connection refused', return: 'Error: Connection refused (retryable). The target service is unreachable. Wait 30s and retry, or check if the target is online with nmap_scan.' Use structured error objects with code, message, and recovery hints.
Score history
Overall score trend
↓ 3 points across a rubric change (v1 → v2)
49/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
49
2026-07-28+
v2
2026-03-09
D
52
-
v1
hashcat_helpread onlysource verified40/100
Get help information for Hashcat.
list_directoryread onlysource verified52/100
List the contents of a directory on the filesystem.
Search for exploits in SearchSploit using the provided query.
shodan_hostread onlyauthsource verified55/100
Get detailed information about a specific host in Shodan.
shodan_searchread onlyauthsource verified53/100
Search for devices in Shodan using the provided query. This uses the standard Shodan query format, which can include filters like port:80, country:US, os:Windows, asn:as214092.
Destructive filesystem operations (write_file, delete_file) exposed without confirmation pattern or permission gates. delete_file is disabled (enabled=False) but write_file allows arbitrary file overwrite with no safeguards. No dry-run support.
Generic error handling: all tools return 'Error during [operation]: {str(e)}' with no guidance for recovery, classification, or retry logic. LLMs cannot distinguish retryable vs fatal errors or know what to do next.
Multiple tools search for exploits/vulnerabilities with overlapping functionality and no clear differentiation: metasploit_search, searchsploit_search, shodan_search, nuclei_scan, search_cve. LLMs waste reasoning cycles deciding which to use. No canonical naming convention.
Parameter descriptions missing or too generic. Examples: 'options' (commix, curl_request) lacks specification of valid flags; 'module_options' and 'payload_options' (metasploit_exploit) are type 'object' with no schema of expected keys; 'hash_type' (hashcat_dictionary) expects integer with no documentation of valid type codes.
No pagination or result limiting on tools that could return unbounded data (metasploit_search, searchsploit_search, shodan_search, search_cve, gobuster_dir). No page/limit/offset parameters. LLMs could exhaust context window with large result sets.
Tool descriptions lack actionable context: when to use tool, what it returns, prerequisites, or next steps. Examples: 'Perform a Commix request' (no indication this is OS command injection testing); 'List all available payloads' (no indication what fields are in the payload list).
No permission gates or scope declarations. Tools like write_file, metasploit_exploit, and sqlmap_* are high-impact (filesystem modification, network attacks, database enumeration) but require no authorization check or audit trail.
Consolidate overlapping search tools. Choose one canonical search interface (e.g. search_vulnerability with mode parameter: 'cve'|'metasploit'|'searchsploit'|'shodan') or clearly differentiate them in descriptions. Example: 'Use search_cve for CVE IDs by keyword; use metasploit_search to find exploit modules; use shodan_search for internet-wide host discovery.'
Add pagination to metasploit_search, searchsploit_search, shodan_search, search_cve, and gobuster_dir. Add parameters: limit (default 20, max 100) and page (default 1). Return total_results and has_next_page. Document in descriptions: 'Returns up to 20 results per page. Use page parameter to iterate.'
Enrich parameter descriptions with concrete constraints and examples. Example for hash_type: 'Hash algorithm type as integer. Common values: 0=MD5, 100=SHA1, 1400=SHA256, 3200=bcrypt. See hashcat_help() for complete list.' Example for wordlist: 'Path to wordlist file (readable, <10MB). Standard Kali paths: /usr/share/wordlists/rockyou.txt, /usr/share/wordlists/dirb/big.txt.'
Add permission gates to sensitive tools (write_file, delete_file, metasploit_exploit, metasploit_session_interact, sqlmap_*). Check if calling agent/user has required scope (e.g. 'write:filesystem', 'run:exploit'). Return: 'Insufficient permissions: write:filesystem required. Contact admin to grant scope.'
Implement structured audit logging for all tools. Log: timestamp, caller_id, tool_name, parameters_sanitized, result_status, error_details. Make logs available via a log_audit_trail tool so teams can review who ran what and when.
Add examples and WHEN-to-use guidance to tool descriptions. Example: 'commix: Test for OS command injection vulnerabilities. Use when you identify a parameter that executes shell commands (e.g. ping, nslookup). Returns: vulnerabilities found (bool), injection points (array), payloads (array).'
For metasploit_exploit and metasploit_session_interact, document expected module_options and payload_options structures. Provide example JSON: 'module_options: {"RHOSTS": "192.168.1.1", "LHOST": "10.0.0.1"}'. Link to metasploit_info and metasploit_payload_info as discovery tools.
Replace os.system() calls with subprocess.run(check=False, capture_output=True) and validate exit codes. Return both stdout and stderr. Handle timeouts explicitly (set timeout=30 by default, let user override).
Add a dry_run parameter to metasploit_exploit and sqlmap_* tools. Return what would happen without executing. Example: 'Dry-run: Would run exploit unix/webapp/wp_admin_shell_upload with RHOSTS=example.com, LHOST=10.0.0.1. Call again with dry_run=false to execute.'
Separate tool concerns: split curl_request into curl_request (GET/POST) and curl_download (fetch file). Split metasploit_exploit (info only) from metasploit_run_exploit (execute). Reduces accidental misuse.
Document security/risk classification for each tool in MCP tool annotations. Use readOnlyHint, destructiveHint, idempotentHint to signal impact. Example: write_file marked destructiveHint=true; curl_request marked readOnlyHint=true (unless used for exfil).
For tools returning file content (read_file, searchsploit_examine), limit response size. Cap at 64KB; offer a tool to get_next_chunk if content is larger. Prevents context window exhaustion.