Kali-MCP has critical structural issues: STDIO-only transport caps it at 50 overall, but definition quality itself is significantly below baseline. While 12 tools are explicitly registered with names and descriptions, they share widespread problems: (1) descriptions are vague and generic, failing to guide LLM selection; (2) parameter schemas lack proper typing and constraints; (3) no output schemas documented; (4) error handling is minimal and non-actionable; (5) dangerous command injection vulnerabilities in sanitization logic; (6) no idempotency or retry guidance. The server is designed for penetration testing (a high-risk domain requiring strict controls), yet implements naive input filtering that can be bypassed. Examples: nmap_scan, nikto_scan, dirb_scan, wpscan_scan, sqlmap_scan all have descriptions under 100 chars that don't explain when to use them relative to each other, and parameters lack format constraints or enums. searchsploit_query description is 95 chars with no guidance on result structure. apt_install and git_clone are WRITE operations with no confirmation, dry-run, or undo capability. The sanitize_input() function removes characters via regex but doesn't validate IPs, domains, URLs, or file paths, an attacker can still inject payloads by encoding them or using characters the regex misses.
Install a package using apt package manager.
Search for available packages in apt repositories.
Perform directory brute force scan using DIRB to find hidden web paths.
Clone a git repository to the specified destination.
Update a git repository by pulling latest changes.
List security tools currently installed in the container.
Perform web server vulnerability scan using Nikto.
No output schemas documented. LLMs cannot infer what fields tools return or structure downstream calls. E.g., nmap_scan, nikto_scan, dirb_scan all return opaque formatted strings with no structured field extraction.
Parameter schemas lack type definitions and constraints. Input schemas visible in code show only 'type: string' with minimal descriptions. No enums for fixed options (e.g., nmap options should be constrained), no format validation for IPs/URLs/paths, no min/max bounds on timeouts.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Perform network port scan using nmap with version detection and default scripts.
Perform quick reconnaissance combining nmap and basic vulnerability checks.
Search exploit database using searchsploit for known vulnerabilities.
Perform SQL injection testing using sqlmap on a target URL.
Perform WordPress vulnerability scan using WPScan to enumerate plugins, themes, and users.
Dangerous command injection vulnerability in sanitize_input(). The function removes characters via regex (';|`$()<>') but: (1) attacker can URL-encode payloads, (2) regex is incomplete (missing !, |, *, &, {}, etc.), (3) tool names still passed unsanitized to subprocess.run() array, (4) no validation that target is a valid IP/domain/URL. E.g., sqlmap_scan('http://evil.com?cmd=whoami') passes through.
WRITE operations (apt_install, git_clone, git_pull) lack confirmation, dry-run, or undo capability. An LLM can silently install arbitrary packages or clone malicious repos without user approval. No idempotency hints or rollback guidance.
Descriptions are vague and generic. E.g., 'Perform network port scan using nmap with version detection and default scripts' does not explain when to use nmap_scan vs quick_recon, what output format to expect, or how to interpret results. Baseline is 194 chars; most descriptions here are 50-100 chars with minimal guidance.
Error handling is non-actionable. E.g., when nmap times out, LLM receives 'Command timed out after 300 seconds' with no guidance on retry, whether to reduce scope, or what to do next. No error classification (retryable vs fatal) or recovery steps.
No pagination or result limits documented. quick_recon calls nmap -F but no description of max results or how to handle large networks. searchsploit_query could return hundreds of exploits, no limit, page, or offset parameter.
Parameter descriptions lack format or constraint guidance. E.g., 'options' param in nmap_scan says 'Nmap options (default: '-sV -sC')' but does not explain valid flags, maximum option length, or examples of common queries. 'wordlist' in dirb_scan is a file path but description does not clarify if it's absolute or relative, required format, or validation.
No permission gates or audit trails. Destructive operations (apt_install can break the system, git_clone can consume disk) are not gated by user confirmation, permission checks, or logging of who called what. No scope declarations (e.g., 'requires: admin').
Parameter names are generic and ambiguous. 'options' in nmap_scan could mean command-line flags, configuration file, or timeout settings. 'enumerate' in wpscan_scan is documented as 'p,t,u' but unclear what these abbreviations mean or if other values are valid. 'data' in sqlmap_scan is vague, POST data, query string, cookies?