MCP server for executing penetration testing tools in a Kali Linux environment. Provides shell command execution, network scanning, vulnerability assessment, and exploitation capabilities through an HTTP API with FastAPI and integrates with LangChain-based AI agents.
The server exposes a single tool 'shell' with moderate schema definition but critical security and usability gaps. Tool naming is generic ('shell' lacks verb prefix per pattern:tool); description exists but lacks actionable usage guidance. Input schema is present with types but has semantic issues. No output schema documentation visible. Critical issue: destructive shell execution without permission gates, confirmation patterns, or comprehensive error handling. The tool accepts arbitrary shell commands without sanitization, creating command injection and path traversal risks. Error handling returns structured output but does not guide recovery or categorize errors (per pattern:error-classification and pattern:recovery-guide). No tool annotations present despite destructive risk classification.
Execute shell commands in the Kali container. Use for file operations (echo, cat, touch, rm, ls) and other CLI tasks. Examples: 'echo hash > /tmp/hash.txt', 'cat /tmp/file.txt', 'ls -la /tmp'
Tool name 'shell' does not follow verb_noun convention. Should be 'execute_shell_command' or similar to clarify intent (pattern:tool baseline: 90% of A+ tools start with action verb).
No input schema documentation in source. While ShellInput model is referenced (command: string, timeout: integer), the actual schema validation logic and constraints (min/max timeout, command format restrictions) are not visible in provided code. Cannot verify parameter descriptions match JSON Schema.
Tool description (118 chars) lacks critical security warnings and actionable recovery guidance. Does not state that the tool is destructive, does not explain prerequisites, does not warn about command injection risks, and provides examples that could be misused (e.g., 'rm' command could be dangerous if chained).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 43 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 31 | - | v1 |
Output schema (ShellOutput) is not documented in source. No visible definition of return type structure, field descriptions, or what fields are guaranteed vs optional. LLM cannot predict downstream data availability.
No permission gates or destructive operation confirmation. Tool executes arbitrary shell commands without verifying caller authority or requiring confirmation (pattern:permission-gate, pattern:confirmation-request). Agents can delete files, modify system state, or exfiltrate data without safeguards.
No input sanitization against command injection or path traversal. While cwd='/tmp' limits scope, shell=True with user-provided command is inherently unsafe. Agents or LLMs can be tricked into passing malicious payloads (e.g., 'cat /etc/passwd; rm -rf /'). No validation of command structure, no blocklist of dangerous patterns.
Error handling is not categorized. Timeout and generic exceptions return structured output but do not classify errors as retryable vs user-fixable vs fatal (pattern:error-classification). LLM cannot determine appropriate recovery strategy.
No tool annotations present. Tool marked as 'DESTRUCTIVE' in metadata but no destructiveHint annotation in schema (MCP 2026-07-28 spec). Agents may not recognize execution risk.
Timeout parameter lacks constraints. No min/max bounds specified. LLM could pass negative, zero, or excessively large timeout values. Should document valid range (e.g., 1-3600 seconds).
No audit logging of command execution parameters. While basic logging exists (logger.info), no secure audit trail records who requested what command, when, with full context needed for compliance and incident response (pattern:audit-trail).