Pentesting MCP Server: intelligent abstraction layer over security tools
Tengu is a pentesting-focused MCP server with 7 tools covering CVE lookups, hash analysis, and social engineering payloads. While all tools have descriptions and input schemas are present, quality is uneven. Descriptions are adequate but lack the clarity and actionability needed for reliable LLM tool selection. Parameters are described but some lack detail on constraints and format requirements. Output schemas are not documented in the source code provided. Error handling and recovery guidance are absent. The tool set spans legitimate security research (CVE lookup, hash identification) and high-risk operations (phishing, payload generation), but the definitions do not consistently clarify prerequisites, permission requirements, or reversibility. Critical issues: (1) No documented output schemas for any tool, LLMs cannot plan downstream operations or structure responses; (2) Descriptions mention external dependencies (seautomate, John the Ripper, Hashcat) but do not state whether these are optional or required, risking silent failures; (3) Error handling and actionable recovery messages are missing; (4) Security-sensitive tools (set_credential_harvester, set_payload_generator) lack explicit permission/authorization checks in definitions; (5) Parameter constraints are mentioned in descriptions but not enforced via schema enums or patterns. Naming is generally clear (verb_noun style) and follows conventions, but some tools blur responsibility boundaries (e.g., hash_crack combines identification, wordlist selection, and tool preference in one call).
Fetch complete details for a specific CVE from NVD and CVE.org. Returns CVSS scores (v2/v3.1/v4.0), CWE mappings, affected products, references, and cross-references to known exploits.
Search CVEs by keyword, product, CPE, or severity. Queries the NVD database for matching CVEs. Results are cached locally for 24 hours to respect API rate limits.
Attempt to crack a hash using a dictionary attack. Uses John the Ripper or Hashcat to perform a wordlist-based attack against the provided hash value.
Identify the algorithm used to produce a hash value. Uses pattern matching to determine the likely hash type(s) based on length, character set, and structural patterns.
Clone a website and capture credentials submitted via the phishing page. Uses SET's Website Attack Vectors → Credential Harvester → Site Cloner module via seautomate.
Generate a social engineering payload for use in authorized campaigns. Uses SET's "Create a Payload and Listener" module via seautomate to generate a payload that, when executed by a target, will establish a reverse connection to the operator's listener.
No documented output schemas for any tool. LLMs cannot infer what fields are returned, preventing proper tool chaining and response parsing. E.g., cve_lookup returns CVSS scores and CWE mappings but the response structure is not documented.
External tool dependencies (John the Ripper, Hashcat, seautomate) are mentioned in descriptions but not documented as hard or soft requirements. LLMs cannot predict silent failures when tools are unavailable. hash_crack will fail if neither john nor hashcat is installed, but this is not stated clearly.
Parameter constraints are informal or missing. cve_search accepts 'severity' values (LOW, MEDIUM, HIGH, CRITICAL) but these are mentioned in the description without an enum in the schema. max_results has a documented max of 100 but no minimum is stated. No validation examples or error messages are provided.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Generate a QR code pointing to a malicious URL for physical social engineering. Uses SET's QRCode Generator Attack Vector via seautomate.
No error handling guidance or recovery paths. Tools have risk classifications (READ_ONLY, WRITE, DESTRUCTIVE) but no documentation of what errors can occur, how to detect them, or what the LLM should do next (retry, ask user, abort). hash_crack may time out or find no match, but no distinction is made or recovery suggested.
Security-sensitive and DESTRUCTIVE tools (set_credential_harvester, set_payload_generator) lack explicit authorization checks and permission declarations in their definitions. No scopes (e.g., 'pentest:social_engineering', 'auth:credential_capture') are defined. LLMs cannot verify authorization before invoking these tools.
Descriptions are present but often lack actionable context. cve_search mentions that 'Results are cached locally for 24 hours' (useful) but does not explain when to use cve_search vs cve_lookup, or how to interpret CVSS severity for prioritization. Descriptions should answer: What does it do? When to use it? What are the prerequisites?
hash_crack combines multiple concerns: tool selection (john vs hashcat vs auto), wordlist provision, type specification, and timeout override. This should be split or strongly documented with parameter dependencies to clarify mutual exclusivity and defaults.
Some parameter names are vague or underscore important distinctions. set_credential_harvester has 'lhost' and 'listen_port' but does not clarify whether lhost must match the attacker's listening IP or if it can differ. Parameter naming should be self-documenting (e.g., 'attacker_ip', 'listener_port').
set_credential_harvester and set_qrcode_attack require URLs to be in a tengu.toml allowlist. This is documented in descriptions but there is no tool to query the allowlist or validate a URL before calling these tools. LLMs may repeatedly call with forbidden URLs and hit errors without understanding why.