Cybersecurity AI assistant with integrated MCP servers for pentesting, OSINT, and Metasploit automation. Provides chat interface with pentesting report generation, multi-chat management, and Metasploit framework integration.
Unburden exposes 8 Metasploit-focused pentesting tools with critical definition quality gaps. All tools have descriptions and parameter schemas present, but descriptions are generic and lack LLM-optimization guidance. Parameter schemas use proper JSON Schema types, but descriptions are minimal and lack actionable constraints. Most critically, tools lack error handling guidance, recovery paths, and security-scoped permission declarations. The tools are dangerously powerful (run_exploit, send_session_command) with destructive/write risk classifications, yet lack confirmation patterns, dry-run support, or audit trail documentation. Naming is reasonably clear (verb_noun pattern mostly followed), but descriptions do not explain WHEN to use each tool vs. similar ones, WHAT happens on failure, or HOW to recover. This is a high-risk tool suite for autonomous agents without strong guardrails.
Attaches to an active Meterpreter session interactively and returns a tmux session name for manual interaction
Checks the connection status to the Metasploit RPC server and returns version information
Generates a standalone Meterpreter payload and saves it to the system
Kills (terminates) a specific Meterpreter session
Lists all active Meterpreter sessions and returns their information
Executes a Metasploit exploit module with specified options and waits for a Meterpreter session to be established
Searches the Metasploit database for exploits matching a keyword or criteria
Destructive tools (run_exploit, send_session_command, kill_session) lack confirmation/dry-run patterns and do not document irreversibility. Agents may execute destructive commands without user consent.
Parameter descriptions lack validation rules, ranges, constraints, or enum values. Free-form string parameters (exploit_path, payload, command, format) invite hallucinated invalid values. LLMs cannot infer valid inputs from vague descriptions like 'Payload type (e.g., ...)' without formalized enums.
Output schemas for most tools are not documented in the source. Responses likely return raw Metasploit JSON structures without field-level description. Agents cannot reliably chain tool calls or extract required downstream IDs without documented output schemas.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 40 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 25 | - | v1 |
Sends a command to an active Meterpreter session and returns the output. Automatically handles reverse shell entry if command contains shell keywords.
Error handling and recovery guidance entirely absent. Tools do not document what errors can occur (e.g., session_id not found, exploit failed, connection timeout), how to classify them (retryable vs. fatal), or suggest recovery actions. Agents hitting errors have no guidance on what to do next.
No audit trail, permission gate, or scope declaration. These tools execute arbitrary Metasploit commands (exploits, shell commands, payload generation) with no documented permission checks, logging requirements, or least-privilege guidance. Agents can execute any Metasploit action without authorization verification.
Tool descriptions do not explain WHEN to use one tool vs. similar ones (e.g., send_session_command vs. attach_session_interactive; generate_payload vs. manually invoking msfvenom). LLMs waste reasoning cycles on disambiguation.
Parameters like 'options' (object type in run_exploit) and 'command' (string in send_session_command) lack structure and validation. 'options' is described as 'Dictionary of exploit options' with only a vague example (RHOSTS, LHOST, LPORT). LLMs cannot infer which options are valid, required, or mutually exclusive.
Descriptions contain examples of valid values (e.g., 'ms17_010', 'smb', 'windows/x64/meterpreter/reverse_tcp', 'exe', 'dll', 'elf') that should be formalized as enums or regex patterns. LLMs may treat examples as exhaustive and hallucinate variants ('windows/x86/meterpreter/reverse_https' when only x64 is available).
No documented rate limiting, timeout, or resource constraints. Long-running tools like run_exploit and send_session_command may hang. An agent in a retry loop could spawn unlimited Metasploit commands without guards.
No pagination, result limits, or filtering for discovery tools. search_exploit may return hundreds of exploits; list_sessions may return hundreds of sessions. Returning all results wastes tokens and risks context window exhaustion.