MCP server for the cybersec-toolkit — query 670+ tools, get recommendations, execute safely
The server defines 15 tools with generally clear naming (verb_noun pattern) and non-empty descriptions. However, several critical gaps limit production readiness: (1) Input schemas are not visible in the provided source code, only tool names, descriptions, and parameter outlines are shown. Cannot verify actual JSON Schema definitions, type constraints, or enum declarations. (2) Output schemas are completely undocumented, no tool describes what fields it returns or the structure agents should expect. (3) Parameter descriptions exist but are inconsistent in depth. Some params lack format/constraint details (e.g., 'tool_name' in check_installed has no guidance on valid names). (4) Error handling is mentioned in descriptions (e.g., 'Uses multiple detection strategies') but no error recovery patterns are visible (e.g., 'If tool_name not found, try list_tools to discover valid names'). (5) Two tools (run_script, run_tool_remote) represent high-risk operations but lack confirmation/dry-run patterns or explicit risk-classification guidance in descriptions. (6) Field naming consistency issues: 'tool_name' vs 'host_name' vs 'host' (manage_remote_hosts uses both 'host_name' and 'hostname') creates mapping ambiguity.
Retrieve the audit log of all tool executions and MCP server operations. Shows what tools were run, when, and what the results were.
Build a guided assessment plan for a security testing scenario. Helps structure the testing workflow with recommended phases, tools, and validation checkpoints.
Check if a specific cybersecurity tool is installed on the system. Uses multiple detection strategies: .versions tracking, PATH lookup, pipx binary name fallback, /opt directory check, and docker image check. When host is provided, checks installation on the remote host via SSH using 'which <binary>'.
Look up information about a specific CVE (Common Vulnerabilities and Exposures) identifier and get tool recommendations for analyzing or exploiting it.
Get detailed information about a specific tool from the registry, including its module, install method, homepage, and current installation status on the local machine.
List available installation profiles (tool collections optimized for specific security roles and workflows).
No output schemas documented for any tool. LLMs cannot plan downstream calls or structure parsing without knowing return types. E.g., does list_tools return {tools: [...], total: int, filters_available: [...]} or just an array? Does check_installed return {installed: bool} or {status: string, version?: string}?
Input schemas not visible in source code. Only parameter names and descriptions are provided; actual JSON Schema definitions (with types, formats, constraints, enums) are not shown. Cannot verify schema compliance or validate that LLMs receive machine-parseable constraints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
List and filter the 670+ cybersecurity tools in the registry. Returns the tools drawn from tools_config.json — the same registry the installer and the advisors share — with the total count, the filters still available, and one entry per tool. Combine the filters to scope the list: module="web" for web tools, method="pipx" for Python-packaged tools, installed_only=True for only what is on this host. Start here to discover what exists before check_installed or get_tool_info.
Configure and manage SSH remote host connections for running tools on remote systems. Supports SSH key-based authentication with configurable ports and usernames.
Get tool installation recommendations based on your use case, skill level, and available disk space.
Execute a safe, multi-step pipeline of shell commands without spawning an interactive shell. Each step is validated: no pipes to shell, no redirects, no expansion, no variables. Useful for chaining tool outputs together safely.
Execute a Python or Bash script for custom logic that tools and pipelines cannot express. Disabled by default (CYBERSEC_MCP_ALLOW_SCRIPTS=0). Enabling this is an explicit full-code-execution opt-in: scripts inherit the MCP server user's filesystem and network access and are not constrained by CYBERSEC_MCP_ALLOW_EXTERNAL.
Execute a single installed cybersecurity tool by name with arguments. Runs the tool as-is and returns its output. For pipeline safety, use run_pipeline for multiple steps. For custom logic, use run_script (if enabled).
Execute a tool on a remote host via SSH. Requires the host to be configured via manage_remote_hosts. Uses the same tool registry and execution safety as run_tool, but operates on a remote machine.
Get tool recommendations for a bug bounty hunting scenario based on the target type, scope, and vulnerability class.
Get tool recommendations for a CTF (Capture The Flag) challenge based on the challenge category.
Parameter naming inconsistency in manage_remote_hosts: uses both 'host_name' (for action identifier) and 'hostname' (for DNS/IP) to mean the same thing. This creates mapping ambiguity and forces LLMs to reason about which param means what. Should use 'host_id' (identifier) and 'host_address' (IP/DNS), or rename consistently.
High-risk tools (run_script, run_tool, run_tool_remote) lack confirmation/dry-run patterns. run_script is IRREVERSIBLE (full code execution) but offers no way to preview what will run or confirm before executing. run_tool_remote offers no dry-run. This invites catastrophic mistakes.
Error recovery patterns are absent. Descriptions mention what tools do but do not guide LLMs on how to recover from failures. E.g., check_installed has no 'If tool_name invalid, call list_tools first' hint. get_tool_info has no 'If tool not found, try list_tools with filters' guidance. This forces agents to guess or fail.
Parameter format/constraint descriptions are inconsistent. Some params (like 'cve_id' in get_cve_info) show example formats ('CVE-2024-1234') but provide no regex, enum, or validation rules in the description text. LLMs cannot parse examples as constraints, they need explicit enum/pattern/range declarations.
Parameter bounds are missing or unspecified. E.g., audit_server's 'limit' param lacks min/max (should be 1-1000?). recommend_install's 'disk_space_gb' is integer but has no range. This allows LLMs to pass invalid values (e.g., limit=999999 or disk_space_gb=-1).
Tool composition chain risks: list_tools returns tools with a 'tool_name' field, but some callers expect different identifiers. E.g., run_tool accepts 'tool_name', but manage_remote_hosts uses 'host_name' for a different concept. No documentation clarifies which tools produce IDs compatible with downstream tools.