MCPSafetyScanner - An MCP server for security auditing and vulnerability scanning
This is a deliberately vulnerable MCP server designed to expose security anti-patterns, not a production tool. The server implements 8 tools across two categories: (1) VulnerableMCP tools (read_file, write_file, execute_shell_command, get_environment_variables) that provide unrestricted filesystem and OS access; (2) mcpsafety scanner tools (list_available_tools, call_a_tool, get_tools, call_tool) that allow remote tool invocation on arbitrary MCP servers. Tool descriptions are minimal (10-60 chars), lack prerequisites or usage guidance. Input schemas are skeletal, most parameters are bare strings with single-line descriptions. No output schemas documented. Error handling is absent, tools return raw exception strings. No parameter validation, no enum constraints, no rate limiting, no security filtering. The 'execute_shell_command' tool exemplifies the problem: it accepts any command string with zero sanitization, invokes subprocess.run(shell=True), and returns combined stdout/stderr. This is a teaching artifact for security scanning, not a safe agent tool.
Call a tool on the MCP Server.
Call Tools From MCP Server
Executes any shell command on the host OS. Provides direct system access.
Exposes all system environment variables, which may contain secrets.
Get Available Tools from the MCP Server
Get Available Tools from the MCP Server.
Reads any file from the local filesystem given an absolute path.
No input validation on 'command' parameter in execute_shell_command. Uses subprocess.run(shell=True), the highest-risk pattern for shell injection. LLMs can be tricked via prompt injection into passing 'rm -rf /' or similar destructive commands. Description does not warn of this risk.
No output schema documented for any tool. Agents cannot plan downstream calls or extract structured data. Responses are untyped strings or dicts with no schema contract.
Parameter descriptions are sparse (under 10 words on average). LLMs cannot infer constraint, format, or valid ranges. E.g., 'command' parameter in execute_shell_command is 4 words; 'args' in call_a_tool is 8 words. Rubric baseline for param descriptions: 72 chars average.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 32 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Writes or appends to any file on the local filesystem.
Naming collisions: 'list_available_tools' (tool #5) vs 'get_tools' (tool #7) both discover tools. 'call_a_tool' (tool #6) vs 'call_tool' (tool #8) both invoke tools. LLMs will waste reasoning cycles disambiguating or pick the wrong one.
Environment variable tool (get_environment_variables) exposes ALL secrets without filtering or redaction. Description warns 'may contain secrets' but provides no guidance on how agents should handle this risk. No parameter to filter or redact sensitive keys.
No error handling guidance. Tools return raw Python exception strings (e.g., 'Error reading file: [Errno 2] No such file or directory'). LLMs cannot parse these or determine if an error is retryable, user-fixable, or fatal.
write_file tool is not marked as destructive or idempotent. Agents cannot determine if repeated calls with same arguments are safe. No mention of atomicity, rollback, or backup behavior.
Tool descriptions are below baseline length. Baseline: 194 chars avg (p10=34, p90=392). These tools: 38-95 chars. Many descriptions are under 50 chars, providing minimal context for tool selection.
Remote tool invocation tools (get_tools, call_tool) accept arbitrary URLs with no validation. No mention of HTTPS enforcement, timeout, DNS rebinding protection, or SSRF mitigation. Agents could be tricked into scanning/invoking internal servers.
Tool parameter 'args' in call_a_tool and call_tool is typed as 'object' with no schema. Unstructured. LLMs have no guidance on what structure to pass. No validation of argument types or required vs. optional fields.