A security proxy and research platform demonstrating MCP server vulnerabilities, tool sandboxing, and attack/defense mechanisms
This MCP Security Proxy server has critical definition quality gaps. While 8 tools are present (fetch, read_file, write_file, list_directory, get_time, store_memory, retrieve_memory, execute_sql), the evaluation reveals serious deficiencies: most tool parameters lack descriptions, output schemas are not documented, error handling is absent from descriptions, and security considerations are not addressed. The server appears to be designed as a test/thesis harness for demonstrating prompt injection vulnerabilities rather than a production-ready MCP server. Descriptions are minimal (e.g., fetch: 'HTTP fetch tool that retrieves content from remote URLs' = 60 chars; execute_sql: 'Execute SQL queries against SQLite database' = 44 chars), and critical security concerns (SQL injection, path traversal) are unaddressed. No evidence of input validation guidance, error recovery patterns, or composition strategies in the visible code.
Execute SQL queries against SQLite database
HTTP fetch tool that retrieves content from remote URLs
Get the current time
List contents of a directory
Read contents of a file from the filesystem
Retrieve data from memory/cache
Store data in memory/cache
Write content to a file on the filesystem
execute_sql tool allows raw SQL injection without input validation or sanitization guidance. No description mentions parameterized queries, validation rules, or dangerous operations to avoid. Critical security vulnerability.
Filesystem tools (read_file, write_file, list_directory) lack path traversal protection description. No mention of sandbox constraints, allowed base paths, or how '..' sequences are handled. Honeypot file at /etc/thesis_secret indicates this is intentional vulnerability for thesis testing, not production code.
All 8 tools lack parameter descriptions. 'path' parameter in read_file/write_file/list_directory has no documentation of format, constraints, or examples. 'query' in execute_sql has no guidance on parameterization or limits. 'url' in fetch lacks guidance on allowed protocols/domains.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 40 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | 2024-11-05+ | v1 |
No output schemas documented. LLMs cannot predict what fields to expect from any tool. fetch likely returns {status, body, headers} but this is inferred, not documented. execute_sql return type unknown (rows array? affected count? result set?).
Error handling descriptions are absent. None of the 8 tool descriptions explain what errors can occur, how to recover, or what the LLM should do next. No mention of 'file not found', 'permission denied', 'connection timeout', 'query syntax error', etc.
Tool descriptions are under 60 chars for most tools. Rubric baseline for A+ tools is 50-200 chars; these are 30-60 chars (get_time, fetch, execute_sql). Too brief to guide LLM selection or explain dependencies.
No permission gates or scope declarations. execute_sql and write_file are destructive but no indication of required permissions or audit logging. Per pattern:permission-gate, destructive tools must verify authority before execution.
fetch tool lacks domain/protocol restrictions. No mention of allowed URLs, timeout constraints, or handling of redirects/SSL errors. Could be exploited for SSRF attacks.
store_memory and retrieve_memory lack idempotency and lifecycle documentation. No description of memory persistence (ephemeral? across restarts?), quota limits, or expiry. No schema for what data structures are supported.
No tool composition guidance. No indication of which tools are commonly chained together or what IDs are returned for downstream use. For example, does list_directory return paths that can be fed directly to read_file?