A comprehensive MCP server with filesystem operations, code execution, browser automation, dynamic tool creation, and integrations with external services
This server has 25 tools with significant definition gaps. Most tools lack complete parameter descriptions and output schemas. Naming is generally adequate (verb-first convention), but descriptions are often bare or generic. Critical issues: execute_python_code and create_tool expose dangerous operations with minimal safeguards; no input validation constraints visible; error handling is absent; security patterns not evident. Per-tool analysis shows 8 tools score below 30, indicating structural problems across the toolkit.
Asks the human user a question and waits for a response.
Calls a dynamically created tool.
Clicks an element matching the selector.
Creates a new dynamic tool.
Deletes a file from the workspace.
Deletes a dynamic tool.
Executes Python code in a sandboxed environment.
execute_python_code and create_tool expose arbitrary code execution with no sandboxing, no input validation, no capability restrictions, and no permission gates. The description 'Executes Python code in a sandboxed environment' is false confidence, no actual sandboxing visible in server.py.
Missing output schemas for 12 tools. Tools like get_page_content, ask_human, open_page, fetch_url, puppeteer_navigate, run_command return no documented schema, LLMs cannot predict what fields to extract or chain calls.
Destructive operations (delete_file, delete_tool, kill_process) have no confirmation step, dry-run mode, or multi-step confirmation. Agents can irreversibly delete files or kill processes with a single tool call.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Fetch content from a URL and return as text
Gets the current time in a specific format and timezone.
Gets the text content of the current page.
Get tools organized by category.
Kill a process by name or ID
Lists files in a directory within the workspace.
Opens a URL in the browser.
Fetch web page content using Playwright (headless browser).
Click an element
Evaluate JavaScript in the browser console
Fill an input field
Navigate to a URL
Take a screenshot
Reads a file from the workspace.
Run a shell command
Takes a screenshot and saves it to the workspace.
Types text into an element matching the selector.
Writes content to a file in the workspace.
Error handling absent across all tools. No guidance on retryability, error classification, or recovery actions. A tool failure returns a raw exception, agents cannot self-correct.
Redundant browser tooling. Both Playwright (playwright_fetch) and Puppeteer (puppeteer_navigate, puppeteer_click, etc.) expose overlapping functionality. LLMs must reason about which to use; no clear canonical choice documented.
Parameter descriptions lacking format constraints and ranges. E.g., open_page accepts 'url' but does not specify format (http:// vs ftp://?), timeout, or error behavior. run_command accepts 'command' with no validation or sandboxing.
No audit trail or logging integration visible. Destructive and sensitive tools (delete_file, write_file, run_command, kill_process) do not log caller, timestamp, or parameters. Compliance and forensics are impossible.
Missing permission gates. No evidence of authorization checks before executing delete_file, delete_tool, run_command, kill_process, or execute_python_code. Any caller can destroy data or crash the system.
call_dynamic_tool uses JSON-encoded string for 'args' parameter instead of structured object. This forces LLMs to manually serialize JSON and is error-prone. Should accept an object with typed fields.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible. LLMs cannot determine which tools are safe to retry, which modify state, or which are read-only without parsing descriptions.