Multi-agent framework with LangGraph runtime, browser automation, code execution, and MCP tool integration for agents
This MCP server has critical gaps in definition quality. While 6 tools are present with basic descriptions and parameter schemas, the definitions are severely underdeveloped for production use. Tool names lack clear verb prefixes (e.g., 'browser_navigate' is acceptable but 'execute_python_code' and 'execute_shell_command' lack clarity about what 'execute' means in context). Descriptions are minimal (10-50 chars) and do not explain WHEN to use each tool, what the LLM should expect, or error recovery paths. Parameter descriptions exist but are terse and often lack constraint information (e.g., no min/max for timeout_seconds, no enum for browser_click's selector). Output schemas are completely undocumented, LLMs cannot plan downstream calls without knowing what fields each tool returns. Error handling is absent, no guidance on retryability, user-fixable errors, or recovery steps. Security concerns are severe: execute_python_code and execute_shell_command are marked DESTRUCTIVE with no permission gates, dry-run options, or confirmation patterns. The server exposes two of the most dangerous tools an agent can invoke without any guardrails beyond a timeout. No evidence of idempotency checks, audit logging, or rate limiting.
Click an element in the active browser session
Fill form fields in the active browser session
Navigate to a URL in the active browser session
Capture a screenshot of the active browser session
Execute a Python script and return its stdout, stderr, and exit code.
Execute a shell command and return its stdout, stderr, and exit code.
execute_python_code and execute_shell_command marked DESTRUCTIVE with zero permission gates, confirmation patterns, or dry-run support. An agent can execute arbitrary code and shell commands without authorization checks or intervention steps.
No output schemas documented for any tool. LLMs cannot determine what fields to expect (e.g., does browser_screenshot return base64, a file path, a URL?). This forces agents to guess downstream field names and breaks tool chaining.
Tool descriptions are 10-25 characters, far below the 50-200 char baseline for production tools. Descriptions do not answer WHAT the tool does, WHEN to use it, or what it returns. Example: 'Execute a Python script and return its stdout, stderr, and exit code.' is 65 chars but does not explain when agents should use this vs alternatives, what dependencies are available, or error recovery.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 42 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Parameter timeout_seconds in both execute_* tools lacks min/max constraints. Schema shows type 'integer' and description mentions '(default 30, max 120)' but does not enforce minimum (0? 1? negative allowed?). LLMs will pass unconstrained values.
browser_click selector parameter accepts free-form CSS strings with no validation or error guidance. An LLM passing an invalid selector gets a silent failure. Description does not explain selector format, valid syntax, or how to debug selection failures.
browser_fill_form fields parameter is type 'object' with description 'Map of CSS selectors to values' but no schema for the object structure, no description of the value types, no guidance on what happens if a selector is not found. LLMs cannot infer proper invocation.
No error handling guidance across any tool. What happens if execute_python_code times out? If browser_click selector is invalid? If browser_navigate URL is malformed? LLMs receive failures with no recovery path and cannot self-correct.
execute_python_code and execute_shell_command provide no environment isolation, sandboxing verification, or resource limits beyond timeout. Agents can install packages, modify files, or attack the host system with no guards.