Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The server defines 5 tools with explicit schemas and descriptions, but lacks rigor in several critical areas. Tool naming follows verb_noun convention (good), but descriptions are generic and lack actionable context for LLM decision-making. Parameter descriptions are sparse, most lack detail on expected format, constraints, or when to use optional params like 'timeout'. Output schemas are not documented; the code returns TextContent and ImageContent but the tool definitions don't declare return types. Error handling is minimal, exceptions are caught and returned as text, but errors lack categorization (retryable vs fatal) or recovery guidance. Security is weak: the server launches an actual headless browser and executes arbitrary JavaScript (puppeteer_evaluate), which is extremely high-risk in an agentic context. No input validation, no rate limiting, no audit logging. The tool composition is reasonable (single responsibility per tool), but the lack of pagination, result limits, and chaining IDs makes multi-step workflows inefficient.
Arbitrary JavaScript execution (puppeteer_evaluate) without input validation, sandboxing, or user confirmation. An agent can inject malicious code, exfiltrate data, or cause side effects. This is a critical security vulnerability for agentic use.
No output schemas documented. Tools return TextContent and ImageContent (visible in code), but tool definitions don't declare return types. LLMs cannot plan downstream calls or validate response structure.
Parameter descriptions lack actionable detail. 'timeout' appears in 4/5 tools but no guidance on what event it times out on (selector finding? navigation? rendering?). No min/max constraints. LLMs guess at reasonable values.
Add explicit output schemas to all tool definitions. Specify that puppeteer_screenshot returns {type: 'object', properties: {message: string, image_b64: string, width: number, height: number}}. This allows LLMs to plan downstream calls.
Expand tool descriptions to 80-150 characters with actionable context. E.g., puppeteer_navigate: 'Navigate to a URL and wait for the page to load. Use this to start a new browser session or move to a different page. If you only have a partial URL or need to search for a link, take a screenshot first to inspect the page.' This guides LLM selection.
Add min/max constraints and format guidance to all parameters. E.g., timeout: {type: 'number', minimum: 1000, maximum: 120000, description: 'Time to wait in milliseconds (default 30000). Must be between 1 and 120 seconds.'}
Implement strict input validation and error recovery. For selectors, validate CSS syntax. For URLs, validate format (http/https). For JavaScript, use an allowlist of safe functions (e.g., document.body.innerText) and reject dangerous ones (fetch, eval, XMLHttpRequest). Return actionable errors: 'Invalid CSS selector "foo[", check bracket matching. Call puppeteer_screenshot to inspect the page.'
Add error categorization to tool responses. Structure error returns as {type: 'error', category: 'retryable|user_fixable|fatal', message: '...', recovery: '...'}. E.g., 'selector not found (user_fixable): CSS selector "#submit-btn" not found. Call puppeteer_screenshot to see the current page state and find the correct selector.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Tool descriptions are generic and lack LLM-optimized context. 'Navigate to a URL' doesn't explain when to call this vs other tools, prerequisites, or what constitutes success. Descriptions should be 50-200 chars and answer: what, when, and why.
No error categorization or recovery guidance. Errors are caught and returned as generic TextContent. LLMs don't know if a failure is retryable, user-fixable, or fatal. E.g., 'selector not found' should suggest calling screenshot to inspect the page state.
Browser state is global and persistent. Multiple tool calls operate on the same page instance. No isolation between agents/users. If agent A navigates to a malicious site, agent B's subsequent click tool might click on unexpected content.
No rate limiting, timeouts, or resource controls. A runaway agent could launch infinite browsers, exhaust memory, or hang indefinitely on slow page loads. The check_page_loaded() function has a 5-second hardcoded max_wait that's not configurable.
No input validation or sanitization. Agent-provided selectors, URLs, and scripts are passed directly to Playwright with no checks for SQL injection, command injection, or path traversal. Selectors like ''; DROP TABLE --' could cause undefined behavior.
No audit logging. Who called what tool, when, and what happened is not logged. This violates compliance requirements and makes debugging/incident response difficult.
No idempotency guarantees. Clicking a button, filling an input, or executing JavaScript are not idempotent. If an agent retries due to a timeout, the action may execute twice (duplicate submissions, accidental double-clicks).
puppeteer_clickpuppeteer_fill
Gate puppeteer_evaluate behind a strict allowlist or remove it entirely. Arbitrary JavaScript execution is too dangerous for production agentic use. If JavaScript evaluation is essential, restrict to read-only operations (document.body.innerText, document.querySelectorAll) and explicitly reject fetch, eval, localStorage, etc. Add a confirmation step via MRTR (Multi Round-Trip Request) requiring user approval before execution.
Implement per-agent/user browser isolation. Use a pool of isolated browser instances with per-request cleanup. Prevent state leakage between concurrent agent calls.
Add rate limiting and timeout guards. Limit puppeteer_navigate to 5 concurrent calls per agent. Set a global timeout (e.g., 60 seconds) for all operations. Return clear timeout errors: 'Navigation timed out after 60 seconds. The page may be unresponsive. Try a different URL or check your network.'
Enable audit logging. Log every tool call with timestamp, agent ID, parameters (excluding secrets), and result. Write to stderr or a syslog endpoint for compliance tracking.
Document all parameters with type, description, and constraints in the schema. Currently many are missing type info or have vague descriptions.
Add documentation on browser prerequisites and lifecycle. Clarify that the first call to any tool launches a browser instance, and specify how/when it shuts down.
Implement idempotency tokens for state-modifying operations. For puppeteer_click and puppeteer_fill, accept an optional idempotency_key parameter. Dedup based on key to prevent double-execution on retries.