Single tool 'playwright_eval' has critical definition gaps. Tool name lacks action verb (should be 'evaluate_javascript' or 'run_javascript'). Description is present but generic and doesn't explain parameter semantics, expected output format, or error handling. Input schema exists but lacks type validation depth. No output schema documented. Tool performs irreversible code execution but provides no confirmation mechanism or error recovery guidance. Parameter 'code' has minimal context about what variables are available, what the return value should be, or how to handle errors. No error classification or recovery hints provided.
Tools (1)
playwright_evalirreversiblesource verified35/100
Evaluate JavaScript code (supports await) in a persistent Playwright context.
Variables sharedState, browser, context, page are maintained between evaluations, others are lost.
To accumulate data, put it into sharedState.
Place screenshots in temp folder
DO NOT USE waitForLoadState('networkidle') or waitForSelector
DO NOT CODE COMMENTS OR WHITSPACE
Tool name does not start with action verb. 'playwright_eval' is vague; should be 'evaluate_javascript' or 'run_javascript_code' to clearly signal LLM intent.
Input schema lacks comprehensive type constraints and descriptions. Parameter 'code' has no length limits, no validation of syntax, no explanation of available context (sharedState, browser, context, page, sessionManager). LLM cannot determine valid code patterns.
No documented output schema. Tool description mentions sharedState and variables but does not specify what the function returns, what fields are in the response, or how to access results. LLM cannot plan downstream operations.
Irreversible operation (code execution, page navigation, state mutation) lacks confirmation mechanism or dry-run support. Tool performs destructive actions without user approval or undo capability.
Recommendations
Rename tool to 'evaluate_javascript' or 'run_javascript_code' to start with an action verb matching LLM intent patterns.
Expand input schema: add maxLength constraint (e.g. 50000 chars), add examples of valid code patterns, document available globals ('browser', 'context', 'page', 'sharedState', 'sessionManager'). Use an enum or pattern to prevent obvious errors.
Document output schema explicitly: 'Returns an object with { result: any, error: null|string, output: string[], executionTime: number }' or equivalent. Specify what happens when code throws.
Add a 'dry_run' boolean parameter (default false) to support confirmation-request pattern. When true, validate code syntax and globals without executing side effects.
Enhance error handling description: 'Errors are returned as { error: string, suggestion: string }. Common errors: SyntaxError (invalid code), ReferenceError (undefined variable, check sharedState), TimeoutError (execution exceeded 30s limit). Always provide a corrected code snippet in the suggestion field.'
Rewrite parameter description: 'JavaScript code (async/await supported) to execute in a persistent Playwright context. Available globals: browser (BrowserContext), context (BrowserContext), page (Page), sharedState (object for state between calls), sessionManager (SingleBrowserSessionManager). Code is executed as an async function, to return a value, make it the final expression or assign to sharedState. Max 50000 chars. NO comments, NO whitespace (minified). Avoid waitForLoadState("networkidle") and waitForSelector, use waitForNavigation() instead. Screenshots should be saved to a temp folder using page.screenshot({path: "/tmp/..."}). Execution timeout: 30 seconds.'
Error handling provides no recovery guidance. Description includes constraints ('DO NOT USE waitForLoadState', 'DO NOT CODE COMMENTS') but does not explain what happens when constraints are violated or how LLM should respond to failures.
Parameter description lacks depth. 'JavaScript code to evaluate' does not explain: (a) available global variables and their types, (b) expected return value format, (c) how to accumulate state across calls, (d) what errors look like, (e) execution timeout, (f) heap/memory limits.
Tool description includes example constraints but no positive examples of valid usage. 'DO NOT USE X' is incomplete, description should show what TO USE instead (e.g. 'Use page.waitForNavigation() to wait for page loads').
playwright_eval
Add example use cases to description: 'Example: navigate to a page with `await page.goto("https://example.com"); await page.waitForNavigation();` then extract data with `return await page.evaluate(() => document.title);`. Store multi-step results in sharedState to access across calls.'
Document which Playwright methods are safe and which are not. Replace 'DO NOT USE' with 'SAFE: page.goto, page.evaluate, page.screenshot, page.click, page.fill. AVOID: waitForLoadState("networkidle"), use waitForNavigation(), and waitForSelector, use waitForFunction() instead.'
Implement per-item error classification in the response: 'Parse code, catch SyntaxError and ReferenceError early (before execution), return { success: false, errorType: "SyntaxError"|"ReferenceError"|"TimeoutError"|"ExecutionError", message: string, suggestion: string } so LLM knows whether to retry, fix code, or escalate.'
Add tool annotation hints to the schema: 'destructiveHint: true' (code can mutate page state), 'idempotentHint: false' (side effects on page), so LLM knows retry semantics.