MCP server for Playwright browser automation and CSSOM inspection
Shakespeare exposes 7 browser automation tools with explicit JSON Schema definitions. All tool names follow verb_noun convention (navigate, screenshot, evaluate, get_computed_styles, query_elements, set_viewport, close_browser), which is good. However, descriptions are uniformly short (15-70 chars) and lack critical contextual information: WHEN to use each tool, WHAT the return value contains, and particularly for 'evaluate', the fact that it executes arbitrary JavaScript (a major security/capability signal). Parameter descriptions exist but are minimal, e.g., 'JavaScript code to execute' does not explain format constraints, error modes, or how output size limits work. Schemas are present for all tools and properly typed, but output schemas are completely undocumented, LLMs have no way to know what 'evaluate' returns, or what fields 'screenshot' embeds in the image response. The 'evaluate' tool is particularly concerning: it accepts a script parameter with no validation guidance, no documentation of how errors are reported, and no mention of the security implications. The 'output_mode' and 'size_limit' parameters show thoughtfulness about result handling, but their interdependencies and the fallback behavior are not clearly explained in the descriptions.
Close the browser instance
Execute JavaScript in the page context and return the result
Get computed CSS styles for an element
Navigate to a URL in the browser
Query elements and get their computed styles
Take a screenshot of the current page
Set the viewport size for responsive testing
Output schemas completely undocumented. Tools return image, text, or structured data with no schema specification. LLMs cannot plan downstream operations or extract needed fields without explicit output documentation.
Tool descriptions are too brief (15 - 70 chars, well below recommended 50 - 200 char baseline) and lack critical context. 'Execute JavaScript in the page context and return the result' does not explain: (1) what happens on errors, (2) what 'result' means (string? JSON? base64?), (3) security implications, or (4) how size limits and output modes interact.
'evaluate' tool accepts arbitrary JavaScript with no input validation, format constraints, or error handling guidance. Parameter descriptions do not explain what happens if script throws, returns undefined, returns a circular object, or exceeds size limits. This is a security and usability gap.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 53 | 2026-07-28+ | v2 |
'close_browser' tool is irreversible (closes browser, likely terminating context) but description does not warn of side effects. Irreversible operations should document consequences and state if confirmation is required. This invites accidental invocation.
Parameter interdependencies not documented. In 'evaluate', 'output_mode', 'output_path', and 'size_limit' are tightly coupled, file mode requires output_path (or defaults to temp); size_limit triggers auto-file mode, but these dependencies are invisible in parameter descriptions. LLMs cannot reason about valid parameter combinations.
No pagination or result limits documented for tools that could return many results (e.g., 'query_elements' with a broad selector could match hundreds of DOM elements). Baseline pattern requires documenting result limits (e.g., 'Returns up to 50 elements; use selector refinement for more.') to prevent context window exhaustion.
Error handling guidance is absent. No tool description explains what errors are recoverable, what they mean, or how to proceed. E.g., 'navigate' may fail with network errors, timeout, or invalid URL, but LLMs get no guidance on which are retryable or how to adjust parameters.