Playwright-based browser automation MCP server with Ollama self-healing for web testing. Provides tools to interact with web pages, capture screenshots, extract interactable elements, and record/replay test scripts with AI-powered selector correction.
This MCP Playwright server has pervasive definition quality issues. Tool descriptions are present but frequently generic and lack LLM-optimized guidance. Parameter descriptions exist but many lack type information or constraints. No output schemas are documented. Critical security issue: the `submit_report` tool accepts a JWT token as a parameter, violating secret-injection patterns. The server conflates session management (init_session, close_session) with browser automation, forcing agents to reason about stateful resources. Tool naming is mostly clear (navigate, click, fill) but the generic nature of some names (get_text, get_all_text, wait_for_selector) duplicates functionality without clear disambiguation. Error handling is largely absent from descriptions. No tool annotations (readOnlyHint, destructiveHint) despite clear risk classifications. Parameter types are declared in schema but descriptions do not restate constraints (e.g., timeout in milliseconds for wait_for_selector). The server uses custom HTTP transport with JSON-RPC, not standard MCP streaming.
Click on an element identified by XPath selector
Close a browser session and clean up resources
Fill an input field with text
Get all text content matching a selector
Get all interactable elements on the current page with their properties (tag, id, name, role, text, placeholder, aria-label, xpath)
Get the text content of an element
Hover over an element identified by XPath selector
CRITICAL: `submit_report` tool accepts JWT token as a parameter ("token":"Bearer eyJ..."). Credentials must never appear as tool parameters, tokens in params leak into logs, traces, and prompt history. Violates secret-injection pattern.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite server classifying tools with explicit Risk levels. Tools marked as WRITE (navigate, click, fill, close_session, submit_report, init_session) should declare destructiveHint=true. READ_ONLY tools should declare readOnlyHint=true.
No output schemas documented for any tool. LLMs cannot plan downstream tool calls or extract the right data without knowing what fields to expect. Pattern requires 'Document the output schema.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 42 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 30 | - | v1 |
Initialize a new browser session for testing
Navigate to a URL in the current browser session
Take a screenshot of the current page
Submit a test report with JWT token to external reporting system
Wait for an element matching selector to appear
Session management (init_session, close_session) is exposed as tools, forcing agents to reason about stateful lifecycle management. Pattern recommends abstracting session state inside the server, tools should be stateless from the agent's perspective.
Tool descriptions lack LLM-optimized guidance. Most descriptions are under 50 characters and do not explain WHAT the tool does, WHEN to use it, or prerequisites. E.g., 'Get the text content of an element' (37 chars) does not explain when get_text vs get_all_text should be used, or what selector format is expected.
Duplicate and ambiguous naming for text extraction: get_text (single element) vs get_all_text (multiple elements) are not clearly distinguished. Pattern requires tool names to make distinction obvious, consider get_text_single vs get_text_all, or better yet, make get_text accept an optional 'match' parameter.
Parameter descriptions do not restate constraints from schema. E.g., wait_for_selector has a timeout parameter with type integer and description 'Timeout in milliseconds (default: 5000)', but does not specify min/max bounds. LLMs cannot read JSON Schema constraint fields, they rely on description text to understand valid ranges.
No error handling guidance in tool descriptions. Pattern requires error responses to tell the LLM what to do next. Descriptions for navigate, click, fill, etc. do not mention error cases (network failure, element not found, timeout) or recovery steps.
sessionId parameter appears in all tools but description is generic ('Session ID identifying the browser instance'). Does not explain that sessionId must be initialized first via init_session, or that IDs persist only within the server's in-process memory. Agents cannot infer lifecycle dependencies from generic descriptions.
screenshot tool description is incomplete (50 chars: 'Take a screenshot of the current page'). Does not specify output format (PNG, JPEG, base64-encoded, raw bytes), dimensions, or whether it includes viewport-only vs full-page capture. LLMs cannot reason about downstream processing without knowing format.