MCP Server for Playwright Testing - A Model Context Protocol server for automating Playwright tests, including UI testing, API testing, test generation from requirements, and network/console capture
Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
PlaywrightMCP exhibits moderate definition quality with structural elements present but significant gaps. All 7 tools have basic schemas and descriptions, but the descriptions are often Chinese-language, brevity-optimized, and lack LLM-friendly guidance. Several tools combine multiple concerns (e.g., generate-test-cases handles both UI and API), violating single-responsibility patterns. Parameter descriptions exist but lack actionable constraints, format guidance, and error recovery hints. Output schemas are not documented anywhere in the visible source. Error handling is minimal, tools return generic isError: true responses without recovery guidance. No pagination support visible despite tools returning results. Tool names lack consistent verb-noun structure (e.g., 'clone-repository' is good, but 'use-local-project' is vague; 'execute-ui-tests' combines execution + result retrieval).
Tool descriptions are all in Chinese and cryptically brief (9 - 18 characters), far below the 10 - 1024 character baseline and 194-character production average. LLMs cannot infer tool purpose or selection criteria from '启动Playwright浏览器实例' or 'Git仓库URL'.
No output schemas documented for any tool. Responses are inferred from code (success: true/false, projectPath, reportID) but not declared in tool definitions. LLMs cannot predict response structure or plan downstream calls without explicit documentation.
Rewrite all tool descriptions in English using LLM-optimized prose. Target 50 - 200 characters per tool, answering: What does it do? When should the LLM call it? What does it return? Example: 'launch-browser' → 'Launch a Playwright browser instance (chromium, firefox, or webkit) for UI testing. Specify headless mode, slowMo delays, and viewport dimensions. Returns success status.'
Document output schemas for all 7 tools. Define the JSON structure of responses (success, projectPath, reportID, testCases, etc.) using JSON Schema. Include in tool definition or as a separate schema reference.
Eliminate tool overlap: merge generate-test-cases and generate-tests-from-spec into a single generate-tests tool that accepts either requirements text OR specPath, disambiguated by which parameter is provided.
Split generate-test-cases into two tools: generate-ui-tests (takes requirements, projectPath, optional baseUrl) and generate-api-tests (takes requirements, apiSpec path, optional baseUrl). Each is clearer for LLM reasoning.
Add numeric constraints to param descriptions: timeout (1 - 300000 ms), slowMo (0 - 5000 ms), depth (1 - 100 for git clone). Format constraints for baseUrl and specPath (must be valid URL or file path).
Refactor execute-api-tests to remove the headers parameter entirely. Instead, require baseUrl and accept standard HTTP method/content-type via enums. Store authentication headers in server-side configuration using environment variables or secure vault. Document that the tool does NOT accept credentials in parameters.
Score history
Overall score trend
↑ 56 points across a rubric change (v1 → v2)
56/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
56
2026-07-28+
v2
2026-03-09
F
0
-
v1
Semantic tool overlap: generate-test-cases (tool 4) and generate-tests-from-spec (tool 7) both generate tests from specifications, but with different interfaces and unclear distinction. LLMs will be uncertain which to call, violating single-responsibility pattern.
Parameter descriptions lack actionable constraints. No ranges for numeric params (timeout, slowMo), no format guidance (URL format for baseUrl/apiSpec), no examples of valid inputs. Produces validation failures and LLM guessing.
execute-api-tests accepts 'headers' as a free-form object. This is a critical security anti-pattern, LLMs may be tricked into passing API keys, Authorization headers, or tokens in tool parameters, which then log into traces. Should use server-side secret injection.
Error handling is minimal. All tools return {success: false, error: message} without categorizing errors as retryable, user-fixable, or fatal. No recovery guidance provided. LLMs cannot decide whether to retry, ask the user, or abort.
Tool naming lacks consistent verb-noun pattern. 'use-local-project' is vague (does it set, switch, or read?). 'execute-ui-tests' conflates execution + result retrieval. Should be more specific, e.g., 'switch-to-local-project', 'run-ui-tests'.
generate-test-cases combines UI and API test generation in a single tool based on an enum parameter (testType). Violates single-responsibility pattern. Should split into generate-ui-tests and generate-api-tests so agents can reason about them independently.
No pagination support visible. Tools that generate or execute test suites (generate-test-cases, execute-ui-tests) may return large results, but no limit, offset, or cursor parameters are defined. Could exhaust context window.
Implement error classification and recovery guidance. Examples: 'Repository already cloned locally. To update, call pull-repository or clone-repository with force:true.' 'Test suite not found. Call list-test-suites to see available IDs.' 'API timeout after 30s. Retry with higher timeout or check server status.'
Add parameter 'limit' (default 20, max 100) and 'offset' to generate-test-cases and execute-ui-tests. Return {success: true, test_cases: [...], total_count: N, next_offset: ...} to support pagination.
Rename tools for clarity: use-local-project → switch-to-local-project (indicates a state change). execute-ui-tests → run-ui-tests (simpler verb). execute-api-tests → run-api-tests.
Add a dependency hint in generate-test-cases description: 'Requires a project context set via clone-repository or use-local-project. Call those first if you don't have a project path.'
Document idempotency guarantees. E.g., 'clone-repository is idempotent, calling it twice with the same URL and branch is safe. Launch-browser can only have one active instance; calling it again closes the previous one.'
Add an optional dry-run parameter to execute-ui-tests and execute-api-tests. When true, validate test suite without running. Lets agents preview before committing.