Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The server provides 11 tools with clear verb-noun naming (start_server, stop_server, get_devserver_logs, browser_navigate, etc.) and consistent parameter documentation. All tools have input schemas with parameter types and descriptions. However, there are notable gaps: (1) output schemas are not explicitly documented in the code, return types are inferred from type hints and Pydantic models but not formally described in tool definitions; (2) error handling is minimal, most Playwright tools catch exceptions and return generic error dicts with only 'status' and 'message' fields, lacking actionable recovery guidance or error classification; (3) parameter descriptions are functional but brief, averaging 40-60 characters, well below the 72-char baseline for A+ tools; (4) no tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk classifications in the spec (start_server and stop_server marked WRITE, browser_click/type/resize/screenshot marked WRITE, get_* tools marked READ_ONLY). The Playwright experimental tools are well-structured and respect pagination/offset/limit patterns, but the core dev-server tools lack detailed response documentation. Missing are per-tool recovery hints, dependency documentation, and confirmation patterns for destructive operations.
Tools (11)
browser_clickwritesource verified74/100
Click an element on the page by selector reference
Error handling is minimal and lacks actionable recovery guidance; exceptions are caught and returned as generic {status: error, message: str} without classification (retryable, user-fixable, fatal)
Recommendations
Add tool annotations to all tools: use readOnlyHint for get_*, destructiveHint for start_server/stop_server/browser_click/browser_type/browser_resize/browser_screenshot, idempotentHint for idempotent operations. This enables LLM safety reasoning.
Document output schemas explicitly for every tool. For example, start_server should declare: 'Returns {success: bool, message: string, server_name: string, status: string}'. Use JSON Schema notation or inline type descriptions.
Expand tool descriptions to 100-200 characters with WHAT, WHEN, and WHY. Example: 'Start a development server by name. Call this after server creation to launch it for testing. Returns server process info and startup status. Requires server to be configured in devserver.yml.' instead of 'Start a development server by name'.
Expand parameter descriptions with constraints, examples, and LLM guidance. Example for 'name' param: 'The name of the server to start (string, 1-50 chars, must match a configured server in devserver.yml; available servers: api-server, frontend, db-worker)' instead of 'The name of the server to start'.
Add min/max constraints to numeric parameters. For limit, offset, width, height: document bounds (e.g., 'limit must be 1-1000, default 100', 'width and height must be 320-1920 pixels').
Implement error classification in Playwright tools: catch and return {status: 'error', code: 'timeout|selector_not_found|navigation_failed', message: '...', retry_after_ms: 5000, next_action: 'Try calling browser_navigate() first or use a different selector'} instead of generic {status: 'error', message: str(e)}.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Add a dry-run parameter to start_server and stop_server (default false) to enable safe pre-flight checks: 'dry_run (bool, default false): If true, validate server exists and would start/stop without executing.' This prevents accidental server disruptions.
Document pagination behavior for get_devserver_logs and browser_console_messages: clarify that offset+limit work together, state maximum returnable items (e.g., 'Logs are buffered in memory; querying beyond 10000 lines is not supported'), and provide cursor-based pagination option if large datasets are common.
Add a tool to discover available server names (e.g., list_available_servers()) so agents don't have to guess valid values for start_server/stop_server name parameter.
For Playwright tools, add session/context management: document whether browser state persists across calls, whether each browser_navigate() starts a fresh session, and how to reset browser context between test runs.