事件驱动确定性快照的浏览器操控 MCP 服务器 (Playwright + Accessibility Tree)
The server implements 34 browser automation tools with reasonable naming (verb_noun convention: browser_*) and basic parameter schemas. However, the definition quality is significantly hampered by: (1) descriptions in Chinese, making them inaccessible to most LLM systems trained primarily on English; (2) inconsistent schema completeness, many tools defined in bench/newtools_real.py rather than main server.py, suggesting incomplete static registration; (3) shallow parameter descriptions that lack actionable guidance on constraints, prerequisites, or recovery paths; (4) no visible output schema documentation; (5) missing explicit error handling guidance; (6) no tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear semantic differences (READ_ONLY vs WRITE vs IRREVERSIBLE vs DESTRUCTIVE); (7) several high-risk tools (browser_evaluate, browser_upload_file, browser_adopt_page, browser_dialog_respond) lack confirmation gates or safety documentation. The schema structures visible for individual tools (e.g., browser_navigate, browser_click) are reasonably complete with type definitions, but this is not universal across the toolset. Overall, this is a competent but unpolished implementation.
采纳(接管)浏览器中打开的页面
点击页面元素
关闭指定的 MCP session
关闭指定的任务(关闭浏览器上下文)
读取控制台消息
对对话框(alert/confirm/prompt)做出响应
关闭弹窗
拖拽页面元素
Tool descriptions in Chinese ('导航到指定 URL', '获取页面的可访问性树快照', etc.) are inaccessible to most LLM systems and violate the internationalization principle. All descriptions must be in English for integration with mainstream AI systems.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
读取页面错误事件
执行 JavaScript 代码并获取结果
在页面中搜索元素
将鼠标悬停在元素上
列出所有打开的页面/标签页
列出所有活跃的 MCP session
导航到指定 URL
后退到前一个页面
读取网络请求信息
获取网络请求的响应体
读取页面性能指标
按下键盘按键
读取页面元素的文本内容
获取页面的截图
滚动页面
滚动到指定元素
选择下拉框选项
获取页面的可访问性树快照
切换到指定的页面/标签页
列出当前 session 的所有任务
向输入框输入文本
上传文件到文件输入框
等待页面中出现指定文本或选择器
等待指定的毫秒数
等待页面导航完成
等待 DOM 稳定(不再变化)
High-risk tools (browser_evaluate, browser_upload_file, browser_adopt_page, browser_dialog_respond) marked IRREVERSIBLE or DESTRUCTIVE lack explicit confirmation gates or safety documentation. The 'confirmed' parameter appears in some tools but is not universally applied or explained. Error handling for unconfirmed operations should be explicit in tool descriptions.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in source. The Risk field in metadata (READ_ONLY, WRITE, IRREVERSIBLE, DESTRUCTIVE) should be formally exposed via MCP tool annotations to help clients understand safety implications and decide retry strategies.
Parameter descriptions are consistently shallow (10-30 characters). Examples: 'task_id' described as '任务 ID', 'selector' as 'CSS 选择器', 'mode' as '快照模式: reading|interactive|full'. Descriptions must explain WHAT each option does, WHEN to use it, and what the default behavior is. No actionable recovery guidance for invalid inputs.
No visible output schema documentation. Tools return results but the structure is not formally documented in the MCP schema registration. LLMs cannot infer what fields to expect. Each tool should document return type, key fields, and pagination support if applicable.
Inconsistent tool registration: tools like browser_snapshot, browser_type, browser_list_pages, browser_switch_page, browser_press_key, browser_hover, browser_select_option, browser_upload_file, browser_navigate_back, browser_drag, browser_dialog_respond are defined in bench/newtools_real.py (test/benchmark file) rather than in the main server.py. This suggests incomplete static registration in the production server and raises questions about whether all tools are actually available at runtime.
Error handling guidance is absent. Tools provide no recovery hints (e.g., 'If selector not found, use browser_find() first', 'Timeout: consider browser_wait_stable() before retrying'). Error responses will not guide LLM recovery.
Parameters like 'selector', 'ref', 'role', 'name' in browser_click are mutually exclusive but no explicit mutual-exclusivity documentation. LLMs will pass multiple simultaneously, causing ambiguity or errors.
Numeric parameters (timeout in milliseconds, amount in pixels) lack min/max constraints. LLMs can pass absurd values (timeout=-1, amount=999999) without validation guidance.
Enum parameters (e.g., 'direction' in browser_scroll, 'mode' in browser_snapshot, 'button' in browser_click) are declared in descriptions but not as JSON Schema enums. LLMs cannot introspect valid values and may hallucinate invalid ones.