Browser automation MCP server based on DrissionPage and FastMCP, providing rich browser operation APIs for AI calls
The DrissionPage MCP server exhibits significant quality gaps. While it provides 27 tools with detailed descriptions emphasizing workflow prerequisites, critical issues emerge: (1) Parameter schemas are visible but lack proper type enforcement in many cases; (2) Tool descriptions, though comprehensive (averaging 150-250 chars), contain emojis and workflow guidance that waste tokens without adding machine-actionable constraints; (3) Output schemas are entirely undocumented, no tool specifies what fields it returns or their types; (4) Error handling is absent, tools lack recovery guidance or error categorization; (5) Security concerns: tools like 'load_cookies', 'save_cookies', 'download_file' and browser control tools (connect_browser, close_browser) expose powerful capabilities without permission gates or audit-trail documentation; (6) Composition issues: 27 tools is excessive for a browser automation server, many could be consolidated (e.g., click_element, input_text, scroll_page are low-level primitives that should be wrapped into higher-level workflows). The server compensates somewhat with rich descriptions and enum constraints on some parameters (selector_type, direction), but lacks the rigor expected of production tools. Per-tool scores average 42 because descriptions exist but output schemas don't, and error handling is minimal.
清空网络日志
点击页面元素(智能优化版)。⚠️ 重要提示:使用此工具前,请务必遵循标准化工作流程:1. 📸 先使用 take_screenshot() 确认目标元素存在 2. 🔍 使用 get_dom_tree() 或 find_elements() 分析页面结构 3. 🎯 基于准确信息构建选择器,禁止猜测元素名称
关闭浏览器
关闭当前标签页
连接到浏览器或启动新浏览器
下载文件
启用网络监控
查找页面元素(查询工具)。⚠️ 核心查询工具:用于精确定位元素!主要用途:1. 🔍 使用CSS/XPath查找元素 2. 📋 获取匹配元素的详细信息 3. 🎯 验证选择器的有效性 4. ✅ 构建精确选择器的基础工具
Output schemas entirely undocumented. Tools (get_page_text, get_dom_tree, find_elements, get_network_logs, get_all_clickable_elements, get_browser_info) return complex data structures with no specification of fields, types, or structure. LLMs cannot reliably extract needed fields or chain calls.
No error handling or recovery guidance. Tools lack error response specifications, categorization (retryable vs fatal), or actionable guidance. E.g., click_element offers no instruction if selector fails or element not found.
Descriptions use decorative emojis (📸, 🔍, 🎯, 🐛, 📝, ⚠️) and workflow prescriptions ('先使用', '请务必') that waste tokens. While helpful for human reading, these reduce parsing clarity for LLMs and pad token count without adding machine-actionable constraints.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
获取页面所有可点击元素信息
获取浏览器信息
获取当前页面URL
获取DOM树结构及其文本内容。⚠️ 核心预处理工具:这是标准化工作流程的第3步!主要用途:1. 🔍 分析页面DOM树结构 2. 📋 获取精确的元素定位信息 3. 🎯 为选择器构建提供精确的元素信息 4. ✅ 验证元素的层级关系和属性
获取元素文本内容(精确定位版)。⚠️ 重要提示:这是预处理工具,用于获取精确的元素信息。使用场景:1. 🔍 在点击或输入操作前,验证目标元素的实际文本内容 2. 📋 获取页面动态内容,如表格数据、状态信息等 3. ✅ 确认元素存在性和可见性
获取网络请求日志
获取页面源码HTML
获取页面完整文本内容(预处理必备工具)。⚠️ 核心预处理工具:这是标准化工作流程的第2步!主要用途:1. 🔍 在操作元素前,获取页面的完整文本信息 2. 📋 为非多模态LLM提供详细的页面内容描述 3. 🎯 帮助构建精确的元素选择器 4. ✅ 确认页面加载完成和内容可用性
获取页面标题
获取截图二进制数据并返回为base64编码资源
在输入框中输入文本(智能优化版)。⚠️ 重要提示:使用此工具前,请务必遵循标准化工作流程:1. 📸 先使用 take_screenshot() 确认输入框存在且可见 2. 🔍 使用 find_elements() 验证输入框的选择器 3. 🎯 基于准确的DOM信息构建选择器
从文件加载cookies
导航到指定URL
创建新标签页
保存当前页面的Cookies
保存页面源码到文件
滚动页面
截取页面截图(标准化工作流程第1步)。⚠️ 核心预处理工具:这是标准化工作流程的第1步!主要用途:1. 🔍 视觉确认:在任何元素操作前,先确认目标元素存在 2. 📋 为多模态LLM提供视觉上下文信息 3. 🐛 调试辅助:操作失败时用于问题诊断 4. 📝 文档记录:保存操作过程的视觉证据
等待指定秒数
Security: Tools for browser control (connect_browser, close_browser), file operations (save_cookies, load_cookies, download_file, save_page_source), and network monitoring (enable_network_monitoring, clear_network_logs) lack permission gates. No scope declarations (e.g. 'write:filesystem', 'write:cookies', 'write:browser'). No audit-trail documentation.
Tool composition fragmented. 27 tools including many low-level primitives (click_element, input_text, scroll_page, wait, close_tab). These should be wrapped into higher-level workflows (fill_form, navigate_and_wait, perform_search) so agents don't waste tokens chaining primitives.
Generic descriptions. Tools like get_page_source, get_page_title, get_current_url, wait, close_tab, clear_network_logs have descriptions under 40 characters. These lack context for LLM selection and do not explain when to use them vs related tools.
Parameter descriptions use Chinese-language text, reducing accessibility for English-speaking LLMs and developers. Translations or multilingual descriptions should be provided.