AI Agent framework for WeChat and multiple messaging platforms with MCP tool integration, browser automation, scheduling, and vision capabilities
CowAgent exposes 2 tools via MCP. Both have schemas and descriptions, but significant gaps exist in parameter documentation, error handling, and composition clarity. The 'browser' tool has a comprehensive schema with 10+ parameters, but many lack detailed constraints. The 'scheduler' tool has descriptions in Chinese (non-standard for LLM interfaces) and mixes user-facing guidance with technical parameters. Neither tool includes explicit error handling guidance, recovery paths, or idempotent operation markers. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite both tools being WRITE operations. The scheduler tool description is exceptionally long (~600 chars) and contains example values that LLMs may reuse literally. Output schemas are not documented for either tool.
Control a browser to navigate web pages, interact with elements, and extract content. Actions: navigate, snapshot, click, fill, select, scroll, screenshot, wait, back, forward, get_text, press, evaluate. Workflow: navigate (auto-includes snapshot with element refs) → click/fill/select by ref → snapshot to verify. Use snapshot as the primary way to read pages. Use screenshot + send to show key results to the user. For login/CAPTCHA/authorization etc., screenshot and ask the user for help. Login state is persisted across sessions (cookies / localStorage are kept in a user profile directory), so once the user logs in to a site, the agent can keep using it without logging in again.
创建、查询和管理定时任务(提醒、周期性任务等)。 ⚠️ 重要:仅当需要「定时/提醒/每天/每周/X分钟后/X点」等延迟或周期执行时才使用此工具。使用方法: - 创建:action='create', name='任务名', message/ai_task='内容', schedule_type='once/interval/cron', schedule_value='...' - 查询:action='list' / action='get', task_id='任务ID' - 管理:action='delete/enable/disable', task_id='任务ID' 调度类型: - once: 一次性任务,支持相对时间(+5s,+10m,+1h,+1d)或ISO时间 - interval: 固定间隔(秒),如3600表示每小时 - cron: cron表达式,如'0 8 * * *'表示每天8点 注意:'X秒后'用once+相对时间,'每X秒'用interval For Web cross-channel delivery, call action='list_recipients' first, then create a task (fixed message or ai_task) using the returned channel_type and receiver.
Generic tool names lack action verbs. 'browser' should be 'navigate_page' or 'control_browser'; 'scheduler' should be 'create_scheduled_task' or 'schedule_reminder'. Generic names force LLMs to read descriptions to disambiguate, wasting reasoning cycles.
Scheduler tool description is written in Chinese (~600+ chars, exceeds recommended 300 max). LLMs trained primarily on English will struggle to parse this for tool selection. Descriptions must be in the LLM's primary language and under 300 chars.
No output schemas documented for either tool. LLMs cannot plan downstream tool calls or extract required fields. Browser tool's 'snapshot' output is opaque, what structure does it return? Scheduler's 'list' action, what fields do tasks have?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 35 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Browser tool parameters lack constraint documentation. 'ref' doesn't explain how agents obtain element refs (via 'snapshot' action?). 'timeout' has no min/max bounds (LLM could pass 1ms or 1h). 'selector' has no CSS selector format guidance.
Scheduler tool contains example values in description ('action='create', task_id='任务ID'). LLMs are known to reuse example values literally rather than adapt them to context, causing failures. Use formal schema constraints (enums, patterns) instead.
Both tools are WRITE operations but lack tool annotations (destructiveHint, idempotentHint). Browser 'click' can trigger irreversible actions (submit form, delete item). Scheduler 'delete' is destructive. LLMs need explicit markers to understand side effects.
No error handling guidance. What happens if browser navigation times out? If scheduler task fails? Tool descriptions don't guide LLMs on recovery (retry, fallback, ask user). Error messages should be actionable, not stack traces.
Scheduler 'message' vs 'ai_task' mutual exclusivity is documented only in prose. JSON Schema should use 'oneOf' or 'not' constraints to enforce this at parse time, preventing invalid tool calls.
Browser tool 'action' enum lacks per-action guidance. Each action (navigate, click, fill, select, scroll, screenshot, wait, back, forward, get_text, press, evaluate) has different required parameters and return types, but the schema treats them as a simple enum. Consider breaking this into multiple tools (navigate_url, click_element, fill_field, take_screenshot) for clarity.
Scheduler 'receiver' parameter described as 'trusted target receiver returned by list_recipients' but LLM cannot know this without calling list_recipients first. Add dependency hint: 'Call list_recipients() first to obtain a valid receiver value.'