An AI agent that uses Google ADK and MCP toolsets to send WhatsApp messages. The agent understands user requests, crafts messages, adds disclaimers, and sends them via the WhatsApp MCP server.
WhatsApp Agent has 3 tools with visible schemas and descriptions. However, multiple critical gaps reduce the score significantly: (1) descriptions are present but generic and under-optimized for LLM selection; (2) parameter descriptions lack specificity about format, constraints, and validation rules; (3) no output schema documentation; (4) missing error handling guidance; (5) parameter naming inconsistencies (sender_phone_number vs recipient); (6) no discussion of idempotency, retry behavior, or side effects for the write operation (send_message). The send_message tool is particularly concerning, it has destructive semantics (sends a message) but no confirmation pattern, dry-run option, or explicit guidance on idempotency. Tool definitions are explicitly visible in the code (agent.py), so they are not inferred.
Lists available WhatsApp chats. Supports filtering by sender phone number, chat JID, and search query.
Lists messages from a WhatsApp chat. Supports filtering by sender phone number, chat JID, and time range (before/after).
Sends a WhatsApp message to a recipient. Takes recipient phone number and message content as arguments.
send_message lacks confirmation pattern and idempotency statement. A destructive operation (sends a message) should support dry-run or require explicit user confirmation to prevent accidental duplicate sends and mis-targeted messages.
No output schemas documented for any tool. LLMs cannot infer what fields are returned, what IDs are available for chaining, or what structure to expect. This forces agents to guess and retry on malformed outputs.
Parameter descriptions lack format and constraint details. 'Filter by sender phone number' does not specify format (e.g. +91XXXXXXXXXX, no +, etc.). 'Filter by chat JID' does not explain what a JID is or valid format. Timestamps (after/before) lack format (ISO 8601? Unix epoch?). This forces LLMs to guess and trial-error.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
send_message parameter names and types raise concerns. recipient is a string phone number; message is a string. However, no guidance on phone number validation (must include + ? must be E.164 format?). No character limit on message. No guidance on which language is supported or whether WhatsApp has any restrictions (length, formatting, special chars).
Parameter naming inconsistency. send_message uses 'recipient' (singular, friendly name). list_messages and list_chats use 'sender_phone_number' (full, formal name). When a tool accepts multiple identifiers (phone number, JID, etc.), names should be explicit and consistent across all tools. This inconsistency forces LLMs to reason about which identifier to use in each context.
No error handling guidance. What happens if recipient is invalid? If message delivery fails? If network is down? Responses must include actionable recovery steps (retry, ask user for clarification, etc.), not just error codes.
list_messages and list_chats lack pagination info. No mention of result limits, page size, or how to fetch additional results. If a chat has thousands of messages, returning all could exhaust context. Tools should declare result limits and offer pagination (limit, offset, next_cursor).
Tool descriptions are generic and could apply to many tools. 'Sends a WhatsApp message to a recipient' is true but does not guide LLM selection. Why use this over a hypothetical send_bulk_messages? When should I call this vs list_messages? Descriptions should clarify unique capability and use cases.