MCP server for WhatsApp integration via Baileys library, enabling message sending, retrieval, contact management, and chat operations through a FastMCP HTTP bridge
WhatsApp-MCP has significant definition quality gaps. While 7 of 8 tools have descriptions (typically 30-60 chars), they are terse and lack context for LLM decision-making. Naming is mostly verb-noun compliant (get_, send_, connect_), but descriptions do not explain WHEN to use each tool or dependencies between them. Input schemas exist for all tools but lack depth: many parameters lack explicit type/format constraints, descriptions are minimal, and several tools omit required parameter documentation. No output schemas are documented in the visible source code. Error handling is absent from the visible definitions. The most critical issue: several parameter descriptions are identical or generic (e.g., 'Optional' with no context), forcing LLMs to guess at constraints. Per-tool analysis reveals consistent pattern: tool description too short (<20 chars deduction), parameters partially described, no output schema, no error recovery guidance.
Initialize WhatsApp connection.
Get chat conversations list.
Get contacts list.
Get messages from a chat or search across all chats.
Check WhatsApp connection state.
Search messages across all chats.
Send message to phone number.
No output schemas documented. Tools return data but LLMs cannot plan downstream calls without knowing response structure. Critical for agents composing tool chains (e.g., get_whatsapp_contacts returns what fields? Does it include contact_id that send_whatsapp_message expects?).
Tool descriptions are too brief (<20 chars) and lack context. 'Check WhatsApp connection state' and 'Initialize WhatsApp connection' do not explain prerequisites, when to call relative to other tools, or what fields returned will contain. LLMs cannot disambiguate between similar operations.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Send message to JID from contacts/chats.
Parameters lack type constraints and enums. 'limit' parameter across multiple tools lacks min/max bounds. 'search' parameter is a free-form string with no format guidance. 'phone' parameter in send_whatsapp_message lacks regex pattern or length constraints. LLMs will pass arbitrary values.
Parameter descriptions are minimal. 'Optional' and bare labels like 'Filter by name/phone (optional)' do not explain what operation is performed or how the parameter affects behavior. Based on rubric baseline (average param description 72 chars), most params are 20-30 chars.
No error handling guidance. Tools provide no recovery hints. If send_whatsapp_message fails because phone is invalid, there is no error message telling the LLM to use search_whatsapp_contacts first to verify the recipient. If connect_whatsapp fails, there is no guidance (retry? check network? scan QR again?).
Missing tool composition support. send_whatsapp_message accepts 'phone' but no guidance on whether to use get_whatsapp_contacts first. send_whatsapp_message_to_jid accepts 'jid' but LLM does not know how to obtain a JID (from get_whatsapp_chats? contacts list?). Tool chain is opaque.
Destructive operations lack confirmation or dry-run. send_whatsapp_message performs WRITE but no confirmation step or dry-run mode. Agent mistakes (e.g., typo in phone number) result in messages sent to wrong recipients with no undo path.
Pagination guidance missing. get_whatsapp_messages and get_whatsapp_chats accept 'limit' and 'offset' but descriptions do not specify default behavior, max safe limit, or whether a total_count is returned. LLMs cannot plan pagination without this metadata.
Tool annotations missing. No toolAnnotations in feature set. Tools that modify state (send_whatsapp_message, connect_whatsapp) should declare destructiveHint:true or idempotentHint:true so agents know retry behavior.