WhatsApp, Gmail watch, and Telegram MCP servers for AI agents
This Gmail watch MCP server demonstrates solid tool organization with clear verb-noun naming (gmail_*) and Zod-based schema validation. However, the implementations reveal significant gaps in parameter descriptions, output schema documentation, and error handling guidance. Most tools lack explicit input schema visibility in parameter metadata, and descriptions, while present, are often too terse to guide LLM selection effectively. The tool composition is reasonable, single-responsibility tools like gmail_watch, gmail_send, gmail_list_watches, and gmail_update_watch are well-separated. However, gmail_send_and_watch combines send + watch creation in one call, which violates single-responsibility; this should be split or the combined behavior justified. Parameter validation is good (email format, enums, min/max constraints via Zod), but descriptions for individual parameters are missing or minimal. Output schemas are not documented in the source, the handler returns formatted JSON strings via `compactJson()` but the schema structure is not declared. Error handling is basic: environment variable checks throw generic errors ('Set GMAIL_WATCH_DEFAULT_WAKE_URL...') without recovery guidance. Security is handled well server-side (secrets via env vars, _hermesOrigin validation), but the `_hermesOrigin` parameter passing is inelegant and not documented for agents. The code uses tool annotations (readOnlyHint, idempotentHint) correctly for gmail_profile and gmail_watch, which aligns with current MCP patterns.
Explicitly close a Gmail logical watch when the sender appears finished and the objective is clear. Does not stop the mailbox-level Gmail users.watch infrastructure.
List Gmail logical watches without secrets.
Read the Gmail account identity used by this MCP.
Send a normal email without opening a watch.
Open a logical watch and send an email in one race-safe workflow. The watch is created before send and linked to the returned Gmail thread. If sending fails, the new watch is closed.
Update an active Gmail logical watch.
Tool naming violates single-responsibility: 'gmail_send_and_watch' combines two actions (send + watch creation). Rubric requires split into separate tools or explicit justification in description.
Output schemas are not documented. Handlers return formatted JSON via `compactJson()` but the structure and fields of each response are not declared. LLMs cannot plan downstream calls or extract chaining IDs without seeing the output schema.
Parameter descriptions are incomplete or missing context. E.g., 'permissions' and 'labelIds' in gmail_watch lack explanation of expected format, valid values, and purpose. LLMs cannot infer.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Open a persistent logical Gmail watch without sending email. Match one Gmail thread or exact correspondent addresses. The infrastructure mailbox watch remains independent.
Error handling lacks recovery guidance. E.g., 'Set GMAIL_WATCH_DEFAULT_WAKE_URL and GMAIL_WATCH_DEFAULT_WAKE_SECRET' is a raw error with no next step for the LLM. Should say 'Configure GMAIL_WATCH_DEFAULT_WAKE_URL and GMAIL_WATCH_DEFAULT_WAKE_SECRET environment variables, then retry.'
_hermesOrigin parameter is opaque to LLMs. Description says 'Reserved Hermes return address, injected automatically' but does not explain what this is, why it's needed, or how the LLM should treat it. If injected automatically, why is it exposed as a parameter?
Tool selection guidance is weak. Descriptions do not explain when to use gmail_send vs gmail_send_and_watch, or when to create a watch independently via gmail_watch. LLMs will guess and often pick wrong.
No pagination support declared for gmail_list_watches. If the watch list grows large, responses could exhaust context. Rubric requires limit/offset/cursor and a total count for list tools.
Parameter types are incomplete in schema definitions. E.g., gmail_update_watch's 'objective', 'permissions', 'subjectContains', 'labelIds' are visible in the input JSON with descriptions but lack explicit type definitions (string, array, etc.). Violates rule that parameters must have type definitions.