The Gmail MCP server has 7 tools with clear action verbs in naming (compose_, send_, search_, query_, list_, mark_, add_). All tools have descriptions and input schemas. However, the implementation has critical gaps: (1) Missing output schema documentation for all tools, only descriptions of return values in docstrings, not structured schemas; (2) Parameter descriptions are brief but generally present; (3) No error handling guidance for LLMs (e.g., when search returns 0 results, when email send fails); (4) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite 3 WRITE tools that would benefit from these; (5) Date validation is present in code but not enforced at schema level (after_date/before_date accept any string, not a date format constraint). The tools are reasonably well-named and compose_email/send_email appropriately separate concerns. Search tools (search_emails vs query_emails) could be clearer in their distinction. Overall, this is a functional but under-specified tool set for an LLM agent.
Add a label to a message.
Compose a new email draft.
Get all available Gmail labels for the user.
Mark a message as read by removing the UNREAD label.
Search for emails using a raw Gmail query string.
Search for emails using specific search criteria.
Compose and send an email.
No output schemas documented for any of the 7 tools. LLMs cannot determine what fields to expect in responses, forcing them to parse unstructured text and plan downstream tool chains blindly.
Three write/destructive tools (compose_email, send_email, mark_message_read, add_label_to_message) lack tool annotations (destructiveHint, idempotentHint). LLMs cannot determine if calling these tools twice is safe, risking duplicate emails or unintended state changes.
Date parameters (after_date, before_date in search_emails) lack format constraint in schema. Validation occurs only in Python code, so LLM cannot see the YYYY/MM/DD requirement upfront. Schema should include 'pattern': '^\d{4}/\d{2}/\d{2}$' and description should state format explicitly.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
search_emails and query_emails serve overlapping purposes, creating tool disambiguation overhead. Description does not state when to prefer one over the other. LLMs must reason through ambiguity, increasing error likelihood.
No error handling guidance. Tools lack descriptions of failure modes and recovery steps. E.g., send_email does not state 'If recipient is invalid, returns error. Consider checking address format before sending.' LLMs cannot self-correct on failures.
list_available_labels lacks discovery guidance. Description does not explain when to call it or what fields are returned. Should state: 'Call before add_label_to_message to get valid label_id values. Returns list of {id, name, type}.'
max_results parameter in search_emails defaults to 10, but this default is not documented in the description. LLM may assume larger limits and be surprised when 10 results are returned.
No pagination guidance for tools returning lists (search_emails, query_emails, list_available_labels). No mention of limits or how to retrieve additional results if result count exceeds max_results.