The Gmail MCP server has 26 tools with moderate definition quality. All tools have basic descriptions (50 - 150 chars), but most descriptions lack actionable detail on WHEN to use the tool or WHAT it returns. Input parameters are defined with types and descriptions, which is positive. However, output schemas are completely absent, no tool documents what it returns, forcing LLMs to infer structure. Parameter descriptions are minimal (5 - 15 words) and do not explain constraints, formats, or dependencies. No tool includes error handling guidance or recovery hints. Naming is generally clear (verb_noun convention), but 'remove-label' and 'apply-label' operations could be better distinguished. The server lacks idempotency markers, confirmation patterns for destructive ops (trash, delete), and batch result per-item status reporting. No tool annotation hints (readOnlyHint, destructiveHint, idempotentHint) are present despite the risk metadata being provided in the evaluation.
No output schemas documented for any tool. LLMs cannot infer what fields a tool returns, forcing them to guess or make wasteful follow-up calls. Example: list-drafts and get-unread-emails likely return arrays, but the schema (field names, types, required fields) is completely absent.
Add output schemas to every tool. Document what each returns: e.g., send-email returns {message_id: string, timestamp: ISO8601}; list-drafts returns {drafts: [{id: string, subject: string, recipient: string, created_at: ISO8601}], total: integer}.
Expand tool descriptions to 50 - 200 characters. Include 'WHEN to use this tool' guidance. Example: 'Create a draft email without sending. Use this to prepare messages before review, or let the user edit the content before sending. Requires recipient email and subject.'
Add pagination parameters (limit: 1 - 100, offset: 0+) to list-drafts, list-labels, list-filters, list-folders, list-archived. Document max result limits and include total count in responses.
Add tool annotations: mark read-only tools (read-email, list-*, open-email, search-*, get-filter) with readOnlyHint=true; mark destructive tools (trash-email, delete-label, delete-filter) with destructiveHint=true; mark idempotent operations (if applicable) with idempotentHint=true.
Add error handling guidance to all mutating tools. Example for send-email: 'If recipient validation fails, the error will state the reason (e.g., invalid email format). Retry with a corrected recipient address. If the recipient does not exist, offer to search for their email via list-contacts or search-emails.'
Document create-filter's criteria and action objects with full schemas. Example: 'criteria: {from?: string, subject?: string, has?: string[]} | action: {addLabels?: string[], archive?: boolean, delete?: boolean, markRead?: boolean}'.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Retrieves unread messages from mailbox. Returns list of message IDs in key 'id'.
list-archivedread onlyauthsource verified55/100
List archived emails
list-draftsread onlyauthsource verified58/100
List all draft emails
list-filtersread onlyauthsource verified55/100
List all email filters
list-foldersread onlyauthsource verified55/100
List all folders
list-labelsread onlyauthsource verified55/100
List all labels
move-to-folderwriteauthsource verified62/100
Move an email to a folder
open-emailread onlyauthsource verified58/100
Opens email in browser given ID.
read-emailread onlyauthsource verified60/100
Read email content
remove-labelwriteauthsource verified58/100
Remove a label from an email
rename-labelwriteauthsource verified60/100
Rename a label
restore-to-inboxwriteauthsource verified62/100
Restore an email to inbox
search-by-labelread onlyauthsource verified60/100
Search for emails with a specific label
search-emailsread onlyauthsource verified62/100
Search for emails using Gmail's search syntax
send-emailwriteauthsource verified65/100
Creates and sends an email message
trash-emaildestructiveauthsource verified62/100
Trash email
Descriptions are too brief and lack actionable guidance. Example: 'List all draft emails' (4 words) does not explain when to call this vs read-email or get-unread-emails, what fields are returned, or pagination limits. Descriptions should be 50 - 200 characters and include 'WHEN to use' context.
create-filter accepts 'criteria' and 'action' as bare objects with no schema. LLMs cannot know what fields these objects require or accept. Should define exact structure: 'criteria { from?: string, subject?: string, ... }' and 'action { addLabel?: string[], archive?: boolean, ... }'.
Destructive operations (trash-email, delete-label, delete-filter) lack confirmation or dry-run patterns. Agents can accidentally destroy data without explicit user approval. Descriptions should warn 'This cannot be undone' and implementation should support confirmation_required or require explicit approval tokens.
No error handling guidance. If trash-email fails because the email does not exist, or send-email fails because the recipient is invalid, LLMs receive no recovery hint (e.g. 'try search-emails first', 'verify recipient exists'). Error responses should guide the LLM to the next action.
Tool annotations missing. No tool is marked with readOnlyHint, destructiveHint, or idempotentHint. This metadata helps agents understand side effects at a glance. Example: trash-email should be marked destructiveHint=true; read-email should be readOnlyHint=true.
No pagination documented for list tools. list-drafts, list-labels, list-filters, list-folders, list-archived return unknown result counts. If an account has 1000 drafts, returning all could blow token budgets. Tools should accept 'limit' and 'offset' or 'page' parameters and document max results.
batch-archive does not specify per-item success/failure reporting. If 100 emails are submitted and 1 fails due to a permission or not-found error, does the operation fail entirely or return partial success? Schema and description must clarify.
Parameter descriptions lack format constraints. Example: 'recipient_id' and 'email_id' are described as generic strings but should specify 'Valid email address (RFC 5321)' and 'Gmail message ID (alphanumeric)' respectively. Without format guidance, LLMs may pass malformed input.
No idempotency guarantees documented. If an agent retries send-email with the same parameters, will it send a duplicate? Agents often retry on ambiguous failures, tools should be idempotent or clearly document non-idempotent behavior so agents add deduplication logic.
Add confirmation pattern to destructive operations. Descriptions should state 'This operation cannot be undone.' Optionally implement a dry-run parameter or require an explicit confirmation token from the user before execution.
Document batch-archive per-item success/failure. Example: 'Returns {succeeded: [email_id], failed: [{email_id, reason: string}]}' so agents know which emails succeeded and can retry or report failures.
Add format/constraint descriptions to all ID and email parameters. Example: 'recipient_id: Valid email address in RFC 5321 format (e.g., user@example.com)' and 'email_id: Gmail message ID, typically a 16-character alphanumeric string.'
Document idempotency for create-draft, send-email, and batch operations. If retrying with identical parameters should be safe, state this clearly. If not idempotent, agents need a deduplication key or mechanism to detect duplicates.
Add natural-language name resolution to tools that accept IDs. E.g., allow apply-label to accept either label_id OR label_name; implement internal lookup. This matches the chat data model where users say 'apply the Work label' not 'apply label ID L123'.
Consider offering a combined 'send_or_draft' tool that accepts a confirmation parameter, allowing users to preview before sending in a single interaction vs two separate tool calls.