Gmail MCP Server exhibits inconsistent quality across 16 tools. Tool naming follows verb-noun conventions well (send_email, draft_email, read_email, search_emails, modify_email, delete_email, list_email_labels, create_label, update_label, delete_label, get_or_create_label, create_filter, list_filters, get_filter, delete_filter, create_filter_template). Descriptions are present for all tools but vary significantly in quality and clarity. Input schemas are defined using Zod with JSON Schema export, providing type information, but parameter descriptions in schemas are inconsistent, some detailed, others generic. Critical gaps: no documented output schemas visible in source code, no pagination support declared for list operations, no error handling guidance visible, no security annotations for destructive operations, and missing parameter validation documentation for several operations. The code shows basic error handling in OAuth flow but no per-tool recovery guidance. Tools like send_email and delete_email lack explicit confirmation or dry-run patterns despite destructive consequences.
Tools (16)
create_filterwriteauthsource verified70/100
Creates a new Gmail filter with specified criteria and actions
create_filter_templatewriteauth50/100
Creates a Gmail filter using a predefined template pattern
No documented output schemas for any tool. LLMs cannot determine what fields to expect in responses, forcing them to infer structure and risking incorrect downstream calls. Source code lacks return type documentation or response shape validation.
Destructive operations (delete_email, delete_label, delete_filter) lack confirmation or dry-run patterns. Agents can irreversibly delete data without explicit safeguards, inviting catastrophic errors.
delete_emaildelete_labeldelete_filter
Recommendations
Add explicit output schema documentation for all 16 tools. For send_email, document: 'Returns { message_id: string, thread_id: string, timestamp: ISO8601 }'. For search_emails, document: 'Returns { messages: [{ id, from, subject, date, snippet }], total_count: number, next_cursor?: string }'.
Implement confirmation pattern for destructive operations. Add optional boolean parameter 'confirm_deletion: true' to delete_email, delete_label, delete_filter with description: 'Prevents accidental deletion. Omit or pass false to receive a confirmation request instead of immediate deletion.'
Add pagination to list_email_labels and list_filters. Include 'max_results' (1 - 100, default 20) and 'page_token' parameters. Return a 'next_page_token' in response to enable cursor-based pagination.
Document error scenarios in each tool description. Example for search_emails: 'Errors: invalid query syntax (user-fixable: check query format), no results found (not an error, return empty list), auth failure (retryable). Returns error object with code and actionable message.'
Replace generic 'object' type for criteria and action in create_filter with explicit schema. Document required fields: criteria: { from?, to?, subject?, hasAttachment?: bool }, action: { addLabelIds?: [string], removeLabel?: bool, archive?: bool }.
Replace free-form 'template' parameter in create_filter_template with enum: 'template: enum["auto_label_promotions", "auto_archive_old", "spam_to_label", ...]. Description: 'Predefined filter template. Use list_filter_templates first to see available options.'
Score history
Overall score trend
↑ 23 points across a rubric change (v1 → v2)
58/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
58
2026-07-28+
v2
2026-03-09
F
35
-
v1
get_filter
read onlyauthsource verified75/100
Gets a specific Gmail filter by ID
get_or_create_labelwriteauthsource verified72/100
Gets an existing label by name or creates it if it doesn't exist
No pagination support visible in list operations (list_email_labels, list_filters). Large result sets will blow context windows and degrade LLM reasoning. No maxResults parameter, cursor, or offset declared for these tools.
Missing error handling guidance in tool descriptions. No indication of what errors are retryable, user-fixable, or fatal. E.g., search_emails does not explain what happens if the query is invalid or if no results are found.
Parameter 'criteria' and 'action' in create_filter are typed as generic 'object' with only vague descriptions ('Filter criteria to match messages'). No constraint documentation, no valid schema, no examples of structure. LLMs cannot infer correct usage.
Parameter 'template' in create_filter_template is a free-form string with no enum or list of valid template names. Description says 'Template name to use' but does not enumerate valid options (e.g., 'spam', 'marketing', 'archive'). LLMs will hallucinate template names.
Missing security annotations (readOnlyHint, destructiveHint, idempotentHint) on tool definitions. Agents cannot determine which tools are safe to retry or which modify state without explicit markup.
search_emails parameter 'maxResults' is typed as 'number' with no constraints (min, max, default). LLMs may pass unreasonably large values (1M+) causing API errors or timeouts. Should specify range (e.g., 1 - 500, default 20).
send_email description does not clarify whether the tool sends immediately or drafts. The presence of both send_email and draft_email creates ambiguity without explicit guardrails in the description to prevent LLM confusion.
Add tool annotations to MCP schema registration. Mark destructive tools: { destructiveHint: true, idempotentHint: false }. Mark read-only tools: { readOnlyHint: true }. Mark create operations: { idempotentHint: false (send_email is NOT idempotent, repeated calls create duplicates) }.
Add numeric constraints to search_emails maxResults: 'number (1 - 500, default 20). Limits result count to prevent context window exhaustion and API timeouts.'
Clarify send_email vs draft_email in descriptions. send_email: 'Immediately sends the email to all recipients. Use draft_email to prepare without sending.' draft_email: 'Saves email in Drafts without sending. Recipients do not receive it until you call send_email with draft_id.'
Declare scope requirements in each tool description. Example: send_email: 'Requires: gmail.modify (write access to emails). Risk: WRITE. Side effect: email is delivered to all recipients.'
Add 'recoveryHint' to error responses. Example: 'Invalid label ID "xyz". Try list_email_labels() first to retrieve valid label IDs, then retry modify_email.'
Document idempotency contract for each tool. send_email: 'NOT idempotent, each call sends a new email. Do not retry without checking for success first.' modify_email: 'Idempotent, repeated calls with same message_id and labelIds produce the same result.'
Add 'When to use' guidance to each tool description. Example search_emails: 'When to use: find specific emails by sender, subject, date, or content. If you only have a recipient name, call search_users() first to get email, then use search_emails.'
Return chaining IDs in all responses. After create_label, return { label_id, label_name, thread_ids_if_any }. After send_email, return { message_id, thread_id } so agent can immediately call read_email or search_emails without extra lookups.