MCP server exposing read-only Gmail inbox tools
This Gmail MCP server demonstrates solid tooling fundamentals with 24 well-structured tools covering read, write, and management operations. Tool naming follows verb_noun conventions (gmail.searchMessages, gmail.createDraft, gmail.archiveMessages). Descriptions are present and reasonably detailed (average ~180 chars), with most tools explaining WHAT they do and WHEN to use them. Input schemas are comprehensive with proper JSON Schema types, constraints, and descriptions. However, there are noticeable gaps: output schemas are not explicitly documented in the tool definitions (only inferred from Gmail API responses), some descriptions lack dependency hints and prerequisites, and a few tools could benefit from more explicit guidance on expected result structures. Error handling is present (tool shows errorReporting=true) but recovery guidance in error messages is not visible in the schema. Tool composition is strong, each tool has one clear responsibility, and output fields (threadId, messageId, etc.) align with input parameters across the tool chain. The server supports natural identifiers (email addresses) alongside IDs, which is excellent for chat-driven workflows.
Add labels to messages or threads. Specify label IDs (use listLabels to discover available labels). Adding the same label multiple times is idempotent.
Archive messages or threads (remove from inbox). Supports both message IDs and thread IDs. Thread IDs are more efficient for archiving entire conversations.
Initiates the OAuth consent flow to connect Gmail
Run multiple Gmail search queries in parallel for faster results. Use this instead of sequential searchMessages calls when you need to search multiple queries. Each query runs independently and results are returned together.
Create a new draft message. Requires gmail.compose scope.
Create a new custom label in the mailbox. Label names must be unique within the mailbox.
Output schemas not documented in tool definitions. Tool descriptions mention response formats (metadata, summary, full) but the actual returned fields, types, and structure are not formally specified. This forces LLMs to infer output structure and risks misuse downstream (e.g., extracting threadId when the response doesn't include it).
Generic descriptions for state-mutation tools (markAsRead, markAsUnread, star/unstar). Descriptions are present but lack guidance on when to use each tool vs. alternatives, potential side effects (e.g., 'markAsRead marks the message read and may update the unread count'), or chaining hints. This is below the baseline of 50-200 chars of LLM-optimized context.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | A | 80 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 27 | - | v1 |
Delete a draft message permanently.
Get attachment information for a message (filename, MIME type, size). Does not download the actual attachment content.
Get a specific draft message by ID.
Get metadata about a specific label (name, message count, unread count). Useful for checking label statistics before operating on labeled messages.
Get a single message by ID. Formats: "metadata" (headers/snippet only, fastest), "summary" (headers + first 2KB of text body, no HTML), "full" (complete message with truncation options). Use "summary" for quick content preview without downloading huge HTML emails.
Get a thread with all messages and metadata. Includes complete message bodies formatted for readability. Returns both individual messages and thread-level metadata (subject, participants, message count).
List all draft messages with pagination support.
List all labels in the mailbox with their IDs and metadata. Returns system labels (INBOX, SENT, DRAFT, SPAM, TRASH) and user-created labels.
List conversation threads with pagination support. Returns metadata about each thread (subject, snippet, participant count, message count). Use with pageToken for efficient pagination through large mailboxes.
Mark messages or threads as read.
Mark messages or threads as unread.
Remove labels from messages or threads. Specify label IDs to remove.
Search messages using Gmail query syntax (e.g., "from:john subject:meeting"). Results include threadId and threadMessageCount to support thread-aware operations. TIP: Use threadIds with archiveMessages to ensure entire conversations leave your inbox.
Star messages or threads.
Returns whether the current user has Gmail connected
Restore archived messages or threads to inbox. Moves archived items back to the inbox label.
Remove stars from messages or threads.
Update an existing draft message. All provided fields will replace the current values.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are not present in tool definitions. Risk field is inferred from the provided metadata (READ_ONLY, REVERSIBLE, WRITE, DESTRUCTIVE) but not exposed to the MCP client via tool annotations. This prevents clients from making informed decisions about tool execution (e.g., warning before delete operations).
No batch variants for common state-mutation operations. Tools like addLabels and removeLabels accept arrays (good), but markAsRead, markAsUnread, starMessages, and unstarMessages require sequential calls when the agent needs to mark multiple items. Batch support would reduce token waste and latency.
Error handling and recovery guidance not visible in schema. While errorReporting=true indicates error messages are returned, the tool definitions do not show what error cases to expect or how to recover (e.g., 'Authorization required. Call gmail.authorize() first if not authenticated.'). This limits LLM self-correction.