Model Context Protocol server for Gmail integration, providing tools to search, read, send emails, list labels, and manage email threads via Gmail API
This Gmail MCP server has fundamental definition quality gaps. All 5 tools have reasonable names starting with action verbs (search_, read_, send_, list_, get_), which is positive. However, tool descriptions are generic and lack LLM-optimized guidance. Most critically, NO OUTPUT SCHEMAS are documented anywhere in the provided source, the server defines input parameters but provides zero information about what these tools return, forcing LLMs to guess at response structure. Parameter descriptions are minimal and lack actionable constraints. Error handling is absent from the tool definitions. The send_email tool lacks any confirmation or dry-run pattern despite being destructive. Tool composition is reasonable (5 focused tools, no mega-tools), but the implementation lacks the rigor expected of production-grade agent tooling.
Get all messages in an email thread/conversation.
List all Gmail labels/folders.
Read the complete content of a specific email.
Search for emails using Gmail query syntax. Examples: - "from:john@example.com" - emails from specific sender - "subject:meeting" - emails with "meeting" in subject - "is:unread" - unread emails - "has:attachment" - emails with attachments - "after:2024/01/01" - emails after a date - "from:john@example.com subject:report" - combined search
Send an email via Gmail.
No output schemas documented for any tool. The server defines what parameters tools accept but provides zero specification of what they return, response structure, field names, types, pagination, counts, or references.
Tool descriptions are too brief and lack actionable context for LLM selection. E.g., 'read_email' description is only 38 chars ('Read the complete content of a specific email.'). Best practice is 50-200 chars with WHEN to use and WHAT is returned.
Parameters lack format/range constraints. 'max_results' has no documented minimum, maximum, or default behavior. 'query' has no guidance on Gmail syntax or error cases. Missing type suffixes: 'message_id' and 'thread_id' should be clearer about whether they accept human-readable names or only system IDs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
send_email tool is destructive (WRITE risk) and lacks confirmation or dry-run pattern. No error handling strategy documented. LLMs should be prevented from sending emails without explicit confirmation or at minimum clear error guidance on failure.
No error handling guidance. Tools have no documented error categories, recovery paths, or actionable error messages. If search_emails returns empty, fails auth, or hits rate limits, there's no specification of what the LLM should do next.
Parameter 'include_spam_trash' (boolean) has no default behavior explicitly stated. When omitted, does it default to true (include spam/trash) or false (exclude)? This ambiguity invites misuse.
search_emails examples in description mention Gmail syntax ('from:', 'subject:', etc.) but do NOT appear as enum constraints or validated patterns. LLMs may hallucinate invalid query syntax.
No pagination or result limits documented. search_emails accepts 'max_results' up to 50, but no guidance on pagination token, cursor, or offset for retrieving additional results beyond 50.