MCP server providing access to Gmail functionality including searching threads, reading messages, managing drafts, and analyzing email content with AI assistance
This Gmail MCP server has multiple critical gaps in definition quality. While tool naming follows verb-noun conventions and schemas are visible in main.go, descriptions are present but generic, parameter coverage is incomplete, and output schemas are entirely undocumented. Critically, there is no visible error handling guidance, no confirmation mechanisms for destructive operations (send_draft, create_draft), and no documentation of what these tools return. The server lacks schema documentation for outputs, meaning LLMs cannot plan downstream operations or extract structured data reliably. Parameter descriptions exist but lack specificity around formats, ranges, and validation rules. For example, 'search_threads' accepts a 'query' parameter described only as 'Gmail search query string' without explaining what syntax is valid or what the response contains. The 'analyze_thread' tool delegates to OpenAI but provides no documentation about the analysis_type enum values or what output to expect. Overall, this reads as a functional prototype rather than a production-grade tool definition.
Analyze a Gmail thread with AI to extract key information
Create a draft message in Gmail, optionally replying to a thread
Download and decode a Gmail attachment
Retrieve existing drafts for a specific thread
Read a specific Gmail message with full content and attachments
Search Gmail threads based on a query
Send a draft message in Gmail
NO OUTPUT SCHEMAS DOCUMENTED. All 7 tools lack documented return types, making it impossible for LLMs to plan downstream operations, extract data correctly, or pass results to chaining tools.
MISSING ERROR GUIDANCE. No tool describes what errors are possible, whether they are retryable, or what the LLM should do next. E.g., send_draft offers no guidance for 'Draft not found' or 'Sending failed' cases.
NO CONFIRMATION MECHANISM FOR DESTRUCTIVE OPERATIONS. send_draft and create_draft modify state irreversibly but lack dry-run or confirmation support. Agents can trigger actions without user approval.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 40 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
analyze_thread delegates to OpenAI API but this dependency is NOT documented. No mention of cost, latency, API key setup, or analysis_type enum validation. The analysis_type param is a free-form string, inviting hallucination.
DESCRIPTIONS ARE TOO BRIEF. 5 of 7 tools have descriptions under 60 characters with minimal context. Current descriptions force LLMs to guess.
PARAMETER CONSTRAINTS MISSING. 'max_results' has no min/max bounds. 'output_format' and 'analysis_type' are free-form strings instead of formal enums. Body content format (HTML vs plaintext) is undocumented. Validation rules are scattered or absent.
NO PAGINATION DOCUMENTATION. search_threads and get_thread_drafts accept max_results but don't document pagination strategy (cursor, offset, limit), result structure, or total count. Large result sets may blow context windows.
MISSING CHAINING DOCUMENTATION. create_draft does not document whether it returns draft_id (needed by send_draft). search_threads does not document whether results include thread_id (needed by read_message, analyze_thread). Tool composition breaks without this info.
send_draft AND create_draft descriptions do NOT state that they modify state.