MCP server providing Gmail integration tools for reading, composing, and managing email threads
This server implements 4 Gmail-related tools with explicit tool registration and Zod schemas. However, it suffers from significant definition quality gaps that limit LLM reliability: (1) Tool names lack clear action verbs and are inconsistent in style (gmail-auth vs gmail-thread-snippets); (2) Descriptions are present but generic and lack actionable recovery guidance; (3) Parameter descriptions are minimal; (4) Output schemas are undocumented, the server returns plain text in a 'content' array, but LLMs have no formal specification of what fields/structure to expect; (5) Error handling is basic and does not guide the LLM on recovery steps. The server does have positive signals: tool registration is explicit, input schemas use Zod with type annotations and basic constraints (min/max for days, email format for address), and the code is syntactically sound. Per-tool average is 58, placing this in the lower-fair range.
Manually authenticate gmail when the tokens are expired or missing or there is an authentication error.
Creates draft responses to specified messages in Gmail.
Get the full messages for a single thread
Get my gmail snippets from threads from the last days. If there are authentication errors, try using the gmail-auth tool.
Tool names lack clear action verbs. 'gmail-auth', 'gmail-thread-snippets', 'gmail-thread-full' do not follow verb_noun pattern (e.g., 'authenticate_gmail', 'list_thread_snippets', 'get_thread_full'). LLMs struggle to infer intent from inconsistent naming.
Output schema is not formally documented. All tools return a plain JSON object with a 'content' array containing text. LLMs have no specification of what fields exist or their types, forcing them to infer structure from examples. This violates the documented-output-schema pattern.
Parameter descriptions are minimal or missing context. 'days' in gmail-thread-snippets has a constraint description ('The number days of emails to read (1 to 60)') but lacks guidance on when to use different values. 'threadId' in gmail-thread-full is generic ('The id of the thread we are trying to read'), no hint that it comes from gmail-thread-snippets.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Error messages do not guide recovery. When authentication fails, the tool returns generic strings like 'Cannot authenticate gmail now, please try again.' LLMs need actionable guidance: 'Authentication failed. Call gmail-auth tool to refresh credentials, then retry.'
No dry-run or confirmation pattern for destructive operations. 'gmail-draft-response' modifies Gmail state (creates a draft). The tool should offer a preview or confirmation step before executing to prevent accidental drafts.
Tool dependencies not documented. 'gmail-thread-full' requires a threadId from 'gmail-thread-snippets', but this is not explicitly stated in descriptions. Similarly, 'gmail-draft-response' expects parameters from 'gmail-thread-full', but no hint is provided.
No pagination support. If 'gmail-thread-snippets' returns many threads, there is no limit parameter or pagination mechanism visible. Large result sets will exhaust context or cause timeout.