An MCP server for Gmail operations including listing, reading, searching, and sending emails with AI-powered reply suggestions
The server has 5 tools with reasonable naming (verb_noun pattern) and descriptions. However, there are several significant gaps: (1) Output schemas are not documented anywhere, the code returns formatted strings (markdown), but no structured schema is declared for responses, violating the 'Document the output schema' pattern. (2) Parameter descriptions exist but are minimal (1 sentence each) and lack actionable format constraints or dependency documentation. (3) No error handling guidance, tools do not document failure modes, recovery steps, or what could go wrong. (4) Security: the server loads Gmail credentials but the code does not show how secrets are injected; presence of 'gmail_client = GmailClient()' at module level suggests credentials are read from environment, which is correct, but no explicit documentation. (5) Input validation is minimal, no handling of malformed gmail_id strings or email addresses. (6) The 'gmail_suggest_reply' tool uses OpenAI (visible in pyproject.toml dependency on 'openai>=2.14.0') but its behavior is opaque; the code calls 'gmail_client.generate_reply_suggestions()' but this function is not shown, making assessment incomplete. (7) gmail_send_email is destructive (WRITE risk) but has no confirmation step or dry-run mode. Per-tool assessment below.
List recent Gmail messages with subject, sender, and preview Returns a formatted string of emails with key details like who sent it, the id of the email, when it was sent, and a preview of the content
Reads an email given its id and returns its sender, subject, body, and date Returns a formatted string of the email with sender, subject, body, and date
Search gmail messages using Gmail query syntax Returns a formatted string of emails that satisfy the query with the same format as gmail_list_messages. Supports queries like 'from:email@example.com', 'subject:meeting'
Send an email or reply to a thread Can send a new email or reply to an existing conversaion by providing thread_id
Generate smart reply suggestions for an email Returns a formatted string of three suggestions with different tones: casual, professional and detailed
Output schemas are not documented. All five tools return formatted strings (markdown) with no declared structure. LLMs cannot infer what fields to extract or how to chain results to other tools.
No error handling guidance. Tools do not document failure modes, recovery paths, or what could go wrong. For example, gmail_read_email does not say what happens if the email ID is not found or invalid. No actionable error messages documented.
gmail_send_email is destructive (WRITE) and non-idempotent but has no confirmation step, dry-run mode, or pre-flight validation. An agent could accidentally send duplicate emails or to wrong recipients without safeguards.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
gmail_suggest_reply depends on an LLM (OpenAI, visible in dependencies) but its behavior, cost, reliability, and failure modes are undocumented. Underlying function gmail_client.generate_reply_suggestions() is not shown in source code, making full assessment impossible.
Parameter descriptions are minimal (1-2 sentences) and lack actionable format constraints. For example, 'query' parameter in gmail_search_messages says 'Gmail search query' but does not explain what syntax is valid, what happens if malformed, or document the default empty string behavior (likely returns all messages, risking context overload).
No input validation or sanitization documented. LLMs could pass malformed email addresses, excessively long subject/body, or invalid gmail_id values without clear error guidance. Agents should be protected from malicious or accidental bad input.
No pagination or result limiting documented for gmail_list_messages and gmail_search_messages beyond the max_results parameter. If a search returns many results, returning all of them could exhaust context. No guidance on whether results are ordered, what happens with large result sets, or how to paginate through results.
No permission/scope documentation. Tools do not declare what Gmail API scopes they require (e.g., 'gmail.readonly' vs 'gmail.modify'). This makes it impossible to audit or configure least-privilege access.