The Gmail MCP server defines 22 tools with generally consistent structure. Most tools have descriptions and basic input schemas. However, there are significant gaps: descriptions are often minimal (10-50 chars), parameters lack detailed validation guidance, output schemas are not formally documented, error handling is absent, and the implementation lacks the structured responses and guidance patterns expected of production tools. The server does NOT use tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite having clear risk categories for each tool, a missed opportunity for protocol compliance. No output schemas are documented in the code, and error responses are generic ('Failed to...' or not shown). The tool composition is sound (each tool does one thing), but the descriptions and schemas need strengthening.
Output schemas not documented. Tools return free-text formatted strings (e.g., 'Found X emails with label Y...\nID: ...\nFrom: ...') instead of structured JSON objects. LLMs cannot parse these efficiently and must extract fields via string parsing, increasing hallucination risk and token waste.
Destructive operations (delete_email, delete_emails, delete_draft, delete_gmail_label) lack confirmation/dry-run support and explicit error guidance. No recovery mechanism if an agent mistakenly deletes content. Error responses are generic ('Failed to delete email with ID {email_id}...').
Convert all responses from free-text strings to structured JSON objects with typed fields. E.g., list_emails should return {"emails": [{"id": "...", "from": "...", "subject": "...", "date": "...", "read": true}], "total": 42, "next_cursor": null} instead of a formatted string.
Add tool annotations to the FastMCP decorator: @mcp.tool(annotations={"type": "read"}) for read-only, @mcp.tool(annotations={"type": "write"}) for write, @mcp.tool(annotations={"type": "destructive"}) for destructive operations. Align with the Risk: field already documented.
Implement confirmation/dry-run for destructive operations. Add an optional 'confirm' parameter to delete_email, delete_emails, delete_draft, delete_gmail_label. When false, return a preview of what will be deleted and ask the agent to confirm before executing.
Expand tool descriptions to 100-200 characters with context: 'When should I use this vs the similar tool? What does it return? Any prerequisites?' E.g., 'Delete a single email by ID. Permanently removes the message; cannot be undone. Use delete_emails for batch deletion. Returns success confirmation.'
Add validation constraints to parameter descriptions. For email_id, add: 'The unique Gmail message ID (alphanumeric string, e.g., 18eb2d74d68f96c5). Obtain via list_emails or search_emails.'
Implement structured error responses with actionable guidance. When send_email fails, return: {"success": false, "error": "Invalid recipient address: 'alice@' is incomplete.", "recoverySteps": ["Verify the email address is complete and valid.", "Retry with a correct address."]}
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 0 points across a rubric change (v1 → v2)
49/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
49
<=2025-11-25
v2
2026-03-09
F
49
-
v1
read onlyauthsource verified67/100
Download and retrieve an attachment from an email.
get_draftread onlyauthsource verified70/100
Get the full content of a specific draft.
get_emailread onlyauthsource verified70/100
Get the full content of a specific email.
get_gmail_labelread onlyauthsource verified67/100
Get details about a specific Gmail label using its ID.
Tool annotations missing. Despite having clear risk semantics (READ_ONLY, WRITE, DESTRUCTIVE, REVERSIBLE), tools do not use MCP tool annotations (readOnlyHint, destructiveHint, idempotentHint). This prevents clients from applying appropriate safety guardrails.
Parameter descriptions are minimal and lack validation guidance. For example, 'email_id' has description 'The ID of the email to retrieve' but provides no format hint, length constraints, or examples. Similar gaps across all ID parameters (label_id, draft_id, message_id, attachment_id).
Tool descriptions are significantly shorter than baseline (averaging ~80 chars vs 194 chars in A+ tools). Many are purely mechanical ('Delete a single email', 'Get details about a specific Gmail label') without explaining when to use them, dependencies, or what the LLM should expect. This reduces LLM selection confidence when multiple tools could apply.
Error handling is missing or generic. Tools return strings like 'Failed to send email. Please check the logs for details.' This provides no actionable guidance for the LLM: should it retry? Is the error transient or permanent? What should the user do?
No pagination or result-limiting guidance in list tools. 'list_emails', 'search_emails', 'list_drafts' accept max_results but do not document whether results can exceed the limit, whether pagination is supported, or what to do if more results exist. Users may expect automatic pagination that doesn't happen.
Parameter naming inconsistencies. 'send_email' uses 'to' (recipient), but 'update_draft' also uses 'to'. Tools use 'email_id', 'draft_id', 'label_id', 'message_id', 'attachment_id', mixing naming conventions. Would be clearer to standardize (e.g., always use 'message_id' internally, expose 'email_id' for user-facing tools).
Document pagination for list tools. Add fields to response: 'total_count', 'next_cursor' (or page_token). If max_results < total, explain how to fetch the next batch. E.g., 'To get more results, pass next_cursor to the next call.'
Standardize parameter naming: use 'message_id' internally (Gmail API standard) but expose 'email_id' in user-facing tools for clarity. Document this mapping.
Add idempotency keys or idempotent operation pattern for send_email, reply_to_email, create_draft to prevent duplicate sends/creates if the agent retries.
Include chaining IDs in responses. When search_emails returns messages, include all fields downstream tools need (e.g., label_ids if add_label_to_email might follow).