This Gmail MCP server has 6 tools with basic schemas and descriptions, but suffers from significant quality gaps. All tool descriptions are present and reasonable in length (35-95 chars, within baseline 34-392), but parameter descriptions are minimal or missing entirely. Most critically, output schemas are completely undocumented, the rubric requires documented return types for A+ tools (100% baseline). Input schemas are visible and mostly complete, but lack proper format/constraint documentation. Error handling is minimal: most functions return generic dicts without guidance for LLM recovery. Tool names are clear (verb-noun pattern), but the email-centric domain leaves composition fragmented, no batch send, no template support, and read-email lacks pagination hints. The server implements logging and prompts but relies on deprecated patterns (server-side logging) and lacks current-spec features like tool annotations, error reporting, or structured output hints.
Retrieves unread messages from mailbox. Returns list of message IDs in key 'id'.
Marks email as read given ID.
Opens email in browser given ID.
Retrieves email contents including to, from, subject, and contents.
Creates and sends an email message
Moves email to trash given ID.
Output schemas completely undocumented for all 6 tools. Callers have no typed contract for what fields/types to expect.
Parameter descriptions are minimal (16-36 chars) and lack format/constraint details. Rubric baseline is 72 chars and expects 'The X parameter (constraints)' format.
get-unread-emails lacks pagination (no limit, offset, or page_size params). Tool could return thousands of emails, blowing context window. Rubric requires paginated-result pattern for list tools.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 17 | - | v1 |
trash-email (destructive operation) lacks confirm-before-execute or dry-run pattern. Description does not state side effects or reversibility despite REVERSIBLE risk marking.
Tool descriptions are action-focused but lack WHEN/WHY guidance, prerequisites, and differentiation from similar tools. Descriptions are 31-95 chars but below LLM-optimized baseline of 50-200 chars with full context.
open-email opens system browser, unusual and brittle for server MCP context. Fails silently in headless environments. No error handling documented.
No error handling guidance. Tools return generic dicts or strings without recovery hints. Rubric requires 'User not found. Try search_users()' style messages.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. Current spec (2026-07-28) recommends these for agent planning.
Parameter 'recipient_id' in send-email is poorly named. 'recipient_id' suggests an opaque system ID, but it actually accepts email addresses. Should be 'recipient_email' per name-clarity pattern.