MCP server giving AI assistants full access to Apple Mail - read, search, compose, organize & analyze emails
The Apple Mail MCP server defines 20 tools with mostly complete schemas and descriptions. Tool naming is clear and verb-first (list_, get_, compose_, search_, etc.), following standard conventions. However, there are significant gaps: (1) several tools lack parameter descriptions or have minimal descriptions under 20 characters, which violates the baseline that 100% of A+ tools have full parameter documentation; (2) output schemas are not explicitly documented in the source, forcing inference; (3) error handling guidance is absent, tools do not indicate what to do if an email is not found or a search fails; (4) several parameters accept free-form strings where enums would be safer (e.g., flag_color in flag_email, analysis_type in analyze_email_patterns, scope in get_statistics). The server avoids storing credentials as parameters (good), but lacks recovery hints and actionable error messages. Average tool score: 62/100.
Analyze email patterns and provide insights for productivity.
Batch process multiple emails with actions like move, flag, or mark as read.
Compose and send a new email.
Create a draft email without sending.
Delete an email by moving it to Trash.
Identify and categorize newsletter emails in an account.
Flag or unflag an email with a color.
Forward an email to new recipients.
Missing or incomplete parameter descriptions. Many tools have parameters with minimal or absent descriptions (e.g., flag_email's 'flag_color' is documented as 'Color to flag with (red, orange, yellow, green, blue, purple, gray, or null to unflag)' but no guidance on which emails use which colors or what happens if an invalid color is passed). Rubric baseline: 100% of A+ tools have descriptions for EVERY parameter.
Free-form string parameters where enums should exist. Tools like flag_email accept 'flag_color' as a bare string; analyze_email_patterns accepts 'analysis_type' as a bare string; get_statistics accepts 'scope' as a bare string. LLMs will hallucinate invalid values. These should be enum constraints (flag_color: ["red", "orange", "yellow", "green", "blue", "purple", "gray", null], analysis_type: ["sender_frequency", "response_time", "volume_trends"], scope: ["account_overview", "sender_stats", "mailbox_breakdown"]).
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 26 | - | v1 |
Search for emails and retrieve full content.
Get unread counts per mailbox for one account or all accounts. When summary_only=True, returns only per-account inbox unread totals (replaces the former get_unread_count tool).
Get the raw RFC 822 source of an email.
Get comprehensive email statistics and analytics.
List attachments for emails matching a subject keyword.
List all emails from inbox across all accounts or a specific account. Replaces the former get_recent_emails tool — use account + max_emails to get recent emails from a single account.
List mailboxes (folders) for one account or all accounts.
Mark an email as read or unread.
Move an email to a specific mailbox.
Reply to an email matching a subject keyword.
Search emails by multiple criteria with flexible filters.
Automatically organize emails into mailboxes based on senders or keywords.
No output schemas documented. The source code shows parameter schemas (input) but does NOT show what fields are returned or what the output structure looks like. Rubric: 'Document the output schema. LLMs need to know what fields to expect so they can plan downstream tool calls.' For example, list_inbox_emails's output is not shown, are emails returned as [{'subject': ..., 'from': ..., 'date': ...}]? What fields does get_statistics return? Without documented output schemas, agents cannot chain tools effectively.
No error handling guidance or recovery hints. Tools do not document what happens when an email is not found, a search returns no results, or an invalid mailbox is specified. Rubric: 'Error responses must tell the LLM what to do next.' For example, move_email should state 'If mailbox not found, call list_mailboxes() to see available mailboxes' or 'If email not found, try search_emails() with a broader keyword.'
No idempotency guarantees documented. compose_email, reply_to_email, and forward_email are WRITE operations that could create duplicates on retry. Rubric: 'Make tools produce the same result on repeated calls with the same input. Agents retry on ambiguous failures, non-idempotent tools risk duplicate side effects.' Should document retry-safety or provide idempotency tokens.
Destructive operations lack confirmation or dry-run mode. delete_email and batch_process_emails with action='delete' can permanently remove emails. Rubric: 'Irreversible operations (delete, send, publish) should support a dry-run or confirmation step.' Neither tool offers a confirmation_required or dry_run parameter to prevent accidental deletion.
Result limit enforcement unclear. list_inbox_emails accepts max_emails (unbounded if 0), search_emails accepts max_results (unbounded?), get_statistics has no limit. Rubric: 'Even if the API allows returning thousands of items, cap results at a reasonable limit (e.g. 20-50) and offer pagination.' No evidence that results are capped or paginated.
Parameter naming inconsistency: some tools use 'subject_keyword' (singular), others 'subject_keywords' (plural). This inconsistency forces agents to remember which tool accepts which form. Rubric: 'When multiple tools operate on the same resource, their names must make the distinction obvious.' Standardize on one convention.