MCP server for IMAP mailboxes: read, search, organise and draft mail — it cannot send
The server implements 11 tools with mostly complete input schemas and descriptions. Tool naming follows verb_noun conventions (list_mailboxes, search_messages, get_message) which is strong. Descriptions are present for all tools and most parameters. However, there are significant gaps: (1) output schemas are not documented, LLMs cannot know what fields to expect from tool results; (2) no tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear read vs. write semantics; (3) no pagination parameters on search_messages despite 'limit' being present (no offset/cursor for large result sets); (4) error handling guidance is absent, tools do not describe what errors can occur or how to recover; (5) parameter descriptions lack constraint details (e.g., what are valid IMAP search criteria? what character limits apply to mailbox names?). The server does show good practices in input schema completeness (most parameters have type and description) and security awareness (using mcp-approval for confirmations on destructive ops). Risk classification is declared but not surfaced in tool metadata. Strengths: schema presence, naming clarity, parameter descriptions. Weaknesses: output schemas, annotations, error recovery guidance, pagination design.
Append text to an existing draft message
Copy messages from one mailbox to another
Create a draft message in the Drafts mailbox
Mark messages as deleted
Download and extract text from message attachments (PDF, DOCX, XLSX, PPTX, ODT, ODS)
Get mailbox metadata and recent message counts
Get the full content of a message including headers, body, and attachments
List all mailboxes in the IMAP account
Output schemas are not documented. LLMs cannot infer what fields (mailbox_id, message_count, sender, subject, body, attachments, etc.) are returned from each tool. This forces agents to guess downstream parameter passing and breaks tool chaining.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are absent from tool definitions. The risk classification is tracked internally but not exposed in the MCP tool metadata. Tools like delete_messages should be marked destructiveHint=true, and read operations should be marked readOnlyHint=true.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 51 | 2026-07-28+ | v2 |
Move messages from one mailbox to another
Search for messages in a mailbox using IMAP search criteria
Set flags on messages (read, starred, spam, etc.)
No pagination design for search_messages. The 'limit' parameter is present but there is no offset, cursor, or next_page_token to handle large result sets. A search that returns 1000+ messages will blow the context window. Baseline: search tools should support limit (1 - 100) and offset or cursor.
Parameter constraints are not documented in descriptions. 'IMAP search criteria' (search_messages) lacks examples or enumeration of valid values. 'Mailbox name' (get_mailbox) does not specify character limits, special characters, or case sensitivity. LLMs cannot generate valid input without this guidance.
Error handling guidance is absent. Tools do not describe what errors are possible (e.g., 'mailbox not found', 'connection timeout', 'invalid IMAP criteria') or how to recover. Baseline: every tool should guide LLMs on retryability and next steps.
Flag names in set_flags are not constrained to a known enum. The parameter description says 'Flag names to set' but does not list valid options (e.g., seen, flagged, deleted, spam, answered). LLMs may pass invalid flag names, causing silent failures.
Tool chaining data is incomplete. get_attachments and other tools accept 'partIds' but it is unclear what data structures previous calls return to populate this. If search_messages does not return attachment metadata with part IDs, the LLM cannot chain get_attachments.
Confirmation workflow for destructive operations is implemented (mcp-approval integration visible in server.ts) but not exposed in tool descriptions. LLMs do not know that delete_messages requires approval, leading to surprises when the tool returns a confirmation request instead of executing.