MCP server for Outlook email access via Microsoft Graph API
The server defines 6 tools with full JSON Schema input specifications and descriptions. However, critical gaps significantly reduce quality: (1) Tool naming violates single-responsibility principle, 'set_access_token' is a configuration concern unrelated to email operations, and should not be a callable tool. (2) Parameter descriptions are present but generic and lack actionable format guidance (e.g., 'OData filter query' without examples or validation rules). (3) Output schemas are NOT documented anywhere in the source, responses are inferred from code, violating the requirement that 'document the output schema.' (4) Error handling is minimal, no recovery guidance, no categorization of errors as retryable vs fatal. (5) No tool annotations (readOnlyHint, destructiveHint) despite having both read and destructive operations. (6) The 'top' and 'skip' parameters lack numeric constraints (min/max). Overall, the server has the structural foundation (names, descriptions, input schemas) but lacks depth in output documentation, error design, and LLM-guidance patterns.
Delete an email
Get detailed information about a specific email
List emails from Outlook inbox
Mark an email as read or unread
Search for emails using a search query
Set the Microsoft Graph API access token for Outlook access
Output schemas not documented. Responses are inferred from code (list_emails returns [id, subject, from, received, preview, isRead, hasAttachments, importance]; get_email returns {id, subject, from, received, body, bodyType, isRead, hasAttachments, importance}). The MCP tool definitions in src/index.ts do NOT include outputSchema, violating the requirement that 'document the output schema so LLMs know what fields to expect.'
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Tools like delete_email (destructive), mark_as_read (write, idempotent), and list_emails (read-only) lack these hints. LLMs cannot infer operation safety class without explicit hints, increasing risk of unsafe tool selection.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
delete_email lacks confirmation/dry-run mechanism. Irreversible operation with no user confirmation step, violating pattern:confirmation-request. A malformed agent plan could delete all inbox contents without recourse.
set_access_token should not be a tool. Authentication/configuration are setup concerns, not chat actions. Exposing tokens as tool parameters (even though the implementation requests them as params) violates pattern:secret-injection. Move token management to server initialization or environment variables.
Numeric parameters lack constraints. 'top' (default 10) has no max stated; 'skip' has no bounds. Unbounded numbers let LLMs pass absurd values (e.g., top=1000000) that break APIs or exhaust rate limits. Should specify: top 1 - 100, skip 0 - 9999.
Minimal error handling. Error responses are generic ('error: message') with no recovery guidance. Pattern:recovery-guide requires errors to say 'what to do next.' E.g., if message_id not found, should suggest 'Try list_emails() first to find valid IDs.'
Parameter descriptions lack actionable format/syntax guidance. 'OData filter query (e.g., "isRead eq false")' gives one example but no formal constraint. 'Search query to find emails' does not clarify scope (subject? body? from? all?). Should specify: 'OData syntax with operators: eq, ne, gt, lt, and, or (see https://... for examples)' and 'Search term(s) matched against subject and body.'
list_emails and search_emails overlap in intent. search_emails with 'query' parameter is essentially list_emails with a 'search' parameter. Redundant tools force LLMs to reason about which to use, wasting tokens and risking wrong selection. Consolidate into one tool with optional search/filter parameters.
No pagination metadata returned. list_emails and search_emails should return 'total_count' and 'has_more' or 'next_cursor' so agents know whether more results exist and can plan multi-page fetches. Currently, no signal whether 10 results is all available or if 1000 exist.