Manage Microsoft Outlook tasks, calendar, email, contacts, and Teams via Graph API.
The Outpost MCP server provides 16 tools for Microsoft Outlook integration (Tasks, Calendar, Email). Strengths: all tools have explicit names with action verbs (task_add, cal_list, mail_read), all have non-empty descriptions (100-200 chars typical), and most input parameters include descriptions and type information. Weaknesses: parameter schemas lack formal enum constraints for enumerated values (e.g., priority: low|normal|high, show_as: free|tentative|busy|oof|workingElsewhere), no documented output schemas beyond return type hints, error handling is implicit (relies on GraphClient._get_client() raising RuntimeError), and no guidance on partial failures or recovery paths. Several parameters accept natural language input (e.g., due dates, start times) but validation logic is delegated to utility functions (parse_natural_date) without explicit constraint documentation in the parameter descriptions themselves. Overall structure is sound but lacks the precision and error recovery guidance expected of production tools.
Add a calendar event to Outlook.
Delete a calendar event from Outlook.
List calendar events from Outlook.
Get the next upcoming calendar event(s) from Outlook.
Get today's calendar events from Outlook.
Update an existing calendar event in Outlook.
List email messages from Outlook.
Enumerated parameters lack formal enum constraints in schema. Priority (low|normal|high), show_as (free|tentative|busy|oof|workingElsewhere), folder (inbox|sentitems|drafts|deleteditems) are documented in descriptions only, not as JSON Schema enums. LLMs cannot parse free-form text to detect valid options and may hallucinate invalid values.
Output schemas are not documented. Tools return list[dict] or dict but do not specify field names, types, or required fields. LLMs cannot plan downstream tool calls that depend on response structure (e.g., task_list returns task IDs, but output schema is not documented, downstream tools expecting task_id cannot verify they will receive it).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 58 | - | v1 |
Read a specific email message from Outlook.
Add a new task to Microsoft To Do.
Mark a task as completed in Microsoft To Do.
Delete a task from Microsoft To Do.
List tasks from Microsoft To Do.
List all task lists in Microsoft To Do.
Create a new task list in Microsoft To Do.
Delete a task list from Microsoft To Do.
Update an existing task in Microsoft To Do.
Error handling provides no recovery guidance. _get_client() raises RuntimeError if not authenticated, but the tool does not document this as a possible failure mode, does not classify it as retryable (it is not), or guide the LLM to run 'outpost setup'. Error responses from Graph API (invalid IDs, rate limits, permission denials) are not documented.
Destructive operations (task_delete, task_lists_delete, cal_delete) have no confirmation or dry-run capability. An LLM may call task_delete('abc123') based on a misunderstood user intent, irreversibly deleting the task. No undo tool is provided.
Natural language parameters (due, start, end dates) lack explicit format constraints in descriptions. The description 'Natural language ('tomorrow', 'next friday') or ISO ('2026-03-15')' provides examples but does not formally state the expected format(s), range validation, or what happens on parse errors. LLMs may pass ambiguous inputs like 'next week' (which week? which day?) without guidance.
Parameters with interdependencies are not documented. cal_add accepts both 'end' and 'duration' but the description does not explicitly state they are mutually exclusive or that end is preferred. If LLM passes both, which takes precedence? The implementation handles this, but the parameter descriptions do not.
No pagination support on list tools. task_list, cal_list, and mail_list do not accept limit/offset or return total counts. If a user has 500 tasks or emails, the entire result is returned, risking context window exhaustion and poor LLM reasoning. Baseline expects limit+offset or cursor pagination for list operations.
Tool descriptions do not clarify when defaults apply or what they are. cal_add defaults duration to 30 minutes if neither end nor duration is given, documented in the docstring but not in the parameter description. task_lists applies a 'default list' if list_name is omitted, which list is it? This forces LLMs to infer or guess.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present. The MCP spec allows tools to declare their side effects via annotations, enabling agents to reason about sequencing and rollback. All tools lack these hints.