An MCP server for Gmail email operations, enabling email reading and draft reply creation
Two tools are explicitly registered in handlers.py with clear names, descriptions, and JSON schemas. Naming follows verb_noun convention (get_unread_emails, create_draft_reply). Descriptions are present and moderately detailed (150-180 chars). Input schemas are properly typed with required fields marked. However, there are significant gaps: (1) no output schemas documented anywhere, LLMs cannot plan downstream tool calls or know what fields to expect; (2) parameters lack critical constraints (limit has no max bound stated despite validation; thread_id and reply_body have no format/length guidance); (3) no idempotency guidance for create_draft_reply, which modifies state; (4) error handling returns plain text responses with no classification (retryable vs fatal) or recovery guidance; (5) no mention of pagination, though fetch_unread_emails accepts a limit parameter suggesting potential for large result sets; (6) no field-name mapping between get_unread_emails output (returns thread_id, id, from, subject, date, body) and create_draft_reply input (expects thread_id), this works but is not explicitly verified in schema.
Create a draft reply to an email thread. The draft will appear in the original email thread in Gmail.
Fetch unread emails from Gmail inbox. Returns sender, subject, date, body, message ID, and thread ID for each email.
No output schemas documented for either tool. Handlers return TextContent with formatted strings, but LLMs cannot infer the structure of returned data (field names, types, presence of IDs for chaining). This violates the requirement that 'LLMs need to know what fields to expect so they can plan downstream tool calls.'
Parameter constraints are partially hidden from the schema. The 'limit' parameter in get_unread_emails has no maximum bound declared in the JSON schema (code validates ≤ 100, but schema shows only default: 10). Clients and LLMs cannot see this constraint without reading implementation code.
create_draft_reply lacks side-effect and idempotency documentation. Description does not state this is a write operation or that repeated calls with identical parameters will create duplicate drafts. Agents need explicit guidance on whether operations are safe to retry.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Error responses are plain text strings with no classification or recovery guidance. The code catches HttpError with 404, 401, 403, 429 status codes and returns human-readable messages, but does not indicate to the LLM which errors are retryable (429), which require user action (401 auth), and which indicate permanent failure (404). LLMs cannot act on unclassified errors.
Parameter descriptions lack detail. 'reply_body' is described as 'The body text of your reply' with no guidance on format (plain text, markdown, HTML?), max length, or character encoding. Per the rubric, descriptions should state 'expected format, range, and allowed values.'
No pagination or result-limiting guidance. Although get_unread_emails accepts a 'limit' parameter, the description does not state the default (10), max (100), or that results are paginated. Per the rubric, tools returning lists should 'accept page/offset and limit parameters and return a total count or next_cursor.'