Moist demonstrates solid foundational quality with clear naming, reasonable descriptions, and explicit schema definitions for all 17 tools. However, several gaps prevent a higher score: (1) descriptions are sometimes generic and lack context about when to use a tool vs. alternatives; (2) error handling guidance is minimal, tools document what they do but not how to recover from failures; (3) no tool annotations (readOnlyHint/destructiveHint/idempotentHint) despite having mixed operation types; (4) parameter descriptions are present but sparse, especially for compound types like labelIds arrays and querystring syntax. The tool names are well-chosen (verb_noun pattern, clear distinctions between send_message/send_draft), and schemas are properly typed with JSON Schema. This is solid C+ / B- work, better than median community servers, but missing patterns that distinguish A-grade tools.
Missing tool annotations despite mixed operation types (READ_ONLY, WRITE, REVERSIBLE, DESTRUCTIVE). Tools like moist_delete_message and moist_delete_draft should carry destructiveHint:true; read-only tools should carry readOnlyHint:true. LLMs need these hints to reason about safety and irreversibility.
Error handling lacks recovery guidance. Descriptions do not state what to do if authentication fails, tokens expire, rate limits are hit, or a message ID is invalid. Tools should return actionable error messages and suggest next steps.
Add readOnlyHint and destructiveHint annotations to all tool definitions. Mark moist_delete_message, moist_delete_draft with destructiveHint:true; mark all read-only tools (moist_list_*, moist_get_*, moist_search, moist_auth_status) with readOnlyHint:true. This enables LLMs to reason about operation safety without reading descriptions.
Document output schemas for all tools in the tool registration. Example: moist_get_message should declare { messageId: string, from: string, to: string[], subject: string, body: string, timestamp: ISO8601, labelIds: string[] }. Schemas enable agents to extract chaining IDs and plan multi-step sequences.
Enhance parameter descriptions with explicit format guidance. For 'query' parameters, provide a link to Gmail search operators or embed examples: 'Gmail search query using from:, to:, subject:, has:attachment, is:unread. Example: "from:user@example.com subject:invoice" '.
Add error classification to tool descriptions. State: 'Returns 400 if messageId is invalid (suggest moist_search to find the correct ID); 401 if authentication failed (call moist_auth_status to check); 429 if rate limited (wait 60 seconds and retry).' This guides agent recovery.
Clarify tool selection boundaries in descriptions. For moist_send_message vs moist_send_draft: 'Use moist_create_draft to prepare an email without sending; call moist_send_draft later. Use moist_send_message to send immediately.' This prevents LLM confusion.
Add mutual-exclusivity and required-together constraints to moist_modify_labels: 'At least one of addLabelIds or removeLabelIds must be provided. Both can be used in a single call to add and remove labels simultaneously.' This prevents ambiguous requests.
Parameter descriptions are minimal, especially for complex inputs. Example: 'query' in moist_search and moist_list_messages says 'Gmail search query syntax' but does not clarify expected format (from:, to:, subject:, has:attachment, is:unread, label:). LLMs need explicit examples or documentation links.
Tool descriptions do not clarify when to choose one tool over a similar tool. Example: moist_send_message and moist_send_draft both send emails, the descriptions should state when to use each (draft first, then send; vs. immediate send).
Composition issue: moist_modify_labels requires both messageId and addLabelIds/removeLabelIds, but it's unclear if both operations are required or if one can be omitted. Parameter descriptions should state mutual-exclusivity or required-together constraints.
Output schemas are not documented in the tool definitions. The eval cannot verify what fields are returned by moist_get_message, moist_get_thread, or moist_list_messages. LLMs need to know the shape of responses to chain tools and extract values for follow-up calls.
For paginated results (moist_list_messages, moist_list_threads, moist_list_drafts, moist_search), add a 'nextPageToken' or similar field to the output schema and document it in the tool description. This guides agents to fetch all results across multiple pages.
Add dependency hints to tool descriptions. Example: moist_send_message 'Requires a valid Gmail account (call moist_auth_status to verify). Recipients must be valid email addresses.' This prevents wasted calls with invalid preconditions.
Expand moist_auth_status description to clarify what 'scopes' and 'token expiry' mean and what to do if either is invalid. Example: 'Returns current email, authorized scopes (read, send, modify), and token expiry time. If expiry is soon, tokens will auto-refresh on next API call. If scopes are missing, re-authenticate with moist_auth_logout then moist_auth_status.'
Document the behavior of moist_trash_message vs moist_delete_message more explicitly: 'trash moves messages to the Trash folder (recoverable for ~30 days); delete permanently removes them (unrecoverable). Use trash for safety; delete only when certain.'