Local MCP server for safe Gmail cleanup (preview, unsubscribe, trash)
The server defines 2 tools with explicit schemas and clear action verbs (preview_cleanup, execute_cleanup). Tool descriptions are substantive (90+ chars) and explain WHAT and WHEN. However, several critical gaps reduce the score: (1) Input schemas lack descriptions for 3+ parameters ('planId', 'confirmPhrase', 'entryIndices', 'maxMessages', 'perEntryMaxMessages'); (2) Output schemas are entirely undocumented, responses are JSON-stringified generic objects with no field definitions; (3) No error handling guidance in descriptions (what if preview fails? what if planId is invalid?); (4) No documented parameter relationships or constraints beyond minima; (5) No idempotency guarantees for destructive execute_cleanup. Both tools are registered with names, types, and input constraints present, which prevents F-level scores. The naming follows verb_object convention well. Descriptions could be LLM-optimized by including dependency hints and recovery guidance.
Execute a previously previewed cleanup plan by moving matched messages to Trash and attempting automatic unsubscribe for list-unsubscribe links. Requires exact confirmPhrase.
Preview a cleanup plan from a sender/query list: counts + small samples per entry + unsubscribe hints. Returns planId + confirmPhrase.
Output schemas completely undocumented. Both tools return JSON objects with no field definitions. LLMs cannot plan downstream calls or extract structured data reliably.
Input parameter descriptions missing for execute_cleanup. 'planId', 'confirmPhrase', 'entryIndices', 'maxMessages', 'perEntryMaxMessages' have no descriptions. LLMs cannot infer what these parameters do or how to populate them.
No error handling guidance in tool descriptions. If a preview fails (e.g., invalid file path), what should the agent do? If planId is invalid, what's the recovery path? Descriptions are silent on failure modes.
Mutual exclusivity between 'filePath' and 'listText' in preview_cleanup is not documented. LLMs may pass both or neither, causing ambiguous tool behavior. Descriptions must state: 'Provide either filePath OR listText, not both.'
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 56 | - | v1 |
No documented constraints on samplePerEntry, maxMessages, perEntryMaxMessages beyond minima. What are the maxima? How do they interact with YAML config caps? Absence of documentation invites invalid LLM inputs.
No documented idempotency guarantee for execute_cleanup. Is it safe to retry with the same planId and confirmPhrase? Or will it double-delete messages? Destructive operations must declare idempotency explicitly.
No dependency hints in tool descriptions. If filePath validation fails, should the agent call a discovery tool first? If planId expires, what's the window? Descriptions are missing action-guiding context.