Multiple MCP servers for Google services (Gmail, Google Drive, Weather)
This server exhibits significant definition quality gaps across all tools. While all 7 tools have basic descriptions and input schemas, they fall short of production-grade standards. Tool names use proper verb_noun patterns (search, list-, get-), but descriptions are often generic (10-50 chars), well below the 194-char production baseline. Parameter descriptions vary widely, some missing critical context (e.g., 'labelIds' in list-emails lacks explanation of what 'INBOX', 'UNREAD', 'SENT' mean). No tool declares output schemas, error handling guidance, or explains when to use similar tools (e.g., search-emails vs list-emails distinction). The Gmail and Google Drive tools handle sensitive OAuth credentials but expose no security-related documentation. Tools lack idempotence declarations, pagination guidance, or recovery instructions for failure cases.
Get weather alerts for a state
Get a specific email by ID
Get weather forecast for a location
Get all Gmail labels
Get emails in inbox
Search for files in Google Drive
Search for emails using Gmail search syntax
Descriptions are too short and lack context. 'Get emails in inbox' (25 chars) and 'Get weather alerts for a state' (32 chars) fall below the 194-char production baseline. Descriptions do not explain WHEN to use the tool, WHAT it returns, or WHEN to call similar tools (e.g., why use list-emails vs search-emails).
Output schemas are not documented. No tool declares what fields the response will contain, what data types they are, or what the agent should expect. This forces LLMs to guess structure and makes chaining tools difficult.
Parameter descriptions lack actionable constraints. 'labelIds' in list-emails says 'Label IDs to filter by (e.g., INBOX, UNREAD, SENT)' but does not explain what these values are, where to get them, or how to call get-labels first. 'state' in get-alerts says 'Two-letter state code (e.g. CA, NY)' but lacks constraint validation text.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
No error handling guidance. Tools do not describe what errors are possible, how to recover, or what the LLM should do next. E.g., search-emails does not explain what happens if the Gmail API rate-limits the request, or get-forecast if coordinates are invalid.
Overlapping tools create ambiguity. Both 'list-emails' and 'search-emails' retrieve emails; descriptions do not explain when to use each. 'list-emails' filters by sender and label; 'search-emails' uses Gmail query syntax. An LLM cannot infer the distinction without explicit guidance.
No pagination or result limits stated in descriptions. Tools like 'list-emails' and 'search-emails' accept 'maxResults' but do not declare a default, maximum, or explain what happens if results exceed the limit. Production baselines cap defaults at 20-50 items.
get-email lacks context. It accepts only an 'emailId' but does not explain where to get email IDs (from list-emails or search-emails), what format the ID is, or what the response contains.
OAuth credentials and authentication are not documented. The server uses Google Cloud authentication but tool definitions do not mention scope requirements, permissions, or what scopes the agent needs. Production tools declare scope-declaration.