MCP server for Google Gmail, Drive, and Classroom integration with vector embeddings via Milvus
This MCP server exposes 7 tools for Gmail and Google Drive access with basic descriptions and parameter schemas. However, it suffers from critical gaps in output documentation, error handling guidance, and parameter completeness. Tool descriptions are present but minimal (11-73 chars), falling below the baseline average of 194 chars. No output schemas are documented anywhere in the code, LLMs cannot plan downstream calls or extract required fields. Error handling returns raw HttpError exceptions instead of actionable recovery guidance. Parameters lack descriptions in most cases and have no constraints or validation hints. The tools are read-only and relatively straightforward, which prevents catastrophic failures, but the server would struggle in multi-step agent workflows due to missing chaining IDs and undocumented return structures.
Fetch recently read (non-unread) emails from the Gmail inbox.
Fetch emails from your Gmail spam folder.
Fetch unread emails from your Gmail inbox.
List all your Google Classroom courses.
List recent files in your Google Drive.
Read the full content of a given email by ID.
Read the content of a Google Drive file by file_id. Supports Google Docs, plain text, and PDFs.
No output schemas documented. Tools return dictionaries with fields like 'id', 'subject', 'from', 'body', but the structure is not declared. LLMs cannot predict return types or plan multi-step workflows. For example, get_unread_emails returns a list of dicts with {id, subject, from, snippet}, but this is inferred from code only, not declared in tool metadata.
Error handling is non-actionable. All tools catch HttpError and return bare error dicts like {"error": str(e)}. LLMs receive raw API error messages with no guidance on what to do next. Per the rubric, errors must tell the agent what to try: 'User not found. Try search_users() with a partial name.' This server provides no recovery path.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Descriptions are minimal (11 - 73 characters). Baseline average is 194 chars. Examples: 'Fetch unread emails from your Gmail inbox.' (43 chars), 'Fetch recently read (non-unread) emails from the Gmail inbox.' (60 chars). These lack context on WHEN to use the tool, WHAT it returns, and how it differs from similar tools (e.g., get_unread_emails vs get_read_emails). Per the rubric, descriptions should guide tool selection for LLMs operating in multi-tool scenarios.
Parameters lack descriptions. The 'limit' parameter appears on get_unread_emails, get_read_emails, get_spam_emails, and list_my_drive_files. Each has a description ('Maximum number of ... to fetch'). However, 'email_id' (read_email) and 'file_id' (read_file_content) have descriptions, but no constraints are documented. Per the rubric, format expectations and valid ranges should be explicit: 'The email ID (RFC 5321 format, e.g. abc123def456)' instead of just 'The unique identifier of the email to read.' Additionally, list_courses() accepts NO parameters, unclear what filters or defaults apply.
Missing chaining IDs in responses. get_unread_emails and get_read_emails return {id, subject, from, snippet}. The 'id' field is returned, which is good. However, if an agent wants to read an email after listing, it needs the ID, which it gets. But for Google Drive, list_my_drive_files is truncated in the source, so the return structure is unknown. If file IDs are not returned, read_file_content cannot be chained. Per the rubric, every response must include IDs and references that downstream tools accept, or agent workflows break.
No pagination support or result limits enforced in descriptions. get_read_emails defaults to limit=50. The description says 'Maximum number of read emails to fetch' but does not specify a hard cap or warn that large limits may exhaust context. Per the rubric, tools returning lists should document pagination behavior and enforce reasonable limits. The baseline expectation for lists is 20 - 50 items max; 50 is on the edge but acceptable if documented as a hard cap.
Credentials hardcoded with placeholder paths. servar.py contains: GMAIL_CREDS_FILE_PATH = 'absolute_path//to//client_creds.json' and DRIVE_CREDS_FILE_PATH = 'absolute_path//to//client_creds.json'. These are not actual paths and will fail at runtime. While not a definition-quality issue per se, it indicates the server is not ready for deployment. Additionally, Credentials are initialized at module load time (gmail_service = get_gmail_service()), blocking the server startup if auth fails. Per the security pattern, credentials should never leak into tool parameters, but hardcoding paths and initializing at startup is also problematic.
Tool definitions are not visible in structured form in the code. The @mcp.tool() decorator is used, but the actual registration and schema definition relies on FastMCP's automatic inference from function signatures. Per the hard scoring rule 'If you cannot see the actual tool definition in the source (only inferred), cap that tool's overall at 50.' The input schemas are inferred from Python function signatures (e.g., def get_unread_emails(limit: int = 5)), which FastMCP likely converts to JSON Schema automatically. However, the output schemas are completely absent, no return type hints, no structured response documentation. This forces the cap on schema scores to ≤35 per tool.