A Python-based spam detection agent that uses LangGraph workflow and OpenAI API to analyze and label Gmail emails
This server has critical definition quality gaps across all dimensions. Tool schemas are partially visible in the source but lack formal MCP registration/definition. Descriptions exist but are generic and lack LLM-optimized guidance. Parameters have basic types but missing descriptions in several cases. No evidence of error handling patterns, permission gates, or security considerations. The codebase shows a LangGraph workflow implementation for spam detection, but tools are not properly registered as MCP tools with formal schemas. Most critically: no MCP server boilerplate detected, no tool registration mechanism visible, making this appear to be a local LangGraph application rather than an MCP server.
Second pass: Use OpenAI for uncertain cases
Get spam score using basic rules
Get or create Gmail API service
Fetch today's emails from Gmail
Label an email as spam
First pass: Filter using basic rules and spam checker
No MCP server implementation detected. Code shows a local LangGraph workflow, not an MCP server with tool registration. No evidence of ServerCapabilities, Tool definitions, or MCP protocol handling.
Tool descriptions lack LLM-optimized guidance. Generic descriptions like 'Label an email as spam' (16 chars) and 'Fetch today's emails from Gmail' (31 chars) do not explain WHEN to use the tool, WHAT it returns, or any prerequisites. Below 50-char minimum for LLM context.
Parameter descriptions missing or incomplete. 'pre_filter' and 'analyze_email' accept a 'state' parameter of type object but provide only high-level descriptions ('EmailState object containing...') without field-level documentation of what properties the LLM must supply.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 37 | <=2025-11-25 | v2 |
No output schemas documented. Tools return complex objects (EmailState, email lists, boolean flags, float scores) but no schema specifies the structure, field types, or required return fields. LLMs cannot plan downstream calls or extract data reliably.
No error handling or recovery guidance. Functions return boolean success flags or raise generic exceptions without telling the LLM what to do next. E.g., 'label_spam' returns True/False but provides no guidance on retry, fallback, or user notification.
API credentials exposed in code. OpenAI API key loaded via os.getenv('OPENAI_API_KEY'), while not a parameter, the key is embedded in the GmailLabeler initialization context. No evidence of secret injection patterns or credential isolation.
'get_todays_emails' and 'get_gmail_service' have empty input schemas ({}).
Tool naming lacks clarity in composition. 'pre_filter', 'analyze_email', and 'get_todays_emails' do not clearly indicate they are part of a stateful workflow. 'pre_filter' and 'analyze_email' suggest generic operations but actually mutate complex state objects. Names should reflect side effects (e.g., 'filter_emails_with_rules', 'classify_uncertain_emails').
No permission gates or scope declarations. Tools like 'label_spam' and 'analyze_email' modify Gmail state and call third-party APIs without any permission checks, audit logging, or scope limiting. An untrusted agent could label arbitrary emails or trigger expensive API calls.
No rate limiting or runaway protection. The workflow hardcodes MAX_API_CALLS=50 but this is enforced only within the LangGraph loop, not as a tool-level guard. A compromised agent or LLM loop could exhaust API quotas or spam Gmail operations.