A RAG (Retrieval-Augmented Generation) system that ingests emails from Gmail, crawls links found in emails, generates vector embeddings, and stores them in Supabase for semantic search and knowledge base queries.
This MCP server has critical structural and definition issues that prevent safe production use. The two tools (ingest_gmail, crawl_links) have non-empty descriptions and input schemas, but lack essential design patterns for security, error handling, and composition. The codebase is a Supabase Edge Function implementation, not an MCP server. No MCP server scaffold, protocol handler, or tool registration mechanism is visible in the provided code. Tools appear to be raw HTTP endpoints that would need wrapper logic to comply with MCP. Per-tool analysis follows.
Retrieves unprocessed links from stored emails, crawls them using Crawl4AI API, generates embeddings for the content, stores link documents in Supabase, and processes child links hierarchically
Fetches emails from Gmail API, extracts text and links, generates embeddings using OpenAI, and stores email data in Supabase database
No MCP server scaffold or protocol implementation detected. Code shows raw Deno HTTP edge functions, not an MCP server. No tools are registered via MCP::Tool or equivalent mechanism. The repository appears to be Supabase edge functions, not an MCP server.
Credentials exposed in environment as plaintext: OPENAI_API_KEY, SUPABASE_URL, SUPABASE_SERVICE_ROLE_KEY, CRAWL4AI_API_KEY. No secret injection abstraction or credential rotation mechanism. Keys are read via Deno.env.get() with empty-string fallback, risking silent failures. Agent traces will log parameter values; sensitive keys could leak.
No input validation or constraint enforcement. crawl_links accepts max_links and older_than_minutes as arbitrary integers with no range checks. An agent could pass max_links=999999 or older_than_minutes=-1. No enum types, pattern constraints, or min/max bounds documented or enforced.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 37 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Error handling is minimal and non-actionable. Both tools return generic 500 errors with only error.message. No recovery guidance, no categorization of errors as retryable vs fatal, no indication of what the agent should try next. Crawler failures for individual links are silently logged but succeed overall; agent cannot distinguish partial failures.
No permission checks, audit logging, or rate limiting. Tools modify Supabase database (insert emails, links, tasks) without confirming caller identity or permission. No audit trail of who invoked which tool when. No rate limits prevent runaway agents from flooding Gmail API or Crawl4AI from thousands of concurrent requests.
Unsafe nested loops and uncontrolled recursion in crawl_links. The processChildLinks function crawls up to 5 child links per parent with no depth limit or cycle detection. If a website has circular links or deep nesting, this could spiral into unbounded crawl operations, exhausting Crawl4AI API quota and Supabase storage.
Output schemas not documented. crawl_links returns {success, processed_count, message}. ingest_gmail implementation is truncated, but no output schema is specified for either tool. Agents cannot plan downstream tool chains when return types are undocumented.
Parameter descriptions are generic and lack constraints. crawl_links max_links description says 'Maximum number of links to process in one execution (default: 50)' but does not specify range, units, or side effects of large values. older_than_minutes lacks guidance on what happens if set to 0 or negative values.
No idempotency guarantees. crawl_links checks if a URL exists before inserting, but a race condition between check and insert could create duplicates if two agents call simultaneously. ingest_gmail lacks deduplication logic entirely. Non-idempotent operations risk duplicate emails/links on retry.
ingest_gmail description is incomplete in provided source; implementation truncated at import statements. Cannot fully assess naming, schema, or parameter quality for this tool.