Your Gemini web app as a service — turn the free Gemini web interface into a programmable HTTP API for image and video generation via Chrome CDP.
Phantom Canvas exposes 8 tools via HTTP API with varying definition quality. While input schemas are present for most tools, descriptions are inconsistent (ranging from 40-160 chars), and critical security/validation concerns exist. Tool naming follows verb patterns but lacks clarity around stateful operations (session/refresh vs session/save). No output schemas are documented in the source code. The server implements an HTTP API wrapper around a browser automation engine (Playwright + Chrome CDP for Gemini web interface), but tooling is not optimized for LLM consumption, parameters like 'id' are opaque, endpoint structures leak HTTP semantics, and error handling relies on HTTP status codes rather than actionable recovery guidance. The /debug/eval tool is a critical security risk (arbitrary JavaScript execution). Average tool score: 38/100.
Health check endpoint returning server readiness status, busy flag, and task queue state.
Poll the status and results of a generation task. Returns task status, generated images/videos, conversation ID, and elapsed time.
Download a generated image or video file from a completed task. Returns binary content with appropriate MIME type.
Execute arbitrary JavaScript code in the browser context for debugging. Returns evaluation result.
HTTP API endpoint to queue an asynchronous image or video generation task. Returns a task ID for polling.
Re-navigate to Gemini to refresh the session and ensure the browser is logged in.
Save the current browser session. Chrome mode persists via user-data-dir automatically.
CRITICAL SECURITY: /debug/eval tool allows arbitrary JavaScript execution in browser context with no validation, rate limiting, or permission gates. This violates pattern:secret-injection and pattern:permission-gate. Any agent with access can exfiltrate cookies, session tokens, or compromise the Gemini account.
OUTPUT SCHEMAS NOT DOCUMENTED: No tool documents what fields are returned from task polling, health checks, or image downloads. Pattern:tool and pattern:response-shaper require structured output schemas so LLMs know what to expect. Example: GET /task/:id returns 'task status, generated images/videos, conversation ID, elapsed time' as prose, not a JSON schema.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 37 | 2026-07-28+ | v2 |
One-shot image or video generation from a text prompt. Supports reference image input (img2img), video generation, timeout control, and conversation continuation.
OPAQUE IDENTIFIERS: Task IDs and image indices are passed as path parameters with no explanation of their format or how to obtain them. LLMs need to understand: is 'id' a UUID? How long? Where does it come from? Parameter naming violates mxe:natural-identifiers and lacks the constraint descriptions required by pattern:tool-description.
HTTP SEMANTICS LEAK INTO TOOL INTERFACE: Tool names are HTTP endpoints (GET /task/:id, POST /generate) instead of action verbs. This violates pattern:tool naming (should be verb_noun like 'get_task', 'poll_task', 'download_generated_image'). LLMs must infer intent from REST structure rather than reading action names.
DUPLICATE/CONFUSING TOOLS: 'generate' (CLI one-shot sync) and 'POST /generate' (HTTP async) do the same thing but with opposite semantics (blocking vs async polling). LLMs will struggle to pick between them. Per pattern:tool and pattern:composition, each tool should have one clear purpose, either consolidate or name them clearly to distinguish sync vs async (e.g. 'generate_sync' vs 'generate_async_queue').
NO ERROR HANDLING GUIDANCE: Tool descriptions do not explain what errors can occur or how LLMs should recover. Example: 'timeout_secs' can fail with a timeout, but there's no actionable message (pattern:recovery-guide). Similarly, 'conversation_id' references are opaque, if a conversation ID is invalid, what should the LLM try next?
SESSION MANAGEMENT IS OPAQUE: /session/refresh and /session/save describe browser state transitions but do not explain WHEN to call them or WHAT happens on error. Are sessions persisted? When does a session expire? What does 'Chrome mode persists via user-data-dir automatically' mean to an agent? LLMs cannot reason about session validity without clear state descriptions.
PARAMETER DESCRIPTIONS LACK FORMAT/CONSTRAINT DETAILS: 'prompt' is described as 'text prompt' with no length limit, format rule, or example structure. 'reference_images' array lacks description of supported formats (PNG? JPEG? Max size?). 'timeout_secs' defaults differ by type but aren't documented (180 vs 300). Per pattern:tool-description and review:param-validation-rules, each param must state format, range, and examples.
MISSING PAGINATION AND RESULT LIMITS: /health returns 'busy flag, task queue state' without defining structure or how many pending tasks are shown. If queue grows to thousands, returning all pending tasks would blow context windows. Pattern:paginated-result and mxe:enforce-result-limits require pagination or a cap with clear documentation.
CALLBACK_URL PARAMETER ENCOURAGES INSECURE PATTERNS: 'callback_url' in POST /generate is a free-text URL with no validation. Agents could be tricked into registering malicious callbacks. Per pattern:secret-injection and pattern:tool-gateway, URLs should not be user-provided parameters, use server-side webhooks config instead.