Record any browser page as GIF or video via MCP — with auto-zoom on interactions. Powered by Playwright + ffmpeg + Remotion.
pagecast has 16 tools with generally clear naming and good platform-specific descriptions, but significant gaps in schema completeness, parameter documentation, and error handling prevent a higher score. 14 of 16 tools have descriptions, but parameter schemas are partially visible, many parameters lack explicit type constraints or detailed range/format documentation. Output schemas are not documented. Error handling is minimal; no recovery guidance or actionable error messages visible in definitions. The 'interact_page' tool uses a complex nested action array structure that lacks clear documentation of what each action type returns or how failures propagate. Compositions are reasonable (record → interact → stop → convert), but tools like 'convert_with_zoom_gif' and 'convert_with_magnify_gif' expose numerous tuning parameters (zoomLevel, transitionDuration, holdPerTarget, etc.) with defaults but minimal guidance on selection. STDIO-only transport prevents remote accessibility.
Stop all active recording sessions and clean up resources.
Convert a .webm video to an optimized GIF using ffmpeg two-pass palette method.
Convert a .webm video to MP4 (H.264). Widely compatible for social media, sharing, and embedding.
Convert WebM to GIF with magnifying glass overlays on interactions. The full viewport stays visible. A zoomed-in lens appears near each interaction.
Convert WebM to MP4 with magnifying glass overlays on interactions.
Convert WebM to GIF with modern tooltip overlays on interactions. Each interaction gets a smooth tooltip annotation.
Convert WebM to MP4 with modern tooltip overlays on interactions. Each interaction gets a smooth tooltip annotation.
interact_page action array structure lacks clear schema documentation. The nested 'actions' parameter is an array of objects with conditional properties (e.g., 'ms' for wait, 'selector' for click, 'text' for type) but no guidance on which properties are required for each action type, what happens on failure, or what the action returns. This forces the LLM to infer or guess behavior.
Conversion tools (convert_with_zoom_gif, convert_with_magnify_gif, etc.) expose many tuning parameters (zoomLevel default 2.5, transitionDuration default 0.35, holdPerTarget default 0.8, panDuration, chainGap, etc.) with only default values shown. No guidance on valid ranges, when to adjust, or what visual results each setting produces. This leaves LLMs unable to reason about parameter selection.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 53 | 2026-07-28+ | v2 |
Convert WebM to zoom-enhanced GIF using FFmpeg crop expressions. Reads the event timeline to know where and when to zoom. Pipeline: input → [zoom crop+scale] → fps+resize → two-pass palette GIF
Convert WebM to zoom-enhanced MP4 using FFmpeg crop expressions.
Perform actions on a recording page (scroll, click, hover, type, press, select, wait, navigate). Actions are performed sequentially and recorded in the video.
List all .webm recordings in the output directory with file sizes and timestamps.
List all active recording sessions with their current status, started time, and viewport dimensions.
All-in-one: open URL, wait for specified duration, stop recording, auto-export to the right format. Use the "platform" parameter and we handle everything: - "Record a demo for my GitHub README" → platform: "github" → 1280×720 GIF - "Record my app for Instagram Reels" → platform: "reels" → 1080×1920 MP4 - "Make a TikTok demo" → platform: "tiktok" → 1080×1920 MP4 - "Record for YouTube" → platform: "youtube" → 1280×720 MP4 Or pass custom width/height/outputFormat for full control.
Open a URL in a browser and start recording video. Returns a session ID. Call stop_recording when done. Instead of specifying width/height, you can use the "platform" parameter: - "Record my app for GitHub README" → platform: "github" (1280×720, GIF) - "Record my app for Instagram Reels" → platform: "reels" (1080×1920, MP4) - "Record my app for TikTok" → platform: "tiktok" (1080×1920, MP4) - "Record my app for YouTube" → platform: "youtube" (1280×720, MP4) - "Record my app for YouTube Shorts" → platform: "shorts" (1080×1920, MP4) - "Record my app for Instagram post" → platform: "instagram" (1080×1080, MP4) - "Record my app for LinkedIn" → platform: "linkedin" (1080×1080, MP4) - "Record my app for Twitter" → platform: "twitter" (1280×720, MP4) Or pass custom width/height for any other size.
Render a cinematic demo video using Remotion (if available). Creates high-quality, animated presentation videos.
Stop a recording session and save the video as .webm file.
Output schemas not documented for any tool. What does record_page return exactly? A sessionId string? An object with sessionId + viewport dimensions? What do convert_with_zoom_gif and similar tools return, file path, success boolean, metadata? LLMs cannot chain calls without knowing the response structure.
No error handling documentation. What happens if a URL is unreachable? If a session ID is invalid? If ffmpeg is not installed? If a selector in interact_page does not match? Tools must guide LLMs on recovery: 'If selector not found, try waiting first with waitForSelector action' or 'ffmpeg not installed, ensure package.json postinstall step ran.'
platform parameter in record_page and record_and_export uses a string enum but does not explicitly list all valid values in the schema. Description lists examples ('github', 'reels', 'youtube', etc.) but schema's enum field shows only placeholder values or is missing. LLMs may hallucinate unsupported platforms.
cleanup tool has no input parameters but is marked DESTRUCTIVE risk. Description says 'Stop all active recording sessions and clean up resources' but does not warn of irreversibility or ask for confirmation. No guard or dry-run option visible.
interact_page action types (waitForSelector, etc.) are documented only as enum values. No description of what each action does, how long it takes, what it returns, or how it affects the video timeline. An LLM must guess whether 'waitForSelector' blocks video or pauses recording.
sessionId is used across multiple tools (interact_page, stop_recording) but is never defined as a reference type or documented as the return value of record_page. An LLM has to infer that record_page returns a sessionId and that this ID is the right value to pass to interact_page.