MCP server for PageBolt — take screenshots, generate PDFs, create OG images, inspect pages, record demo videos with Audio Guide narration, from AI coding assistants like Claude, Cursor, and Windsurf.
PageBolt MCP exposes 10 well-defined tools with strong schema completeness and clear descriptions. All tools have explicit parameter schemas with type definitions and descriptions. However, output schemas are undocumented (responses inferred from API behavior rather than formally specified in the tool definition), and error handling guidance is minimal. Tool naming is consistent and action-oriented (take_, generate_, create_, observe_, act_, etc.), following verb-noun conventions well. Descriptions average ~150-200 chars and are generally LLM-optimized. Security is solid (API key injected server-side, path traversal protection in safePath()). Composition is good, each tool has a single responsibility. The main limitation is lack of documented output schemas and error recovery guides, which prevents agents from planning multi-tool chains with confidence.
Goal-driven automation: give a URL + plain-English goal, runs an observe→plan→act→verify loop and returns a trace. Requires Starter+ plan.
Generate social card images (Open Graph) from templates or custom HTML
Convert a URL or HTML to PDF and save to disk
Poll the status of an async job (e.g., video encoding or page action). Returns status, queue position, output URL, or error.
Convert a page-agent/browser-use action trace into a re-runnable PageBolt sequence (pairs with observe_page format:'flatdomtree'). Returns the optimized step sequence and optional screenshot. Free (no quota cost).
Agent-optimized page observation: id-indexed elements, page-type classification, suggested actions, optional content/ARIA/screenshot/console. Set format:'flatdomtree' for browser-use / page-agent dom_text + selectors map
Output schemas are not documented. Tools define input parameters comprehensively but do not specify what fields/types they return. This forces LLMs to guess response structure, breaking downstream tool chaining and multi-step planning. Example: take_screenshot returns base64 or file path, but the schema does not document this choice or output type.
No error recovery guidance in tool descriptions. Errors like 'Page load timeout' or 'API rate limit' occur but tools do not explain what the LLM should do next (retry, adjust timeout, try different URL, etc.). This leaves agents stranded on failures.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Record a browser automation sequence as a video (MP4, WebM, GIF). Can optionally add Audio Guide narration. Supports async mode for long renders (jobs polling).
Execute a previously-saved PageBolt sequence by name or ID
Capture a URL, HTML, or Markdown as PNG/JPEG/WebP with optional styling (frame, background, shadow, padding)
Pixel-level visual comparison between two URLs or screenshots. Highlights the differences.
Tool descriptions lack explicit state-modification warnings. Tools like generate_pdf, record_video, and act_on_page make external API calls with side effects (file writes, job creation, automation execution), but descriptions do not clearly state 'This tool modifies state' or 'This may incur API quota costs.' Agents cannot distinguish safe from dangerous operations.
Some parameter descriptions lack actionable constraints. Example: act_on_page's 'goal' parameter is described as 'Plain-English description' but does not explain what length, complexity, or format optimizes success. observe_page's 'include' parameter is documented as 'object' but does not specify valid keys or values.
Async job handling (record_video + get_job_status) lacks pagination and result limits. get_job_status returns single job status, but if an agent enqueues many videos without tracking IDs, there is no way to list or search jobs. No guidance on job ID retention or cleanup.