Fast, containerized screenshot service using Chromium with HTTP API and MCP server support
This MCP server provides two screenshot tools with comprehensive parameter schemas and reasonable descriptions. Both tools have well-structured input schemas with type definitions, enums, and constraints. Descriptions are present and adequate (90-150 chars), though could be more directive about when/why to use them. Parameters are typed and mostly described. The main weaknesses are: (1) no output schema documentation, (2) no error handling guidance in descriptions, (3) no idempotency or retry guidance, (4) tool compositions could be more clearly distinguished. The server shows good baseline quality but lacks production-grade polish around error recovery and result documentation.
Capture a screenshot of a webpage. Returns base64-encoded image data. Supports cookie injection for authenticated pages. Use this for capturing web pages, especially long/full-page screenshots.
Capture a screenshot and save it to a file. Returns the file path. Supports cookie injection for authenticated pages.
Output schemas not documented. Tools return base64-encoded image data and file paths, but LLMs cannot plan downstream operations without knowing the exact response structure (field names, types, or pagination). Required for tool chaining.
No error handling or recovery guidance in tool descriptions. Agents don't know what to do if a page fails to load, a selector times out, or cookies are invalid. Descriptions should include: 'If the page fails to load, the tool will return an error with the failure reason. Check the URL format and ensure the domain is accessible.' Also missing timeout guidance and retry strategy.
Tool names lack clarity about execution scope. 'screenshot' is generic; 'screenshot_to_file' requires agents to decide: return base64 or save to disk? No guidance on when to use each. Better names: 'capture_screenshot_as_base64' and 'capture_screenshot_to_file' or a single 'screenshot' tool with an optional 'output_path' parameter (only write if provided).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 58 | - | v1 |
Parameter 'extract_dom' is complex (nested object with 5 sub-properties) and under-explained in the description. The description says 'Extract DOM element positions and text alongside screenshot. Enables hybrid text identification with Vision AI.' This is vague about WHEN to use it or WHAT the output structure is. Should clarify: 'Enable to extract bounding boxes and text content for each DOM element, useful when Vision AI recognition needs fallback text. Returns array of {selector, text, x, y, width, height}.'
Overlapping responsibility between 'screenshot' and 'screenshot_to_file' violates single-responsibility principle. Both capture pages; only the output method differs. This creates ambiguity for LLM tool selection. Recommend: merge into a single 'screenshot' tool where 'output_path' is optional, if omitted, return base64; if provided, save to file and return the path. Reduces cognitive load and eliminates false choice.
Cookie injection security: descriptions don't mention that sensitive cookies (auth tokens, session IDs) should not be passed via parameters in production, as they may be logged or exposed in traces. Should note: 'For production use with sensitive cookies, inject via server-side environment variables rather than parameters.'
No idempotency guarantees. Taking a screenshot at the same URL twice may produce different results if page content is dynamic or time-based. Tool descriptions should clarify: 'Screenshots are point-in-time captures; repeated calls may return different images if the page content changes.'