MCP Screenshot tool for AI-powered image analysis and verification
The MCP Screenshot server registers 10 tools with visible schemas and descriptions, but exhibits significant quality gaps compared to production baselines. Most tools have verb_noun naming (capture_screen, describe_image) which is good, but descriptions average ~80-120 characters, below the 194-char production baseline. Critically, parameter descriptions are sparse or missing entirely for many inputs: 'region', 'zoom_center', 'zoom_factor', 'model', 'threshold', and batch configuration arrays lack meaningful guidance on format, constraints, or expected values. Schema completeness varies: capture_screen and capture_webpage define integer/string/array types with defaults, but describe_image exposes a bare 'model' parameter with no enum or constraint guidance, inviting hallucinated model names. The batch_capture and batch_describe tools accept generic 'screenshots' and 'image_paths' arrays with no nested schema documentation, making it impossible for an LLM to know what structure to pass. Error handling is minimal, the code snippet shows a try/catch that logs errors and returns {'success': False, 'error': str(e)}, which is generic and non-actionable. No recovery guidance, no categorization (retryable vs fatal), no constraint feedback. Tool composition is reasonable (each tool does one thing), but several tools combine concerns in ways that invite misuse: verify_d3 accepts both 'url' and implicit chart detection, requiring the LLM to reason about when to provide vs omit chart_type. Output schemas are undocumented, code shows result dicts with 'success', 'error', and file paths, but LLMs cannot infer what fields describe_image_content or batch_describe will return. No pagination or result limits documented for batch operations. Overall, the server reads as a competent engineering effort that exposes real functionality, but falls short of production-grade tool design patterns.
Add annotations to a screenshot image.
Capture multiple screenshots in batch.
Describe multiple images in batch.
Capture a screenshot and describe it in one operation.
Capture a screenshot of the screen or a specific region.
Capture a screenshot of a webpage using headless browser.
Compare two screenshots and highlight differences.
Parameter descriptions missing or vague for 'model', 'region', 'zoom_center', 'zoom_factor', 'threshold', 'annotations', and 'screenshots' array schema. LLMs cannot infer valid values or format without explicit guidance.
'model' parameter in describe_image, capture_and_describe, verify_d3, and batch_describe is a free-form string with no enum or constraint. Should specify available models (claude-3-5-sonnet, gpt-4-vision, etc.) as an enum or document the default model explicitly.
Batch tools (batch_capture, batch_describe) accept 'screenshots' and 'image_paths' arrays but provide no nested schema or documentation of expected array structure. How should batch_capture array elements be formatted? What fields does each element require?
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Describe an image using AI vision model.
Get available screen regions for capturing.
Verify a D3.js visualization meets expectations.
capture_and_describe requires either 'url' OR 'file_path' but does not document this mutual exclusivity. LLMs may pass both, causing ambiguity or failure.
Error handling returns generic {'success': False, 'error': str(e)} with no recovery guidance, error categorization, or constraint feedback. LLMs cannot self-correct from invalid model names, missing files, or network failures.
Output schemas are not documented in tool descriptions. LLMs cannot predict what fields (e.g., describe_image_content returns) they will receive, forcing them to reason about response structure.
Tool descriptions are 50-120 characters on average, below the production baseline of 194 chars. Most descriptions lack context on WHEN to call the tool vs similar tools and any prerequisites.
capture_and_describe combines two concerns (capture + describe). While composition is sometimes useful, the tool should clarify that it can EITHER capture from a URL OR describe an existing file, not both in one call.