MCP server for web fetching and automation using Chromium headless browser
This server implements 5 web automation tools with complete input schemas and reasonable descriptions. All tools have clear verb-noun naming (fetch_page, screenshot, interact, extract_data, get_link) and documented parameters with type constraints. However, descriptions lack the depth needed for optimal LLM selection, they read more like technical specs than agent-optimized guidance. No output schemas are documented, which forces LLMs to guess at response structure. Error handling is not visible in the provided code, and there are no recovery guides or actionable error messages shown. The interact tool is the only one with WRITE semantics clearly identified, but no confirmation or dry-run pattern is visible.
Extract structured data from a web page using CSS selectors. Returns JSON with extracted text content or attribute values. Supports single element, multiple elements, and attribute extraction.
Fetch a web page using Chromium headless browser and return content as markdown or HTML. Useful for reading web pages, extracting content from dynamic sites that require JavaScript rendering.
Get link href and text from a web page. By default extracts href without clicking. When click=true, follows the link and returns the final URL. Useful for extracting download links, checking redirect destinations, or getting link text.
Interact with web page elements using Chromium headless browser. Execute actions like click, fill, select, scroll, and wait in sequence. Returns the final page content as markdown after all actions are executed.
Take a screenshot of a web page using Chromium headless browser. Returns base64-encoded image. Supports PNG and JPEG formats, full page screenshots, and element screenshots via CSS selector.
No output schemas documented. LLMs cannot infer what fields to expect in responses, forcing them to guess at response structure and breaking downstream tool chaining.
Descriptions are too technical and lack LLM-optimized guidance. They describe WHAT the tool does but omit WHEN to use it and what the expected behavior is. Average description length ~150 chars, acceptable by baseline, but quality is surface-level.
No visible error handling or recovery guidance. Code does not show how the server handles network timeouts, invalid selectors, missing pages, or JavaScript execution failures. Without recovery hints, LLMs cannot self-correct.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
The interact tool modifies state (WRITE risk) but lacks confirmation or dry-run support. No confirmation_request pattern visible to prevent accidental destructive actions.
Parameters could accept natural-language selectors or aliases but force rigid CSS selectors. For example, interact requires exact CSS selectors rather than accepting 'click the submit button' and resolving internally. This breaks the chat-data-model pattern.
No rate limiting or abuse prevention visible. Agents could trigger unlimited browser instances, causing resource exhaustion. No backpressure or quota mechanisms shown.