MCP server that gives AI assistants web crawling superpowers via crawl4ai
crawl4ai-mcp demonstrates solid foundational quality with verb-prefixed naming conventions and comprehensive parameter schemas. However, critical gaps exist in parameter descriptions, output schema documentation, and error handling guidance. The codebase shows strong engineering practices (proper logging, async handling, type hints) but falls short of production-grade agent tool standards. Most tools have well-structured input schemas but lack explicit output schema documentation and error recovery guidance that LLMs rely on for orchestration.
Check the status of the Chromium browser and Playwright installation. Returns current readiness state (ready, repairing, or failed) with details.
Close and clean up a persistent browser session.
Recursively crawl a website from a starting URL, discovering and crawling linked pages up to a maximum depth. Supports domain filtering, URL pattern matching, and keyword-based relevance scoring.
Crawl multiple URLs concurrently with configurable batch size and concurrency limits. Returns results for each URL with shared session state.
Fetch and extract content from a URL with configurable parsing strategies (LLM-based, CSS, XPath, regex, JSON). Handles JavaScript rendering, markdown conversion, and smart content filtering.
Create a persistent browser session that preserves cookies, localStorage, and DOM state across crawl_url calls. Returns a session_id to use in subsequent calls.
Output schemas are undocumented for all 9 tools. LLMs cannot predict response structure, plan downstream tool calls, or extract required fields for chaining. This violates the fundamental requirement that agents understand tool composition.
Parameter descriptions lack format constraints and interdependencies. For example, 'concurrent_requests' has bounds (1-10) only in description text, not schema; 'extraction_strategy' in crawl_url has options documented but no guidance on which strategy to choose based on data type.
Error handling is implicit rather than explicit. No tool description explains what to do if the crawl fails, session is invalid, browser repair times out, or extraction produces no matches. Agents are left to guess recovery strategies.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Extract all links from a URL, optionally filtering by pattern or domain. Returns href, text, and type (internal/external) for each link.
List all available crawl profiles with their effective settings and any ignored/invalid keys.
Detect and repair a missing or broken Chromium browser installation. Downloads the Playwright-expected build via crawl4ai-setup or playwright install.
Parameter interdependencies and mutual exclusivity are undocumented. In crawl_url, parameters like 'js_code', 'css_selector', 'xpath', 'regex', and 'extraction_strategy' likely have constraints (e.g., only one extraction method per call), but this is not stated in descriptions.
crawl_deep lacks guidance on result limits. The tool accepts max_pages but description does not state whether pagination is supported or how large result sets are handled. Without this, agents cannot predict context impact.
Sensitive parameters are accepted but not validated in descriptions. For example, 'llm_model' and 'llm_temperature' in crawl_url could require API credentials, the tool description should clarify whether user-supplied LLM models are supported or only curated options.