Unofficial MCP server wrapper for crawl4ai - Advanced web crawling through Model Context Protocol
The server exposes 20 tools with STDIO transport, wrapping the crawl4ai library. Tool definitions are present in crawl4ai_mcp/core/tool_guides.py with names, descriptions, and input schemas. However, quality varies significantly. Strong points: most tools have descriptive names (verb-prefixed), detailed parameter descriptions for complex tools like crawl_url and batch_crawl, and documented input schemas with type constraints. Weak points: 9 tools (deep_crawl_site, crawl_url_with_fallback, extract_structured_data, extract_youtube_transcript, get_youtube_video_info, batch_extract_youtube_transcripts, search_google, search_and_crawl, batch_search_google) have only placeholder descriptions ('Crawl multiple related pages from a website', 'Crawl URL with fallback strategies...') and NO visible input schemas in the provided source. Output schemas are documented for a few tools (e.g., crawl_url mentions metadata-only response, file path return), but most lack explicit return type documentation. Error handling exists (timeout, overwrite rejection, max URL limits) but lacks actionable recovery guidance. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite many tools being read-only or idempotent. Security: no credentials in parameters (good), but no permission gates or scope declarations.
Crawl multiple URLs with fallback. Max 3 URLs per call. Use output_path (directory) to persist full per-URL markdown + index.json; the return shape stays a list, each success item gets an output_file key.
Extract transcripts from multiple YouTube videos
Batch search Google with multiple queries
Extract web page content with JavaScript support. Use wait_for_js=true for SPAs. Use content_offset/content_limit to paginate the response. Use output_path to persist the full unsliced content to disk as markdown and receive a slim metadata-only response.
Crawl URL with fallback strategies for difficult sites
Crawl multiple related pages from a website
9 of 20 tools (45%) have only placeholder descriptions under 20 characters and zero visible input schemas. Tools like 'extract_youtube_transcript', 'search_google', 'extract_structured_data' are completely underdefined. This violates pattern:tool-description and pattern:tool.
Output schemas are documented in tool descriptions (e.g., 'response is slimmed to metadata+file path') but no explicit structured return type documentation exists. LLMs cannot plan downstream tool chains without knowing the exact fields returned (e.g., output_file field in batch_crawl). Violates pattern:response-shaper.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Process large content with chunking and BM25 filtering. Use output_path to persist chunks + summaries to disk as JSON and receive a slim response.
Extract entities (emails, phones, etc.) from web pages. Use output_path to persist the full entity extraction output to disk as JSON and receive a slim response.
Extract structured data from web pages using schema
Extract YouTube video transcript
Get LLM configuration information
Get available search genres for filtering
Get supported file formats (PDF, Office, ZIP) and their capabilities.
Get comprehensive tool selection guide for AI agents
Get YouTube API setup guide
Get YouTube video information
Extract specific data from web pages using LLM. Use output_path to persist the full extraction output to disk as JSON and receive a slim response.
Convert PDF, Word, Excel, PowerPoint, ZIP to markdown. Use output_path to persist the full unsliced converted markdown to disk and receive a slim response.
Search Google and crawl top results
Search Google with genre-specific filtering
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear semantics. crawl_url, batch_crawl, search_google, etc. are all read-only; get_supported_file_formats, get_youtube_api_setup_guide, get_search_genres are pure info tools. Agents cannot infer safety without these hints. Violates current MCP spec (2026-07-28) for tool annotations.
Error handling lacks actionable recovery guidance. Tools mention 'timeout', 'overwrite rejection', 'max 3 URLs' but do not guide the agent what to do next. Example: if batch_crawl fails on 1 of 3 URLs, the response doesn't say 'retry individually' or 'reduce concurrency'. Violates pattern:recovery-guide.
No pagination parameters for multi-result tools (search_google, batch_search_google, get_youtube_video_info). If results are unbounded, agents cannot retrieve partial sets without overloading the context window. Violates pattern:paginated-result.
Several complex tools (batch_crawl, intelligent_extract, enhanced_process_large_content) support output_path (file persistence) and include_content_in_response flags, but the interaction and response shape are described informally in descriptions rather than via formal output schema. LLMs must infer: when output_path is set, are results still in the response? Are they sliced by content_limit? This is underdocumented and leads to misuse. Violates pattern:response-shaper.