Web scraping and crawling tools for AI agents using Crawl4AI with Model Context Protocol integration
This server has 4 tools with complete input schemas and mostly good descriptions. Naming follows verb-noun patterns (scrape, crawl, crawl_site, crawl_sitemap). Schemas use Pydantic with HttpUrl validation and proper type constraints. However, several parameters lack descriptions, output schemas are documented via Pydantic models but not explicitly registered in tool definitions, and error handling guidance is minimal. The tool descriptions are clear about what they do and when to use them (10-15 chars each is too short per the rubric baseline of 194 chars), but stay focused on core functionality. Security is reasonably sound (public URLs validated, no secrets in params), but the server lacks tool annotations (readOnlyHint/destructiveHint) which are current MCP spec features.
Breadth-first crawl up to max_depth starting from seed_url. Returns markdown per page by default. If output_dir provided, persists to disk and returns metadata only (avoids context bloat). Respects same_domain_only and allows include/exclude regex patterns.
Site crawler that persists results to disk. Returns manifest path and output directory. Does not return page content to avoid context bloat.
Crawl URLs discovered from sitemap.xml/robots.txt. Persists results to disk and returns manifest path.
Fetch a single URL with Crawl4AI. Returns markdown + links by default. If output_dir provided, persists to disk and returns metadata only (avoids context bloat).
Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint) in MCP tool registration. crawl_site and crawl_sitemap have WRITE side effects but are not annotated. Current MCP spec (2026-07-28) expects these hints on the Tool object.
Pydantic model descriptions exist internally but are not visible in the MCP tool registration code excerpt. The source shows model definitions (ScrapeArgs, CrawlArgs, etc.) with Field descriptions, but the actual types.Tool registration for each parameter is truncated. Cannot verify that all parameter descriptions are being passed to the MCP server registration.
Error handling provides no recovery guidance. Tools return generic MCP errors without actionable next steps. Per the rubric (pattern:recovery-guide), error responses should tell the LLM what to do next (e.g., 'Timeout: retry with longer timeout_sec' or 'URL rejected: ensure it is a public HTTP/HTTPS URL').
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Tool descriptions are concise but lack WHEN-to-use and prerequisite guidance. 'Fetch a single URL with Crawl4AI' is clear but does not explain when to choose scrape vs crawl, or what preconditions must hold (e.g., 'Crawl only for breadth-first discovery; use scrape for single-page extraction'). Per the rubric (pattern:tool-description), descriptions should state WHAT, WHEN, and prerequisites.
Some parameters lack clear constraints in descriptions. E.g., crawler and browser params accept Dict[str, Any] with no guidance on valid keys or structure. An LLM has no way to know what overrides are valid. Per the rubric, parameter descriptions must include expected format, range, and allowed values.
Output schemas (ScrapeResult, CrawlResult, ScrapePersistResult, CrawlPersistResult) are defined as Pydantic models but are not explicitly registered in the MCP Tool definition visible in the code excerpt. The source code is truncated at the list_tools() function. Cannot confirm that output schema documentation is being communicated to the MCP client.
No pagination or result limits specified for crawl and crawl_site tools when max_depth > 1. crawl returns a flat list of CrawlPage objects; no mention of pagination support. Per the rubric (pattern:paginated-result), tools returning lists should offer page/offset + limit and return a total count or next_cursor.