Web browsing superpowers for AI agents: one MCP tool to scrape, extract, crawl, map, and search any site — self-hosted, no API keys, with a smart auto-fallback ladder.
PyreCrawl demonstrates solid tool naming (verb-first: scrape, extract, map_site, crawl, search, batch_scrape, deep_research, document, monitor, search_papers, session, cache, health) and comprehensive descriptions covering WHAT and WHEN for most tools. Strengths: clear use-case guidance in tool descriptions (e.g., scrape explicitly states 'Use this when the user shares a URL and wants its content'), escalation strategy well-documented (fast→stealth→llm ladder), and batch operations exposed (batch_scrape, deep_research). Gaps: (1) schemas are present but parameter descriptions lack specificity, e.g., 'prefer' parameter describes modes but omits what happens on failures or cost tradeoffs; (2) output schema documentation is absent, markdown field is mentioned but shape, max length, and field presence guarantees are not documented; (3) error responses not specified, no guidance on what error structure LLMs should expect, recovery hints, or retryability classification; (4) session/cache tools are write-capable but lack confirmation or dry-run patterns; (5) several parameters lack constraints, timeout as integer has no min/max, limit has no stated cap despite being mentioned as defaulting to 200. The tool descriptions themselves (50-150 chars, averaging ~120) fall within production baseline but could be tightened. Parameter annotations average 40-80 chars, acceptable but some (js, wait_for) are verbose. Schemas visible in source are properly typed (string, integer, boolean, object, array) with Pydantic-backed validation, but output structures are not formally declared in the tool registration.
Scrape many URLs in parallel through the smart ladder. Deduplicates input URLs, serves cache hits instantly, scrapes the rest with a bounded thread pool, and returns per-URL results (never raises).
Inspect and manage the response cache. Use to view cache statistics, clear entries, or adjust cache settings.
Crawl a site breadth-first, scraping multiple pages. Use when the user wants to understand a site's content (research, comparison, catalog scraping). Crawls from root following internal links, respecting robots.txt.
Multi-pass research loop: search → scrape → refine → repeat. Use for complex research questions. Runs multiple search + scrape cycles, refining the query each time based on what was found.
Extract structured text from a document (PDF, DOCX, PPTX, images). Use when the user provides a file URL and wants text extraction. Supports Markdown, HTML, and JSON outputs.
Scrape + structured extraction using a CSS-based JSON schema. Use when the user wants structured data (tables, lists, product info) extracted from a page. Define a CSS schema to target specific elements. The schema is a JsonCssExtractionStrategy schema: { "name": "PageItems", "baseSelector": "div.item", "fields": [{"name": "title", "selector": "h2", "type": "text"}, ...] } Returns parsed JSON in `data`.
Output schemas not documented in tool definitions. Tools return 'markdown', 'html', 'title', 'status', 'meta' fields but no formal schema specification exists in the registration. LLMs cannot plan downstream tool chains without knowing return field names and types.
Parameter constraints missing: 'timeout' (integer) lacks min/max bounds; 'limit' (integer) in map_site, search, search_papers has no stated maximum despite description mentioning defaults like 200 or 10; 'iterations' in deep_research unbounded; 'max_concurrency' in batch_scrape has no limit. Unbounded numbers invite LLM hallucination of unreasonable values (timeout=999999, max_concurrency=1000).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
Health check and dependency versions. Use to verify the server is running and check component versions.
Enumerate all internal URLs reachable from `root`. Use when the user wants to map a site's structure or find all pages before crawling. Often paired with `crawl` or `batch_scrape`.
Monitor a URL for changes and notify on significant updates. Use when the user wants to track a page's content over time. Returns diffs between snapshots and a change summary.
Scrape a single URL → LLM-ready markdown. Use this when the user shares a URL and wants its content (read, analyze, summarize, extract). Auto-escalates through fast→stealth→llm when blocked.
Search the web, scrape results → LLM-ready markdown. Use for user queries like "find X" or "latest news on Y". Returns a list of URLs with extracted title + summary.
Search academic papers and extract summaries. Use for research queries targeting peer-reviewed literature. Integrates with arXiv, Google Scholar, and other academic sources.
Manage persistent browser sessions with cookies and state. Use for multi-step workflows: login → interact → scrape. Sessions are keyed by name and persist cookies across calls.
Error handling guidance absent. Tools mention they handle Cloudflare blocks (fast→stealth→llm escalation) but do not specify what error structures LLMs will receive, whether errors are retryable, or what hints should guide recovery. Session and cache tools (write-capable) offer no dry-run, confirmation, or destructive-action warnings.
Session tool (action parameter) is action-based (open, click, fill, type, eval, screenshot, close) but lacks documentation of preconditions (does 'click' require a prior 'open'?), state machine (can you click before opening?), and error cases (what if selector not found?). Multi-step workflows are fragile without explicit ordering and dependency documentation.
Parameter relationship undocumented: in scrape, 'js' and 'wait_for' are labeled '(stealth only)' but the tool does not formally constrain or validate that 'prefer' must be 'stealth' or 'auto' for these to work. If LLM passes js with prefer='fast', tool will silently ignore it or error unclearly.
Monitor tool lacks duration and alarm semantics: 'interval_seconds' and 'max_checks' are numeric but no guidance on total monitoring window, whether the tool blocks/polls or returns immediately, or what 'change summary' format is returned. An agent cannot plan retry or timeout logic without this.
Composition risk: multiple search tools exist (search, search_papers, deep_research) with overlapping intent. An LLM must reason about when to use search vs deep_research vs search_papers. No explicit disambiguation provided (e.g., 'search = general web; search_papers = academic only; deep_research = iterative multi-pass research').