AI-native, agent-first web scraping library with deterministic replay, exposed as MCP tools
Scrapo demonstrates solid tool design with well-structured schemas, clear naming conventions, and comprehensive parameter documentation. All 7 tools follow verb_noun patterns (scrapo_scrape, scrapo_crawl, etc.). Input schemas are complete with type definitions, constraints (minItems, maxItems, min/max for integers), and defaults. Descriptions are substantive (100-250 chars), explaining WHAT each tool does and WHEN to use it. However, output schemas are not documented, LLMs cannot see what fields to expect from responses. Error handling guidance is absent; tools don't indicate which errors are retryable or how to recover. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite all tools being READ_ONLY. Parameter descriptions could be more prescriptive about format/constraints (e.g., 'max_tier: integer 0-4, where 0=HTTP, 1=HTTP+session, 2=browser, 3=stealth, 4=agent' is clear but not in the description text itself).
Scrape an explicit list of URLs concurrently (not a recursive crawl). Per-URL error isolation; returns one result per input URL in order.
Recursive crawl from one or more seed URLs. Returns crawl_id and stats.
Field-level diff between two recorded runs.
List recent recorded runs, optionally filtered by URL.
Discover the URLs on a site WITHOUT scraping their content. Merges each origin's sitemap.xml with a bounded same-host link crawl; returns a sorted, de-duplicated, SSRF-filtered URL list. Use this to see what's on a site before crawling.
Re-run extraction over a previously archived run's HTML without re-fetching the live page.
Output schemas not documented. LLMs cannot see what fields responses contain (e.g., does scrapo_scrape return 'markdown', 'chunks', 'via', 'not_modified'?). This forces agents to guess field names and risks failed downstream tool calls.
No error handling guidance. Tools don't indicate which errors are retryable (e.g., network timeout vs. invalid URL), what the LLM should do next, or how to recover. Agents will retry blindly or give up.
Tool annotations missing. All tools are READ_ONLY but lack readOnlyHint annotation. This prevents clients from optimizing caching, retry logic, or UI presentation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
Fetch a single URL through Scrapo's tier router and return clean markdown + provenance-tagged chunks. Optionally extracts typed JSON. Re-scraping a URL uses a conditional GET, so an unchanged page comes back fast (not_modified=true). Known sites with a clean public API (Wikipedia and its Wikimedia sister projects) are auto-routed through that API instead of the bot-walled page — so a Wikipedia URL just works, no CAPTCHA, no max_tier needed; the response reports via="api:<provider>" when this happens. Set api_first=false to force-scrape the real page.
Descriptions for scrapo_replay, scrapo_diff, scrapo_list_runs are minimal (under 80 chars). They lack context on WHEN to use these tools vs. alternatives or what they return. E.g., 'scrapo_replay: Re-run extraction over a previously archived run's HTML without re-fetching the live page.', does this return the same fields as scrapo_scrape? Is it for debugging or production use?
Parameter descriptions lack prescriptive format guidance. E.g., 'max_tier' description states the enum values but not in the description text itself, LLMs must infer from the schema. 'wait_for' says 'CSS selector' but doesn't explain what happens if the selector never matches (timeout? error?).