Streamable HTTP transport for claude.ai. Exposes Scrapling's fetchers as MCP tools so Claude can scrape pages that web_fetch fails on (GitHub, Cloudflare-protected, etc.)
Shark-no-Kari demonstrates solid tool design with clear, action-verb naming (fetch_*, extract_*, get_*) and well-structured schemas. All 5 tools have descriptions (194 - 280 chars, within baseline 34 - 392 range) and complete input schemas with typed parameters. However, output schemas are undocumented, LLMs cannot predict response structure. Error handling is minimal; tools lack recovery guidance. Parameter descriptions are good but could specify constraints (e.g., max URL length, timeout behavior). No tool annotations (readOnlyHint, idempotentHint) despite all being read-only operations.
Fetch a page and extract multiple elements using CSS selectors. Returns structured JSON with the results.
Fetch and parse an RSS/Atom feed. Filters entries by cutoff_datetime and optional skip_terms, strips HTML from summaries, returns compact JSON.
Fetch a web page using a fast HTTP request with stealth headers. Good for static pages, GitHub raw content, docs sites, etc.
Fetch YouTube video transcripts/captions by URL.
Fetch a page using a real headless browser with anti-bot evasion. Bypasses Cloudflare Turnstile, bot detection, JS-rendered pages. Slower than fetch_page (~5-15s) but much more capable.
Output schemas undocumented. LLMs cannot predict response structure (JSON fields, types, pagination). fetch_page returns markdown or HTML; extract_elements returns JSON; fetch_feed returns parsed entries, but none are formally documented.
No tool annotations. All 5 tools are read-only (safe to retry, no side effects), but readOnlyHint is absent. Agents cannot distinguish safe tools from destructive ones without explicit hints.
Minimal error handling. No recovery guidance. If fetch_page fails (timeout, 404, bot detection), the response does not suggest next steps (try stealth_fetch_page, check URL, etc.). Agents cannot self-correct.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | <=2025-11-25 | v2 |
get_youtube_transcript description is minimal (5 words). Does not explain when to use it, what format transcripts are in, or what happens if captions are unavailable. Violates 10 - 1024 char guideline.
Parameter constraints underspecified. css_selector, selectors, and wait_seconds lack format/range guidance. E.g., wait_seconds has no min/max; selectors dict structure is vague ('Dict mapping field names to CSS selectors').