An undetectable, powerful, flexible, high-performance web scraping library with MCP server capabilities for browser automation, HTTP requests, and content extraction
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
Scrapling MCP server has 11 well-defined tools with consistent naming patterns (verb_noun, clear resource-focused names). All tools have descriptions and input schemas with types. However, there are significant gaps: (1) No output schemas documented for any tool, critical for agent chaining; (2) Parameter descriptions exist but many lack actionable constraints (ranges, formats, validation rules); (3) No error handling guidance or recovery patterns visible; (4) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk classifications; (5) Missing dependencies and prerequisites in descriptions. Tools are operationally sound but lack the rigor for production agent systems.
Tools (11)
batch_fetchread only50/100
Fetch multiple URLs concurrently using a browser session
close_sessionwritesource verified67/100
Close an open browser session
extract_dataread only50/100
Extract structured data from a page using CSS selectors or XPath
fetchread onlysource verified75/100
Fetch a URL using a dynamic browser session and extract content
fetch_sessionread only50/100
Fetch a URL using an existing persistent session
interactwrite50/100
Interact with page elements using a browser session (click, type, fill forms)
No output schemas documented for any of the 11 tools. Agents cannot anticipate response fields, forcing them to infer structure from trial or context. This violates the pattern:tool and pattern:response-shaper requirements.
Parameter descriptions lack actionable constraints. E.g., 'wait' is a number with no min/max bounds (0 - 3600?). 'concurrent_count' defaults to 5 with no range. 'timeout' is undocumented for realistic ranges. LLMs cannot self-correct without constraint guidance.
Document output schemas for all 11 tools. For 'fetch', describe: {content: string, url: string, format: string (markdown|html|text), metadata: {title?: string, language?: string}}. For 'open_session', describe: {session_id: string, session_type: string, status: string}. For 'list_sessions', describe: {sessions: Array<{session_id: string, session_type: string, status: string}>, total: number}.
Add numeric constraints to all numeric parameters. E.g., 'wait: number (0 - 3600 seconds, default 0)', 'concurrent_count: integer (1 - 50, default 5)', 'timeout: number (1 - 300 seconds, default 30)'.
Implement tool annotations in the tool registration. Mark 'fetch', 'stealthy_fetch', 'static_fetch', 'fetch_session', 'batch_fetch', 'screenshot', 'extract_data' as readOnlyHint=true. Mark 'open_session', 'close_session', 'interact' as destructiveHint=true. Mark 'fetch' and 'fetch_session' as idempotentHint=true (same URL returns same content).
Add error handling guidance to descriptions. E.g., 'fetch: Returns error if URL is unreachable (try with disable_resources=true), timeout expires (increase wait or timeout param), or selector matches nothing (verify css_selector is correct). For Cloudflare blocks, use stealthy_fetch with solve_cloudflare=true.'
Clarify tool selection. Rename 'fetch' → 'dynamic_fetch' to match 'stealthy_fetch' naming convention. Add to each: 'Use dynamic_fetch for JavaScript-heavy sites (default). Use stealthy_fetch if dynamic_fetch is blocked. Use static_fetch for simple HTML sites (faster, no browser overhead). Use fetch_session to reuse an open session from open_session().'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 68 points across a rubric change (v1 → v2)
68/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
68
<=2025-11-25
v2
2026-03-09
F
0
-
v1
screenshotread onlysource verified75/100
Take a screenshot of a URL or existing session
static_fetchread only50/100
Fetch a URL using HTTP session without launching a browser
stealthy_fetchread onlysource verified73/100
Fetch a URL using a stealth browser session that bypasses anti-bot protections
Missing tool annotations despite clear risk classifications. 'fetch' and 'stealthy_fetch' should carry readOnlyHint=true. 'close_session' and 'interact' should carry destructiveHint=true. Annotations guide agent safety reasoning.
No error handling guidance or recovery patterns. E.g., 'fetch' might fail on timeout, network error, or bad selector, but no description of expected error format or how agents should recover. Agents will retry blindly without understanding the failure.
Semantic overlap between 'fetch', 'fetch_session', 'dynamic_fetch' (implied via session_type enum). Descriptions do not clarify when to choose one over another. LLMs will guess and invoke the wrong tool.
Parameter 'selectors' in extract_data is type 'object' with generic description. No example structure given (is it {field_name: selector_string}? {field_name: {selector, attribute}}?). Agents will guess and likely pass wrong payloads.
No pagination or result limits documented. 'batch_fetch' with urls array and 'list_sessions' have no maximum item counts. LLMs could request 10,000 URLs or sessions, blowing context or overwhelming the server.
Dependency hints missing. E.g., 'fetch_session' requires a session_id, but descriptions don't guide agents to call 'open_session' first. 'extract_data' with 'selectors' object lacks examples of valid structure.
fetch_sessionclose_sessionextract_datainteract
Document the 'selectors' object structure in extract_data. E.g., 'selectors: {field_name: css_selector_string, ...}, e.g., {product_name: "h1.title", price: "span.price"}. Each field is extracted via CSS selector and returned in the result object.'
Add result limits and pagination hints. E.g., 'list_sessions: Returns up to 100 sessions. For large workloads, implement pagination with offset/limit params.' 'batch_fetch: urls array limited to 50 items; concurrent_count capped at 20 to avoid server overload.'
Add dependency notes. E.g., 'fetch_session requires a session_id from open_session(). To create a new session: (1) call open_session() to get session_id, (2) call fetch_session() with that session_id.' Similar note for 'close_session' and 'interact', both require session_id from open_session().
Document mutually exclusive parameters. E.g., 'fetch: url is required; session_id is optional. If session_id is provided, reuses existing browser context; otherwise creates new temporary context (closed after fetch completes).'
Enhance descriptions with examples and use cases. E.g., 'screenshot: Useful for visual content (charts, tables, UI screenshots). Returns image as base64 or file path depending on output_format param. For full-page screenshots, set full_page=true (may be slow for long pages). For quick viewport captures, use full_page=false (default).'