A powerful, standalone web scraping toolkit using Playwright and various parsers, exposed as an MCP server for scraping, content extraction, metadata analysis, and autonomous web crawling.
This MCP server provides 22 tools with adequate naming and schema structure, but suffers from significant gaps in parameter descriptions and output schema documentation. Tool names follow verb_noun conventions (configure_*, get_*, list_*, clear_*, etc.), which is correct. However, many parameters lack detailed descriptions explaining their purpose, constraints, and valid ranges. Output schemas are not documented in the provided source, the rubric requires explicit documentation of what each tool returns. Configuration tools (configure_scraper, configure_stealth, configure_runtime) have dense parameter lists without constraint details. Job management tools (start_job, poll_job) use JSON-string payloads instead of structured parameters, which complicates schema validation and forces the LLM to construct JSON. Error handling descriptions are absent. The server is technically functional but falls short of production-grade definition quality.
Cancel a running async job.
Clear the response cache. Use when cached data may be stale.
Clear scraping history.
Delete one host profile record from the profile store.
Clear a browser session (cookies, storage). Use for fresh starts.
Configure host-profile auto-learning behavior.
Apply runtime override values without restarting the MCP server.
JSON-string parameters (configure_runtime, start_job, set_host_profile, run_playbook) force the LLM to manually construct JSON payloads. This bypasses schema validation, increases error likelihood, and complicates tool composition.
Output schemas are not documented for any tool. The rubric requires explicit schema documentation so LLMs can plan downstream tool calls and extract the right data. Tools like get_config, get_host_profiles, poll_job return unknown structures.
Many parameter descriptions are too brief (<20 chars) or generic. E.g., 'payload_json' in start_job lacks format guidance. 'overrides_json' in configure_runtime lacks schema. Parameters should document allowed values, constraints (min/max), and format expectations.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Configure browser settings.
Configure stealth mode and robots.txt compliance.
Get response cache statistics (hits, misses, size).
Get current configuration settings.
Get recent scraping history.
Return host profile learning store (all hosts or one host).
Finds, intelligently filters, and prioritizes URLs from a website's sitemap. This version is upgraded to handle sitemap index files (which point to other sitemaps) as well as regular sitemaps. It dynamically finds a URL from the input.
List recent async jobs and their statuses.
List all saved browser sessions.
Start a fresh browser session, clearing all existing sessions.
Get current status of a job started by `start_job`.
Reload runtime settings from config files.
Execute an Autonomous Crawl using a Playbook.
Set active host routing profile from JSON payload (admin override).
Start a long-running job and return immediately with a `job_id`. Supported job types: `batch_scrape`, `deep_research`, `run_playbook`, `batch_contacts`.
No error handling guidance. Tool descriptions do not explain what to do if a call fails. E.g., start_job does not say what happens if the job_type is invalid, or how to retry if the job times out.
configure_scraper has 18 parameters with minimal descriptions. Many lack enum/constraint specification: browser_type (string, which values?), native_fallback_policy (string, which values?), native_context_mode (string, which values?). LLMs will hallucinate invalid values.
No permission/scope declarations. Tools like clear_history, new_session, cancel_job perform destructive actions without documented authorization checks or scope requirements.
Job management lacks pagination guidance. list_jobs accepts 'limit' but does not document max/min bounds, default behavior, or whether a 'next_cursor' or offset is returned for large result sets.
No confirmation pattern for destructive operations. new_session 'clears all existing sessions' but offers no dry-run, confirm_before_execute, or undo capability. Agents may accidentally destroy session state.