Web scraping MCP server that handles both static HTML and JavaScript-rendered pages with data extraction capabilities
WebScraper MCP provides 5 well-structured tools with complete JSON schemas and detailed descriptions. All tools have clear verb-based names (scrape_url, extract_data, extract_first, batch_scrape, crawl_website) and comprehensive parameter documentation. However, several quality issues prevent a higher score: (1) Output schemas are not explicitly documented in the code, responses are inferred from return type hints and docstrings but lack formal schema declarations; (2) No error handling guidance, when scraping fails, errors are returned as simple dicts without actionable recovery hints (e.g., 'Try checking the URL format' or 'Enable javascript=True for dynamic sites'); (3) No pagination support for batch_scrape or crawl_website, large result sets could exhaust context; (4) Parameter descriptions contain example values ('e.g. ["h1", "a.link", "#content"]') which LLMs may reuse literally; (5) No input validation guidance, descriptions don't specify format constraints, ranges, or why a parameter matters. All parameters are properly typed with descriptions, and tool names follow verb_noun convention. The fastmcp framework provides decent structure, but the implementation lacks production-grade polish around error recovery, result limits, and output consistency.
Scrape multiple URLs efficiently. Args: urls: List of URLs to scrape javascript: Set to True if the sites need JavaScript rendering Returns: List of scraping results for each URL
Crawl a website to discover its structure and pages. Args: start_url: Starting URL max_pages: Maximum pages to crawl (default 50) max_depth: Maximum link depth (default 3) same_domain_only: Stay on same domain (default True) Returns: Site map with discovered pages and statistics
Scrape a webpage and extract specific data using CSS selectors. Args: url: The webpage to scrape css_selectors: List of CSS selectors (e.g., ["h1", "a.link", "#content"]) attributes: List of attributes to extract for each selector (e.g., ["text", "href", "text"]) If not provided, defaults to "text" for all selectors javascript: Set to True for JavaScript-rendered sites Returns: Dictionary with extracted data for each selector Example: extract_data( url="https://example.com", css_selectors=["h1", "a"], attributes=["text", "href"] )
Extract the first matching element from a webpage. Useful for getting single values like page title, main heading, etc. Args: url: The webpage to scrape css_selector: CSS selector for the element (e.g., "h1", "title", "meta[name='description']") attribute: What to extract - "text" for content, or attribute name like "href", "content", "src" javascript: Set to True for JavaScript-rendered sites Returns: Dictionary with the extracted value Example: extract_first(url="https://example.com", css_selector="title", attribute="text")
Output schemas not formally documented. Return types are inferred from Python type hints and docstrings, but there is no explicit JSON schema declaration for tool responses. LLMs cannot predict response structure without seeing it in code or documentation.
No error recovery guidance. Error responses are bare dicts with 'success': false and 'error': str(e). LLMs receive no actionable hints, they don't know if the error is retryable, user-fixable, or fatal. Example: 'Connection timeout' with no hint to try a different URL or enable javascript=True.
No pagination or result-limit enforcement. batch_scrape and crawl_website accept unbounded lists/depths. A crawl with max_depth=3 and max_pages=10000 could return gigabytes of HTML, exhausting context and causing timeouts. No mention of result limits in descriptions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Scrape a webpage and return its HTML content. Args: url: The webpage URL to scrape javascript: Set to True for JavaScript-rendered sites (slower but handles dynamic content) wait_seconds: How long to wait for JavaScript to load (only used when javascript=True) Returns: Dictionary with html content, status code, and load time
Parameter descriptions include example values that LLMs may reuse literally. Example: 'List of CSS selectors (e.g., ["h1", "a.link", "#content"])', LLMs often treat 'e.g.' as a template and pass those exact values in subsequent calls rather than adapting to context.
Undefined behavior for partial failures. batch_scrape and crawl_website do not document what happens if some URLs fail or some links are unreachable. Do they return partial results (49 of 50)? Or fail entirely? This ambiguity forces the LLM to guess the error model.
No idempotence or replay guarantees documented. Scraping is inherently non-idempotent (page content changes, JavaScript renders differently), but tools don't state this. If an LLM retries after a timeout, it may scrape a different version of the page, causing silent inconsistencies.