Local web scraping service with dynamic page support for AI agent integration. Exposes web scraping functionality as MCP tools via stdio transport and FastAPI REST API.
The WebScraper MCP has three well-named tools (scrape_url, scrape_batch, validate_selectors) with clear verb_noun patterns and action-oriented naming. Tool and parameter descriptions are present and reasonably detailed. However, schemas are partially visible but incomplete in critical ways: output schemas are defined as Pydantic models (ScrapeUrlResult, BatchScrapeResult, ValidationResult) but the actual JSON Schema representations are not shown in the source. Input parameters have types and descriptions in the function signatures, but the schemas registered with the MCP framework are not visible in the provided code. Error handling guidance is minimal, tools raise RuntimeError without actionable recovery hints for agents. Parameter constraints (e.g., URL validation, selector format) are not documented in descriptions. The server lacks error classification, rate-limiting, and per-request logLevel support expected in current protocol specs.
Scrape multiple URLs and return combined results.
Scrape a single URL and extract structured data.
Test CSS selectors against a URL to validate they find elements.
Output schemas defined as Pydantic models but actual JSON Schema representation not visible in code. While ScrapeUrlResult, BatchScrapeResult, and ValidationResult are defined, it is unclear if these are properly registered and serialized by FastMCP with complete field documentation.
Error handling lacks actionable recovery guidance. Tools raise RuntimeError('Scraping failed: {status.progress}') and RuntimeError('Result file not found'), which give agents no next steps. Errors should classify as retryable vs. user-fixable and suggest alternatives (e.g., 'Try with force_dynamic=True', 'Check the URL format').
Parameter descriptions lack format/constraint details. 'custom_selectors' is described as 'Optional CSS selectors...' but does not specify selector syntax constraints, valid patterns, or what happens if a selector is malformed. URL parameter has no validation hint (e.g., 'Must be a valid HTTP/HTTPS URL').
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
No pagination support or result-size limits documented. scrape_batch returns combined results but provides no limit, offset, or continuation token. For a batch of 100+ URLs, LLM context could be exhausted. Description should state: 'Results are capped at 50 URLs per call; use offset/limit for pagination.'
Tool composition: scrape_batch does not support dry-run or preview mode. For irreversible operations (or those that consume rate limits), agents should be able to validate selectors first via validate_selectors, but no tool composition guidance or chaining hints are provided.
validate_selectors response schema shows 'sample_matches' as Dict[str, List[str]], but description does not clarify what 'sample text' means, are these snippets, full element HTML, or truncated content? Sample data in descriptions would help, but the response field definitions are imprecise.