FastAPI-based MCP server for web scraping with progressive fallbacks. Supports HTTP requests, BeautifulSoup parsing, and headless/headful browser automation via Playwright. Includes search capabilities via Google and DuckDuckGo.
The MCP Scraper Server exhibits mixed definition quality. Tool naming follows verb-first conventions (search_*, scrape_*, browser_*) which is good, but parameter descriptions are sparse or missing in critical areas. Descriptions exist for tools but many are terse (e.g., 'Close current browser session' is only 28 chars). Input schemas are present and include types, but lack granular constraints (enums, ranges). Output schemas are not documented, tools return JSON but the structure is not formally declared. Error handling is minimal: validation exists for URLs and input size, but error messages lack recovery guidance or classification. The tool definitions are visible in server.py with explicit @mcp.tool() registration, so inference is not required.
Click element by selector
Close current browser session
Evaluate JS in browser
Get text content from page
Navigate to URL in browser session (creates if none)
Take screenshot of page or element
Extract clean content from HTML using Trafilatura
Missing output schema documentation. Tools return JSON but the response structure is never formally declared. Tools like search_query, scrape_url, and browser_navigate all return json.dumps() output, but LLMs cannot infer field names, types, or which fields are IDs for chaining to downstream tools.
Terse or missing parameter descriptions. browser_click selector parameter has description 'CSS selector for element to click', but browser_get_text and browser_close have empty input objects with no parameter guidance. browser_close description is only 28 characters ('Close current browser session'), offering no context on side effects or prerequisites.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 49 | - | v1 |
Concurrent scraping of multiple URLs
Scrape a single URL with progressive fallback chain
Concurrent search for multiple queries
Perform search using Google and DuckDuckGo with fallbacks, returns list of URLs
No enum constraints for multi-valued inputs. browser_navigate, scrape_url, and search_query all accept free-form strings. No validation or enum hints guide the LLM on valid formats, retry strategies, or expected value ranges. max_retries defaults to 3 and num_results defaults to 10, but no min/max constraints are declared.
Error handling lacks recovery guidance. is_valid_url() and HTML size checks raise HTTPException with messages like 'Only HTTP/HTTPS URLs are allowed...', which is actionable, but most tools return bare success/failure. browser_click returns 'Clicked {selector}' or 'No browser session', the latter does not suggest calling browser_navigate first. No retry classification or next-step hints.
Stateful browser session management with implicit session IDs. Tools rely on a global _sessions dict keyed by uuid, and browser_click/browser_evaluate/browser_screenshot/browser_get_text/browser_close all use list(_sessions.keys())[-1] to select 'the last' session. This is fragile: concurrent requests will interfere. No session ID parameter is exposed to the LLM, so the agent cannot reason about or manage multiple parallel sessions.
scrape_url and scrape_multiple function names shadow module-level function definitions. server.py imports 'scrape_multiple' from fallback.py and also decorates a tool @mcp.tool() async def scrape_multiple(...). This creates a naming collision that will cause a runtime error when the second definition overwrites the import.
Insufficient validation and error classification. browser_evaluate includes a simplistic regex check (startswith('//'), contains 'prompt(' or 'alert(')) but does not validate that the script is legal JavaScript. An LLM can still pass valid JS that queries internal state or triggers unintended side effects. Error responses ('Script contains potentially malicious patterns') do not classify as retryable or suggest alternatives.
Missing parameter description for max_concurrent defaults. search_multiple and scrape_multiple both accept max_concurrent with defaults (5 and 10), but no description explains what happens if this exceeds available resources, or what range is safe (1 - 100?). LLMs may pass extreme values.