Model Context Protocol (MCP) server for web scraping integration with Claude Desktop and other MCP clients. Provides advanced web scraping capabilities including support for Instagram, Bluesky, and generic websites with image and video downloading, metadata extraction, and deduplication.
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The server defines 3 tools with substantial parameter schemas, but critical gaps in descriptions, error handling, and schema documentation significantly limit usability for agentic calls. Schemas are present but output structures are not documented. Descriptions exist but are often vague and lack context for tool selection. No evidence of error recovery guidance, parameter validation rules, or input constraints beyond basic type declarations. The tools expose file-write operations (risk: WRITE) without permission gates or destructive operation guards. Tool naming is reasonable (scrape_*) but parameter design has issues around identifier types and validation.
Output schemas not documented. The CallToolResult responses have no declared structure, so LLMs cannot plan downstream steps or extract returned IDs/paths. Critical for composition.
No error recovery guidance. Tools accept parameters like output_dir, timeout_seconds, max_files with no validation rules stated. If a user passes output_dir='/invalid/path' or timeout_seconds=0, the LLM has no guidance on acceptable values or what to do on failure.
Destructive write operations lack confirmation or dry-run. These tools write files to disk (risk: WRITE). No evidence of dry_run parameter, permission gates, or confirmation steps. An agent could download gigabytes of unintended content without warning.
Document the output schema for each tool. Example: 'Returns {"downloaded_files": [{"path": string, "width": int, "height": int, "metadata": object}], "total_files": int, "duplicates_moved": int}'. This enables LLMs to extract paths and plan next steps.
Add error recovery guidance to tool descriptions. Example: 'On network timeout, the LLM should retry with a longer timeout_seconds. On permission denied, check that output_dir is writable and user has disk space.'
Introduce a dry_run parameter (boolean, default false) to scrape_web_images. When true, return a count and sample URLs of what would be downloaded without actually writing files. This lets agents preview results before committing.
Specify valid ranges and constraints in parameter descriptions. Example: 'timeout_seconds (default 100): Maximum seconds to wait for page load. Must be 10 - 600. Lower values may fail on slow connections; higher values delay failure detection.'
Clarify tool selection guidance. Rewrite scrape_web_images description to: 'Scrape images and videos from any website (generic). Use scrape_instagram_profile for Instagram accounts (includes stories and metadata); use scrape_bluesky_profile for Bluesky accounts. Requires network access and writeable output directory.'
Normalize username parameter expectations. State clearly: 'Instagram username without @ prefix (e.g., "john_doe" not "@john_doe")' and 'Bluesky handle with or without .bsky.social suffix (e.g., "alice.bsky.social" or "alice")'. Or accept both formats and normalize inside the tool.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Parameter descriptions lack actionable constraints. 'timeout_seconds' has no min/max. 'max_files' (default 1000) has no guidance on reasonable bounds, can an agent pass 1,000,000? 'min_width' and 'min_height' default to 512 but no explanation why or what breaks if set to 0.
Tool descriptions are too generic. 'Scrape images and videos from websites with advanced options' does not explain WHEN to use this vs scrape_instagram_profile or scrape_bluesky_profile. No guidance on prerequisites (authentication, stealth mode requirements) or side effects (bandwidth, rate limits, legal considerations).
Parameter identifier design forces guessing. 'username' in scrape_instagram_profile and scrape_bluesky_profile lacks clarity: Is '@username' accepted or just 'username'? Description says 'without @' for Instagram and 'with or without .bsky.social' for Bluesky, inconsistent and error-prone.
Missing parameter relationships documentation. scrape_web_images has 'download_images' and 'download_videos' both defaulting true, what happens if both are false? Are they mutually exclusive? 'use_stealth_mode' may require setup, is auth_config.json pre-populated?
No audit trail or permission gates. These tools make network requests and write files on behalf of an agent. No evidence of logging who called what, with which parameters, or when. No permission checks (e.g., 'agent can only scrape domains in allowlist').
Missing tool-chaining identifiers. After scrape_web_images downloads images, what does it return? If it returns a list of downloaded file paths, are those absolute paths, relative to output_dir, or URLs? Tools that follow may need to process these files, the response must enable that.
Deduplication options lack clarity. scrape_web_images offers 'hash_algorithm' enum with 'average_hash|phash|dhash|whash|none' but no guidance on which to use for what content type, performance tradeoffs, or false-positive rates. An LLM has no basis to choose.
scrape_web_images
Document parameter relationships and prerequisites. Example: 'If include_stories=true, authentication via auth_config.json is required. Omit this parameter if cookies are not configured.'
Add logging/audit hooks. Wrap tool calls with structured logging: {"timestamp": ISO8601, "tool": tool_name, "params": sanitized_params, "agent_id": agent_id, "result": {success, files_downloaded, errors}}. Log to stderr or a file for compliance.
Expand responses with chaining data. After scrape_web_images, return: {"success": bool, "files": [{"relative_path": string, "absolute_path": string, "url_source": string, "width": int, "height": int, "hash": string}], "stats": {"downloaded": int, "deduplicated": int, "failed": int}, "next_action": "Files ready in /path. Call process_images or archive_to_zip if needed."}. This guides multi-step workflows.
Add enum guidance for deduplication. Document: 'hash_algorithm options: average_hash (fast, perceptual, ~5% false positives for thumbnails), phash (slower, better for cropped images), dhash (difference-based, good for color variations), whash (wavelet-based, best for JPEG compression artifacts), none (no dedup). Default average_hash is recommended for speed. Use phash for art/photo galleries.'
Implement permission gates. Check that output_dir is writable and agent has 'scrape:write' scope before executing. Return a clear error if denied: 'Permission denied: agent lacks scrape:write scope. Contact admin to grant this permission before retrying.'