MCP Research Server with unified search and scrape capabilities for discovering URLs, crawling websites, scraping content, analyzing images, and managing proxies
The server provides 17 tools with generally well-documented purposes and reasonable parameter schemas. Most tools have clear, actionable descriptions (50-200 chars range matches A+ baseline). Parameter schemas are present and typed across all tools. However, several tools lack output schema documentation, and error handling guidance is minimal. The tool naming follows verb-noun patterns consistently. Parameter descriptions are detailed and include constraints (enums, ranges, min/max). The main gaps are: (1) no documented output schemas for any tool, (2) missing error recovery guidance in descriptions, (3) limited explanation of tool interdependencies and workflow ordering.
Analyze an image using advanced AI vision models with comprehensive understanding capabilities. Only supports remote URL.
Clear all blacklisted domains - unblock them immediately. This resets the blacklist for ALL domains, allowing them to be scraped again. Useful if domains were incorrectly blacklisted or if you want to retry after fixing connection issues.
Deep crawl a site following links (BFS or Best-First strategy). This is a DEEP CRAWL tool - it discovers and crawls pages by following links. Use this AFTER map_domain when you need actual page content.
Fetch documentation from a URL and convert to clean Markdown. CRITICAL WORKFLOW - This is a TWO-STEP process: 1) FIRST CALL: Fetch the llms.txt URL (from docs_list_sources). This returns an INDEX of markdown links, not the actual documentation. 2) READ THE INDEX: The returned markdown contains links. 3) SECOND CALL: Call this tool AGAIN with the specific documentation URL to get the actual content.
List all available documentation libraries and their llms.txt URLs. START HERE to discover which documentation libraries are available. This returns a list of llms.txt endpoints that act as indexes to documentation content.
No documented output schemas for any tool. Tools return unstructured dicts (e.g., {"total": ..., "domains": ...}) without specifying field types, presence guarantees, or nested structure. LLMs cannot plan downstream tool calls or extract data reliably without knowing the response shape.
Missing error recovery guidance. No tool description explains what to do if the tool fails or which errors are retryable. E.g., 'scrape' does not mention what happens if a domain is blacklisted, or how to recover. Agents have no guidance for error handling.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 14 | - | v1 |
List all tracked domains with their preferred scraping methods
Extract structured JSON data using pre-built schemas. Requires URL and extraction schema selection.
List available extraction schemas for use with the extract tool.
Discover URLs from a domain using sitemaps or Common Crawl. This is a URL DISCOVERY tool - it finds URLs without crawling them. Use this BEFORE scraping to understand a site's structure.
Manually rotate to the next proxy in the rotation list. Advances to the next proxy in the list (round-robin) regardless of the configured rotation strategy.
Show current proxy configuration and rotation stats. Returns the proxy list, current proxy, rotation strategy, exclusion list, and whether proxy is enabled.
Test proxy connectivity by checking the exit IP. Makes requests through the configured proxy to detect the external IP. Compares against a direct (no-proxy) request to verify the proxy is working.
Search the web and scrape the top results in one call. This is the fastest way to research a topic. It searches multiple engines, then scrapes the top results concurrently so you get actual page content — not just snippets.
Clear all domain tracking data. This resets all learned scraping methods and blacklist entries. Use this to start fresh.
Scrape a URL and extract clean markdown content. The server learns which scraping method works best per domain and automatically uses it on future requests.
Search the web using multiple search engines. Returns titles, URLs, and short snippets. Use the 'research' tool instead if you want full page content along with results.
Get scrape statistics and metrics for monitoring. Shows performance metrics including total scrapes, success rate, average duration (p50, p95, p99 percentiles), breakdown by scraping method, and top failing domains
Tool interdependencies and workflow ordering not documented. 'docs_fetch_docs' description mentions a two-step process (fetch llms.txt, then fetch specific URL), but there is no guidance on calling 'docs_list_sources' first. Similarly, 'map' and 'crawl' have an implied workflow not explicitly stated in descriptions.
Destructive and reversible tools ('reset', 'clear_blacklist') lack confirmation or dry-run patterns. These operations permanently modify state (reset clears ALL domain tracking). Description mentions 'cannot be undone' but tool has no confirmation step.
Some parameter descriptions rely on example values (e.g., 'e.g., "pending"' for status) rather than formal enum constraints. This invites LLMs to hallucinate values outside the intended set.
'docs_fetch_docs' parameter description is verbose and tutorial-style, burying the actual requirement (URL is required, can be llms.txt or specific page). Description should be concise and focus on the parameter's role, not the workflow.