Context-optimized MCP server for web scraping functionality with support for markdown conversion, HTML extraction, text extraction, link extraction, and optional Perplexity AI integration
Mixed quality across 8 tools. Core scraping tools (scrape_url, scrape_url_html, scrape_url_text, scrape_extract_links) have well-defined input schemas with proper types and reasonable descriptions, but descriptions are somewhat generic and lack context about when to choose one tool over another. Cache management tools (cache_stats, cache_clear_expired, cache_clear_all) have minimal or missing descriptions. The perplexity tool has adequate schema but poor naming (no verb prefix) and relies on external API key injection. No tools declare output schemas explicitly. Error handling guidance is absent across all tools. The server lacks tool annotations (readOnlyHint, destructiveHint, idempotentHint) that would clarify tool behavior to LLMs.
Clear all entries from HTTP cache. WARNING: This will remove all cached responses.
Clear expired entries from HTTP cache.
Get HTTP cache statistics.
Engages in a conversation using Perplexity to search the internet and answer questions. Accepts an array of messages (each with a role and content) and returns a chat completion response from the Perplexity model.
Scrape one or more URLs and extract all links.
Scrape one or more URLs and convert the content to markdown format.
Scrape raw HTML content from one or more URLs.
Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint). LLMs cannot distinguish safe reads from destructive operations like cache_clear_all without explicit hints. This is a critical clarity gap given the presence of destructive cache tools.
No output schemas documented. Tools return results but LLMs cannot plan downstream operations (e.g., what fields does scrape_url return? what structure do the links have in scrape_extract_links?). This forces agents to guess at response structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 62 | - | v1 |
Scrape one or more URLs and extract plain text content.
Cache management tools lack meaningful descriptions. cache_stats has only 'Get HTTP cache statistics.' (30 chars), cache_clear_expired has only 'Clear expired entries from HTTP cache.' (37 chars). These are below the 10 - 1024 char guideline minimum useful length (~50 chars). LLMs cannot determine when to invoke these tools.
Perplexity tool lacks action verb in name. Named 'perplexity' instead of 'search_perplexity' or 'query_perplexity'. This violates the verb_noun naming convention (get_, create_, search_, etc.) that helps LLMs infer intent from the name alone before reading descriptions.
Scrape tool descriptions are generic and lack context for choosing among variants. 'Scrape one or more URLs and convert the content to markdown format' (scrape_url) doesn't explain when to use this vs scrape_url_text vs scrape_url_html. LLMs waste reasoning cycles or pick wrong variants.
No error handling guidance. Tools accept parameters like css_selector, render_js, timeout, but provide no recovery hints if these fail (e.g., 'If JavaScript rendering fails, try without render_js=true', or 'Invalid CSS selector, check syntax with browser console first'). LLMs get no actionable error recovery path.
Perplexity API key is exposed as environment configuration (PERPLEXITY_API_KEY). While tools don't accept it as a parameter (good), the tool description doesn't clarify that the server must be pre-configured with valid credentials. An agent cannot detect or communicate API key setup failures to the user.
Overlapping tool responsibility. scrape_url, scrape_url_html, and scrape_url_text all accept the same core parameters (urls, timeout, max_retries, css_selector, include_headers, render_js) but return different formats. The naming convention makes the distinction clear, but parameter duplication and lack of guidance on when to use each creates friction.