Agent Scraper MCP presents a solid web-scraping toolkit with clear naming conventions and reasonable parameter structure. All 6 tools are explicitly registered with the @mcp.tool() decorator and include docstrings. However, definitions lack critical LLM-optimization details: parameter descriptions are minimal or absent in several cases, output schemas are not documented (tools return JSON strings without specifying expected field structure), and error handling guidance is completely missing. Tool names follow verb_noun convention (scrape_url, extract_links, screenshot_url) which is positive, but parameter documentation does not include format constraints, ranges, or dependency relationships. No tool includes recovery guidance for common failure modes (network errors, malformed selectors, rate limits). The 'format' enum in tool_scrape_url is good practice, but tool_scrape_structured's 'selectors' parameter lacks validation guidance and tool_search_google provides no rate-limit or relevance warnings.
Tools (6)
tool_extract_linksread onlysource verified72/100
Extract all links from a URL with their text.
tool_extract_metaread onlysource verified63/100
Extract metadata from a URL (title, description, OG tags, favicon, etc).
Output schemas completely undocumented. All 6 tools return JSON strings but callers cannot see expected field names, types, or optionality. LLMs must infer structure from docstring text alone, causing silent misinterpretations.
Parameter descriptions insufficient for LLM reasoning. 'selectors' in tool_scrape_structured has type (object, additionalProperties) but no guidance on CSS selector syntax, error recovery (partial match vs. fail), or result structure. 'filter' regex in tool_extract_links lacks flavor/flags documentation. Missing constraints like max URL length, max selector count, rate limits.
Document output schema for every tool. Use JSON Schema to specify expected response fields, types, and optionality. E.g., tool_scrape_url returns {"title": string, "content": string, "url": string, "metadata": {"author": string|null, "publishDate": ISO8601|null}}.
Expand parameter descriptions to include constraints, examples, and recovery guidance. E.g., 'format' param: 'Output format: text (plain text), markdown (GitHub-flavored), or html (raw HTML). Default: markdown. Use markdown for most LLM processing; use html only if you need CSS selector precision or styling info.'
Add error handling sections to every tool docstring. E.g., 'Returns error if: (1) URL is invalid or unreachable (retryable after 30s), (2) Content encoding is not UTF-8 (not retryable, suggest extracting as HTML instead), (3) Rate limited by target (not retryable, backoff 60s).'
For tool_scrape_structured, document selector behavior: 'If a selector matches multiple elements, returns first match. If no match, returns null for that field. If selector is invalid CSS, tool fails with error message listing valid fields.'
For tool_extract_links, document regex behavior: 'filter parameter accepts Python regex (re module). Case-sensitive by default. E.g., filter="^https://" to include only HTTPS links. Invalid regex raises error with suggestion.'
For tool_search_google, add usage guidance: 'This tool searches Google and returns organic results. Google may block bots; if rate-limited, wait 60s before retrying. Do not use for production systems relying on 100% availability; consider search API alternatives. Respects robots.txt and serves from Google cached results when available.'
Score history
Overall score trend
↑ 1 points across a rubric change (v1 → v2)
68/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
C
68
2026-07-28+
v2
2026-03-09
C
67
-
v1
Zero error handling guidance. No tool describes recovery for common failures: network timeouts, malformed URLs, non-UTF8 content, rate limiting (Google search), JavaScript-rendered content not visible in HTML. LLMs have no guidance on whether errors are retryable, fatal, or user-fixable.
No guidance on mutual exclusivity or parameter dependencies. e.g., tool_scrape_url format='html' may be incompatible with certain downstream processing. tool_screenshot_url full_page=true may cause excessive memory. tool_search_google num_results beyond 10 may hit rate limits. No documentation of these constraints or LLM warnings.
Example values embedded in parameter descriptions ('e.g. {"title": "h1.title", "price": ".price"}') risk LLM literal reuse instead of adaptation. Use enum constraints or regex patterns instead of examples in descriptions.
tool_search_google does not document Google ToS risks, rate limiting, or bot-blocking behavior. Legal and operational concerns should be surfaced in tool description so agents understand when this tool is appropriate.
tool_search_google
Add composition guidance in tool descriptions. E.g., 'After extracting links with tool_extract_links, use tool_scrape_url on each link to fetch content. For structured data extraction across multiple URLs, batch with tool_scrape_structured to reduce latency.'
For tool_screenshot_url, document timeout and render behavior: 'Viewport rendered with Chromium. JavaScript executed for up to 30s. If page does not stabilize, times out and returns partial screenshot. full_page=true captures entire scrollable height but may timeout on infinite-scroll pages.'
Add data retention and privacy notes: Do these tools cache content? Log requests? Are they GDPR-compliant? Document any PII concerns when scraping user-generated content.
Consider adding a result limit warning to tool_extract_links and tool_search_google: 'Returns up to 100 links/results. For larger datasets, paginate or filter. Returning 500+ items risks context window exhaustion and degraded LLM reasoning.'