Oxylabs MCP server providing web scraping, AI-powered data extraction, and search capabilities
Mixed quality across 10 tools. Strengths: comprehensive parameter schemas with enums and sensible defaults, well-structured descriptions that explain WHEN and WHY to use tools, clear output format options (markdown, json, csv, toon). Weaknesses: output schemas are NOT documented (critical gap for LLM chaining), error handling is implicit rather than explicit, descriptions lack specifics about what data is returned and how to use it downstream. Parameter naming is clear (ai_crawler, ai_scraper, etc. start with verbs) but parameters themselves have some ambiguity, e.g., 'schema' in multiple tools is nullable but the relationship between 'output_format' and 'schema' is only explained via prose, not schema constraints. Tool composition is weak: tools like ai_crawler, ai_scraper, ai_browser_agent, ai_search, ai_map, and universal_scraper overlap significantly in purpose, and there's no clear guidance on which to use when. No evidence of output structure documentation, pagination, or error recovery patterns. Security: no visible secrets/auth handling in the sampled code, though tools accept geo_location parameters.
Run the browser agent and return the data in the specified format. This tool is useful if you need navigate around the website and do some actions. It allows navigating to any url, clicking on links, filling forms, scrolling, etc. Finally it returns the data in the specified format. Schema is required only if output_format is json, csv or toon. 'task_prompt' describes what browser agent should achieve
Tool useful for crawling a website from starting url and returning data in a specified format. Schema is required only if output_format is json, csv or toon. 'render_javascript' is used to render javascript heavy websites. 'return_sources_limit' is used to limit the number of sources to return, for example if you expect results from single source, you can set it to 1.
Get map information and data based on the provided URL.
Scrape the contents of the web page and return the data in the specified format. Schema is required only if output_format is json or csv. 'render_javascript' is used to render javascript heavy websites.
Search the web based on a provided query. 'return_content' is used to return markdown content for each search result. If 'return_content' is set to True, you don't need to use ai_scraper to get the content of the search results urls, because it is already included in the search results. if 'return_content' is set to True, prefer lower 'limit' to reduce payload size.
Output schemas are completely undocumented. Tools accept schema parameters and return data in json/csv/markdown/toon formats, but nowhere in the visible code is the STRUCTURE of the returned data defined. LLMs cannot plan downstream tool calls or extract specific fields without knowing what they will receive. This blocks the most critical LLM use case: chaining tools together.
Severe tool overlap and ambiguous selection. Tools ai_crawler, ai_scraper, ai_browser_agent, ai_search, ai_map, and universal_scraper all crawl/scrape web content with nearly identical parameters and purposes. The descriptions do not clearly explain when to use one vs another. E.g., when should an LLM choose ai_crawler over ai_scraper? The distinction is buried in descriptions ('starting url' for crawler, 'contents of the web page' for scraper) and is not actionable. This forces the LLM to guess or try multiple tools.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Scrape Amazon products. Supports content parsing, different user agent types, domain, geolocation, locale parameters and different output formats. Supports Amazon specific parameters such as currency and getting more accurate pricing data with auto select variant.
Scrape Amazon search results. Supports content parsing, different user agent types, pagination, domain, geolocation, locale parameters and different output formats. Supports Amazon specific parameters such as category id, merchant id, currency.
Generate JSON schema based on the user prompt.
Scrape Google Search results. Supports content parsing, different user agent types, pagination, domain, geolocation, locale parameters and different output formats.
Get a content of any webpage. Supports browser rendering, parsing of certain webpages and different output formats.
Conditional parameter requirements are documented only in prose. Tools declare 'schema' as nullable with default null, and descriptions say 'Schema is required only if output_format is json, csv or toon.' This constraint is not machine-parseable (no dependentSchemas or conditional validation). If an LLM calls ai_crawler with output_format='json' but omits schema, the tool will likely fail, and the error guidance is not visible in the tool definition.
generate_schema has an extremely short and vague description: 'Generate JSON schema based on the user prompt.' This does not explain when to call it, what it returns, or how it integrates with other tools. A LLM will not know whether this generates a schema for ai_crawler, or for some other use case. Description is 69 chars, at the floor of the 10-1024 baseline but lacks actionable context.
No error recovery guidance. The tool definitions provide no hints about what errors might occur (rate limits, blocked sites, invalid geo_location codes, timeout on slow sites) or what the LLM should do next. For example, ai_crawler has render_javascript=false by default with a note 'Unless user asks to use it, first try to crawl the page without it. If results are unsatisfactory, try to use it.', this is a retry hint in the parameter description, not a documented error pattern. Production tools should define error classes and recovery strategies explicitly.
Geo_location parameter accepts ISO country codes but descriptions provide minimal guidance. The parameter says 'Two letter ISO country code to use for the crawl proxy', but what happens if an invalid code is passed? Does it default to US? Raise an error? The LLM has no guidance. Additionally, some tools accept 'California, United States' as an example (google_search_scraper, amazon_search_scraper) while others say 'Two letter ISO', creating inconsistency.
Pagination and result limits are not explicitly handled in most tools. ai_search has 'limit' (max 50), google_search_scraper has 'limit' and 'pages', amazon_search_scraper has 'pages' and 'limit', but there is no documentation of what happens when results exceed the limit, whether pagination is automatic, or how the LLM should paginate through large result sets. No 'next_cursor' or 'total_count' fields are documented.