Decodo MCP Server - A web scraping and AI interaction server for Amazon, Google, Bing, Reddit, ChatGPT, Perplexity, and more
This is a web scraping toolkit with 18 tools that exhibit consistent patterns but significant gaps in quality. All tools follow a verb_noun naming convention (amazon_bestsellers, google_search, scrape_as_markdown) which is good. However, the server has critical issues: (1) Descriptions are present but vary in quality, some like 'Scrape Amazon Bestsellers list with automatic parsing' (54 chars) are on the short side; (2) Input schemas are visible and use Zod for type definition, which is solid, but parameter descriptions are sometimes generic ('Geo location' vs 'Geographic location as country code or ZIP code'); (3) Output schemas are NOT documented anywhere, tools return JSON text without declaring the structure, making it impossible for LLMs to plan chaining; (4) No error handling guidance, tools call external APIs with no visible retry logic, timeout handling, or recovery guidance; (5) No per-tool validation or error categorization. The codebase shows basic schema construction (e.g., amazon-pricing-tool.ts registers inputSchema with Zod) but omits output documentation entirely. Most tools are READ_ONLY with no destructive hints needed, which is appropriate. Tool names are mostly clear (search, scrape, get patterns work well), but lack context about what data each returns, e.g., both 'google_search' and 'bing_search' exist but descriptions don't clarify when to use one vs the other. Composition is weak: no evidence of chaining support or how results from one tool feed into another. Overall, this lands in 'Fair/C range' due to present but incomplete descriptions, missing output schemas, and no error recovery guidance.
Scrape Amazon Bestsellers list with automatic parsing
Scrape Amazon Product pricing information with automatic parsing
Scrape Amazon Product page with automatic parsing
Scrape Amazon Search results with automatic parsing
Scrape Amazon Seller information with automatic parsing
Scrape Bing Search results with automatic parsing
Output schemas are completely undocumented. All tools transform responses via transformResponse() and return JSON text in a 'text' field, but the actual structure of the data object is never declared. LLMs cannot plan tool chaining or downstream data extraction without knowing what fields exist.
Parameter descriptions are often too generic. 'Geo location' appears in multiple tools but is never specified as a country code, ZIP code, latitude/longitude, or ISO 3166-1 code. 'Enable JavaScript rendering' (jsRender param) lacks context on why an LLM would choose true vs false. Without specificity, LLMs default to false or hallucinate values.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-06-11 | F | 43 | - | v1 |
Search and interact with ChatGPT for AI-powered responses and conversations
Scrape Google Ads search results with automatic parsing
Scrape Google AI Mode (Search with AI) results with automatic parsing
Scrape Google Lens image search results with automatic parsing
Scrape Google Search results with automatic parsing
Scrape Google Travel Hotels search results
Search and interact with Perplexity for AI-powered responses and conversations
Scrape a specific Reddit post
Scrape Reddit subreddit results with automatic parsing
Scrape a Reddit user profile and their posts/comments
Scrape the contents of a website and return Markdown-formatted results
Capture a screenshot of any webpage and return it as a PNG image
No error handling or recovery guidance. All tools call external APIs (Decodo SDK scraping endpoints) with no visible timeout, retry, rate-limit, or error categorization. If a scrape fails, the LLM receives a raw error with no guidance on whether to retry, adjust parameters, or abandon the request. See amazon-pricing-tool.ts and others: await sapiClient.scrape() with no try/catch or error transformation.
Tool descriptions lack 'when to use' context. Multiple search tools exist (google_search, bing_search, google_ads, google_ai_mode) with no documentation explaining when to prefer one over another. Descriptions state WHAT the tool does but not WHEN an LLM should select it vs a similar tool. This forces the LLM to guess or trial-and-error through multiple tools.
Missing pagination guidance. amazon_search, bing_search, google_ads, and google_travel_hotels accept 'pageFrom' parameter but do not document whether results are capped, whether pageFrom is 0-indexed or 1-indexed, or what the total result count is. LLMs cannot reliably paginate without knowing these constraints.
Tool names for Reddit tools are inconsistent with retrieval tools. 'reddit_post', 'reddit_subreddit', 'reddit_user' all accept a 'url' parameter (not an ID or search query) but are named as if they are noun-only. Naming should be verb_noun: e.g., 'scrape_reddit_post' or 'fetch_reddit_user' to clarify they retrieve specific resources by URL, not search.