MCP Server for the DiscOverflow Web Scraper. Runs a Playwright scraper as a persistent MCP server with a warm browser between requests for zero cold-start penalty.
Single tool 'scrape_website' has a reasonable description (156 chars) and proper naming convention (verb_noun). Input schema is present with type definition for 'url' parameter, but lacks parameter-level description in the schema itself. Output is formatted text rather than structured JSON, limiting composability with downstream tools. No error handling guidance, no output schema documentation, and no tool annotations. The description is informative but could be more precise about what 'structured content blocks' means and what format the output takes.
Scrapes a dynamic JavaScript-rendered website using a headless browser. Returns the page title, headings, all structured content blocks (like cards, plans, products, services), links, and the full page text. Use this tool whenever you need to fetch real, current data from any website.
Output schema not documented. Tool returns formatted string (_format_data result) but LLMs have no way to know the structure of 'headings', 'structured_blocks', 'links', and 'body_text' sections. Downstream tools cannot parse or chain results reliably.
No error handling guidance. Tool can fail on network errors, invalid URLs, or browser timeouts, but error responses only return 'Error scraping: {error}' with no guidance on whether to retry, try a different URL, or investigate prerequisites. LLM cannot decide next action.
URL parameter lacks description in schema. While the tool docstring mentions 'e.g. https://discoverflow.co/', the input schema definition in JSON has no 'description' field for the 'url' parameter. LLMs see only the type, not constraints or format expectations.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Tool returns unstructured formatted text (8000 char limit on body_text). This is lossy and wastes tokens on formatting rather than semantic content. Structured JSON output (title, headings array, blocks array, links array, body_text) would enable composability and parsing by downstream agents.
No tool annotations. Tool is read-only but lacks readOnlyHint annotation, preventing clients from optimizing caching, retry strategies, or user warnings. Current spec (2026-07-28) supports tool annotations, should be adopted.