MCP server exposing PyScrappy scrapers as agent tools. Each tool wraps a PyScrappy scraper and returns typed ScrapeToolResult with declared output schemas and validated structuredContent.
PyScrappy MCP defines 3 tools with strong, detailed descriptions and complete input schemas. All tools feature comprehensive, LLM-optimized descriptions (200-500+ chars) that clearly state WHAT the tool does, WHEN to use it, and any prerequisites, well above the 194-char production median. Parameter descriptions are present and actionable. However, output schemas lack explicit documentation in the tool definitions themselves, and no error handling or recovery guidance is embedded in tool descriptions. Schemas are JSON-typed but could benefit from explicit enum constraints on mode/interval parameters. All tools are read-only (no state mutation), which is good, but the lack of error categorization and recovery patterns prevents a higher score.
Fetch stock market data from Yahoo Finance and return it as a dict. The returned shape depends on `mode`: - "quote": {"symbol", "currency", "exchange", "price", "previous_close", "volume", "day_high", "day_low", "fifty_two_week_high", "fifty_two_week_low"}. - "history": {"symbol", "period", "rows": [{"date", "open", "high", "low", "close", "volume"}, ...]}. - "profile": {"symbol", "name", "currency", "exchange", "market", "timezone", "instrument_type"}. Behavior: Makes a live network request to Yahoo Finance on each call; no browser is required and nothing is cached or persisted. If the symbol is unknown or Yahoo returns no data, "rows" is an empty list (mode="history") or the remaining fields are None (mode="quote"/"profile"); no exception is raised for an empty result. Args: symbol: Ticker symbol as a string. Example: "AAPL". No default (required). mode: String selecting what to fetch; one of "quote", "history", "profile". Example: "quote". Default: "quote". period: String history window, used only when mode="history" and ignored otherwise; one of "1d", "5d", "1mo", "3mo", "6mo", "1y", "2y", "5y", "10y", "ytd", "max". Example: "1y". Default: "1mo". interval: Candle size for history bars, used only when mode="history"; one of "1d", "1wk", "1mo". Example: "1wk". Default: "1d". Usage: Use for a single ticker's live quote, OHLCV history, or company profile; confirm the ticker is valid with list_available_scrapers or a quote lookup before requesting history.
Scrape any HTTP(S) URL and return a ScrapeToolResult whose `data` holds one object per page containing extracted text (with word_count), links, images, tables, and page metadata. Fetches the page over the network and parses the HTML; no data is stored or mutated. By default it makes a plain static HTTP request, so pages built client-side with JavaScript come back nearly empty. When that is detected, the returned `errors` list gets a hint to retry with render_js=true; set render_js=true to render with a headless browser instead (requires the pyscrappy[browser] extra). On empty or failed results, `data` is [], `count` is 0, and `errors` describes the problem rather than raising. Returns: ScrapeToolResult with fields: data (list of per-page dicts holding text, links, images, tables, metadata; also the keys named in `selectors` when provided), count (int number of items in data), scraper (str backend name), source_urls (list of str URLs actually fetched, one per page), and errors (list of {url, message} for non-fatal problems). Args: url: String, the page URL to scrape including scheme, e.g. "https://example.com/products". Required, no default. selectors: Optional dict mapping output field name to CSS selector to extract specific values into each data item, e.g. {"title": "h1", "price": ".amount"}. Default None (returns only the standard text/links/images/tables/metadata). max_pages: Integer, follow "next"-style pagination up to this many pages, e.g. 3. Default 1 (scrape only the given URL). render_js: Boolean, render JavaScript with a headless browser backend, e.g. True. Default False; allowed values True or False, and True needs the pyscrappy[browser] extra installed. Use this for arbitrary or unsupported sites; for common sources prefer the purpose-built siblings (scrape_wikipedia, scrape_stock, scrape_news, search_amazon, etc.), which return cleaner fields. Gotcha: if results look empty on a modern site, re-call with render_js=true or scrape the underlying data endpoint the page fetches.
Output schemas not documented in tool definitions. ScrapeToolResult is defined in code but not exposed as a formal tool output schema annotation. Agents cannot see the shape of returned data without reading implementation.
scrape_url mode parameter lacks explicit enum constraint in schema. Accepts 'full'|'paragraphs'|'headers' but schema shows type:string without enum. LLMs may hallucinate invalid values.
No error recovery guidance in descriptions. Tools note that failed results return errors in a list, but descriptions do not guide agents on what actions to take (e.g., 'if render_js is false and results are empty, retry with render_js=true' is buried in scrape_url description but not actionable from LLM's perspective).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 76 | 2025-06-18+ | v2 |
Fetch a Wikipedia article by title or search term and return its text content. Makes a live network request to Wikipedia, resolving the query to the best-matching article and extracting its body. The shape of the returned text depends on `mode`: "full" returns the entire article as one string; "paragraphs" returns the article split into a list of paragraph strings; "headers" returns a list of the article's section heading strings (its table of contents). If no article matches the query, an empty result is returned (empty string for "full", empty list for "paragraphs" or "headers"). Args: query: String. Article title or search term. Example: "Model Context Protocol". No default (required). mode: String, one of "full", "paragraphs", or "headers". Selects the return shape as described above. Example: "paragraphs". No default (required). Usage Guidelines: Use when you need the content of a known Wikipedia topic; pick "headers" first to survey structure, then "paragraphs" or "full" to pull the text. Requires network access and returns empty on a miss, so verify the query resolved before relying on the output.
scrape_stock period/interval parameters have allowed values ('1d', '5d', '1mo', etc.) documented as example text in description rather than as formal enum constraints in JSON schema. Prevents LLM from discovering valid options via schema alone.
No tool-level annotations (readOnlyHint, idempotentHint). All three tools are read-only and idempotent, but this is not formally signaled in the MCP tool definition.