Web scraping server using BeautifulSoup to fetch and parse HTML pages, storing results in SQLite
This server has significant gaps in tool design quality. While schemas are visible and mostly well-formed, descriptions lack sufficient detail for LLM decision-making, parameter annotations are sparse, and error handling does not guide recovery. The tool names are reasonable but descriptions often fall below the 50-200 character optimal range for LLM-optimized discovery. Most tools lack parameter descriptions entirely (e.g., scrape_pages_store_sqlite has no description for 'num_pages' parameter). The server exposes database semantics directly (SQLite table names, raw page_id references) rather than abstracting to user-friendly concepts. Error responses are minimal ('ok': False, 'error': str(e)) and do not guide the LLM on retryability or recovery steps.
Fetch a single URL and return HTTP status and the page <title>.
Health check for the MCP server. Returns a success message if running.
Read-only access to the `page_text` table. Parameters: - page_id: filter by specific page id (optional) - contains: return rows where text content LIKE %contains% (optional) - limit: max rows (default 50, max 500) - offset: pagination offset
Scrape one or multiple pages with BeautifulSoup and store into SQLite. If the URL contains '{page}', it will be replaced by page numbers starting at 1. Otherwise a 'page' query parameter will be added or replaced.
Tool 'scrape_pages_store_sqlite' exposes database implementation details (SQLite table names, page_id foreign keys) rather than abstracting to user-friendly concepts. LLMs must reason about which table to query rather than natural language intents like 'extract all product prices'.
Parameter 'num_pages' in scrape_pages_store_sqlite lacks description, type annotation, and constraints. Schema shows 'integer' type but no guidance on valid range (1-100? 1-1000?), default behavior (replaces or adds page param?), or what happens when URL lacks pagination capability.
Error handling across all tools returns bare exception strings ('error': str(e)) without classification (retryable? user-fixable? fatal?), recovery guidance, or actionable messages. LLM cannot determine if network timeout should trigger retry or if invalid URL is permanent.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 9 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Tool 'scrape_pages_store_sqlite' name combines two responsibilities: scraping (reads) AND storing (writes). Should split into 'scrape_page' (returns parsed content) + 'store_scraped_content' (writes to DB). Current naming violates single-responsibility principle.
Tool 'read_page_text' exposes raw database query parameters (page_id integer PK, LIKE operator semantics, limit/offset pagination) that require SQL knowledge. Should expose 'search_page_content(text, page_reference)' with natural filtering, not 'contains' LIKE syntax.
No output schema documented for any tool. LLMs cannot plan downstream operations or extract fields without reverse-engineering from code. E.g., does 'scrape_pages_store_sqlite' return page_id only, or full parsed content (title, links, images)?
fetch_page_title description does not explain what to do when page has no title or fetch fails (e.g., 'Returns empty string if no <title> tag found; retry on 5xx status; returns error object on timeout'). LLM must infer behavior from code, not description.