Web scraping MCP server with Playwright support for extracting content from websites and WordPress APIs
This server has a single tool (extract_site_links) with critical definition gaps. The tool name is clear and action-oriented (extract_*), but the input schema is severely underdocumented. The input parameter 'url' has a type (string) and a basic description in Japanese, but lacks validation constraints (URL format, length limits, protocol requirements). No output schema is documented anywhere in the visible code, we cannot see what structure is returned or what fields the LLM should expect. The tool description is in Japanese ('公式サイトからheader/footer/navのリンクを抽出し、仮想サイトマップを作成する') and provides intent but is vague on what exactly is returned and when to use this tool vs. alternatives. The code shows heavy scraping logic (BeautifulSoup, Playwright, aiohttp) but the tool definition does not expose this complexity or document constraints (timeout, max page size, SSL handling, retry behavior). No error recovery guidance is visible, if a URL is unreachable or scraping fails, the LLM has no hint what to do next. The visible code includes extensive utility functions for Rakuraku/Google Sheets/Cloud Gym integrations, but none of these appear as registered MCP tools, suggesting they are either helper functions or incomplete tool definitions. Only extract_site_links is registered; the others are inferred from code presence, which violates the 'explicit registration' requirement.
公式サイトからheader/footer/navのリンクを抽出し、仮想サイトマップを作成する
No output schema documented. The tool description does not specify what fields or structure are returned (list of links, objects with URL+text, sitemap format, etc.). LLMs cannot plan downstream calls or extract relevant data without knowing the response shape.
Input parameter 'url' lacks validation constraints. No mention of URL format requirements, protocol (http/https), length limits, timeout behavior, or what happens with invalid/unreachable URLs. The description is in Japanese with a single example; English is expected for universal LLM compatibility.
No error recovery guidance. The tool will fail on unreachable URLs, timeouts, or large pages, but the description does not explain what the LLM should do (retry, check URL, pass smaller page). Error responses will likely include stack traces instead of actionable guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Tool description is minimal and in Japanese. At ~20 characters in English equivalent, it fails the 10-1024 character guideline for effective LLM selection. No explanation of when to use this tool, prerequisites (e.g. site must be publicly accessible), or what 'virtual sitemap' means.
Partial tool definitions appear in code (Rakuraku WordPress API, Cloud Gym API, Google Sheets integration) but are NOT registered as MCP tools. These appear as utility functions (_rakuraku_wp_get, _cloud_gym_fetch, etc.) rather than exposed tools. Either these should be removed or properly registered with full schemas and descriptions.