Advanced Web Crawling Platform with Deep Analysis and MCP Server
The Crawlemoon MCP server provides 16 well-named tools with consistent verb-noun patterns (deep_analyze, discover_apis, extract_content, etc.). All tool descriptions are present and substantive (ranging 60-250+ characters), meeting the 10-1024 character baseline. Input schemas are fully defined with proper JSON Schema structure, including types, enums, and descriptions for most parameters. However, there are notable gaps: (1) output schemas are NOT documented anywhere in the provided code, only input schemas are visible; (2) several parameters lack descriptions (e.g., 'endpoint' in introspect_graphql has minimal context); (3) error handling guidance is absent, tools do not document what the LLM should do on failure; (4) no security/permission scopes declared despite tools performing sensitive operations (JavaScript deobfuscation, API discovery, bot detection); (5) composition risk: tools like extract_smart and analyze_javascript operate on externally-hosted websites without explicit consent warnings; (6) no pagination or result-limiting documented despite tools like list_recordings potentially returning large datasets. The server demonstrates solid naming discipline and parameter constraint patterns (enums for depth, format, provider), placing it in the C+ to low-B range. Definition quality is above-average but falls short of production-grade due to missing output documentation and error recovery guidance.
Analyze bot detection and anti-scraping measures on a website. Identifies fingerprinting techniques, rate limiting, CAPTCHA systems, and other protections.
Analyze and deobfuscate JavaScript code on a website. Extracts function signatures, API calls, and security tokens from JavaScript bundles.
Parse and analyze XML sitemaps. Extracts all URLs, priorities, change frequencies, and generates crawling strategies.
Intercept and analyze WebSocket connections on a page. Captures messages, connection details, and message patterns.
Perform comprehensive deep analysis of a website including network traffic, JavaScript analysis, and security detection. Returns detailed insights about APIs, protection mechanisms, and fingerprinting techniques.
Delete a recording by ID. This action cannot be undone.
Output schemas not documented. Tools describe inputs but provide no guidance on response structure. LLMs cannot plan downstream calls or extract fields they need without trial-and-error.
No error handling or recovery guidance. Tools do not document failure modes, required permissions, or what the LLM should do on errors (retry, ask user, substitute tool).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 67 | <=2025-11-25 | v2 |
Identify all technologies used on a website including frameworks, libraries, CDNs, analytics tools, and payment processors.
Discover all REST and GraphQL APIs on a website, including hidden and undocumented endpoints. Useful for API reverse engineering and documentation.
Extract and structure content from a website. Returns clean text, HTML, markdown, and extracted structured data like headings, links, tables, and metadata.
Use LLM-powered extraction to intelligently extract structured data from websites. Supports OpenAI, OpenRouter, Groq, and Ollama.
Generate a crawler definition from a recorded session. Returns a YAML or Python crawler specification that can be executed to automate the same interactions.
Get current statistics of the browser pool including size, alive instances, and proxy pool status.
Retrieve a previously recorded session by ID. Returns all events, network traffic, and state snapshots.
Extract complete GraphQL schema from an endpoint using introspection. Returns schema, queries, mutations, and subscriptions.
List all available recordings with their metadata (duration, creation time, etc.).
Record user interactions with a website. Captures all events, network traffic, DOM changes, and state snapshots. Returns a recording ID that can be used to replay or analyze the session.
No permission scopes declared. Tools perform sensitive operations (JavaScript deobfuscation, API discovery, bot detection evasion) without declaring required permissions or warning about ethical/legal risks.
Destructive operation (delete_recording) lacks confirmation/dry-run pattern. Agent could permanently delete recordings without user approval.
Parameter descriptions inconsistent or minimal. 'endpoint' in introspect_graphql, 'url' parameters across tools lack format guidance (HTTP vs HTTPS, required structure). 'extraction_schema' in extract_smart provides no guidance on what structure is expected.
No result pagination or limiting documented. Tools like list_recordings could return thousands of recordings. No max result count, offset/cursor, or guidance on handling large datasets.
extract_smart requires 'extraction_schema' as arbitrary JSON object with no validation or schema guidance. LLM must infer the correct structure, risking malformed payloads.
record_session defaults to 300-second duration with 3600 max, but no guidance on what happens if recording is interrupted, times out, or fills memory. No timeout error handling documented.