Professional Model Context Protocol server with 31 web scraping, crawling, deep-research, and autonomous-extraction tools. Returns clean Markdown and structured JSON for Claude, Cursor, and any MCP client. Defaults to local Ollama for LLM extraction (no API key needed); OpenAI/Anthropic available as opt-in. Includes a unified multi-format scrape tool, an autonomous agent, pre-built site templates, and Camoufox stealth browsing.
CrawlForge presents 31 tools with inconsistent definition quality. While descriptions are present for all tools, they are frequently under-described (most are 20-80 chars), lack actionable guidance for LLM selection, and many tools conflate multiple concerns into single operations. Input schemas are minimal, most tools accept only 1-3 parameters with basic type declarations but no constraints (enums, ranges, patterns). Output schemas are entirely undocumented: the source code shows no response shape definitions anywhere. No tool provides error guidance, parameter relationships, or recovery hints. The 'scrape' tool (tool 6) and 'agent' tool (tool 31) exemplify scope creep, 'scrape' accepts a 'formats' array and returns 'markdown, html, rawHtml, text, links, metadata, branding, screenshot, json, highlights, question', a kitchen-sink design that forces the LLM to reason about format dependencies rather than the server providing structured variants. The 'browser_session' (tool 23) tool accepts an 'action' parameter as free-form string ('create, close, execute') with no enum constraint, inviting hallucination. Naming is generally action-verb compliant (fetch_, extract_, scrape_, search_, list_, etc.), but several tools lack clear differentiation: 'extract_text' vs 'extract_content' vs 'extract_structured' are confusingly similar and lack descriptions explaining when to use each. The codebase shows tool definitions in scripts/tool-sweep.mjs and server implementations in src/, but actual schema registration is not visible in the provided source, tool-sweep.mjs is referenced but not included, making it impossible to verify whether schemas actually exist in the code or are inferred from the summary. This caps per-tool scores at 50 for tools where registration cannot be directly verified.
Autonomous agent for multi-step web research and data extraction
Analyze content
Process multiple URLs
Manage persistent browser sessions
Deep crawl websites
Multi-stage research tool with autonomous browsing
Extract main content from webpage
Extract embedded JS state (Next.js __NEXT_DATA__, etc.)
Output schemas completely undocumented. No tool documents what fields or structure are returned, forcing LLMs to guess field names for downstream calls or data extraction.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2025-06-18+ | v2 |
Extract all links from webpage
Extract SEO metadata from webpage
Extract structured data using an LLM
Extract clean text from webpage
Extract data using an LLM
Fetch raw HTTP content from a URL
Generate LLMs.txt for a website
Get results from a batch scrape job
List available Ollama models
Multi-language support
Map website structure
Process documents (PDF, web, etc.)
Read a large tool result by handle (internal tool for pagination)
Search Reddit for posts and comments
Unified multi-format scrape tool that returns content in markdown, html, text, links, metadata, json, highlights, and answer formats
Extract data using CSS selectors
Extract repository details using a site template
Browser automation with Playwright
Search the web (requires API key)
Check where a domain ranks in Google organic results
Anti-detection scraping with Camoufox stealth browser
Generate summaries of content
Content change tracking baseline creation and comparison
Descriptions are too short (most 20-80 chars) and lack LLM-actionable guidance. They state WHAT the tool does but not WHEN to use it, prerequisites, or why an LLM should prefer it over similar tools. Example: 'Extract clean text from webpage' gives no hint about output format, which parser is used, or when to call this instead of extract_content.
Multiple tools perform nearly identical operations with confusing names (extract_text, extract_content, analyze_content; extract_structured, extract_with_llm, scrape all extract data differently). Descriptions do not explain differences, forcing LLM to guess which to call.
No input validation or enum constraints. The browser_session 'action' param accepts freeform strings ('create', 'close', 'execute') with no enum, invites hallucination. The scrape 'formats' array has no validation for format values (markdown, html, rawHtml, text, links, metadata, branding, screenshot, json, highlights, question, 11 undocumented options).
No error handling guidance. Tools do not document what errors are possible, whether they are retryable, or what the LLM should do next. Example: 'URL not found' returns no hint to search for an alternative or check domain validity.
Scope creep in composite tools. The 'scrape' tool (tool 6) accepts 'formats' array and returns multiple unrelated representations (markdown, html, text, links, metadata, screenshot, json). This forces LLM to reason about format combinations instead of using single-purpose tools. The 'agent' tool (tool 31) absorbs entire research workflows.
Tool registration details not visible in provided source. scripts/tool-sweep.mjs and tool-selection-eval.mjs are referenced but not included. Actual schema registration, field types, and structured output definitions cannot be verified.
No parameter relationships documented. When a tool accepts multiple parameters (e.g., batch_scrape with 'urls' array, scrape_with_actions with 'actions' array), no description explains what valid actions are, what URL formats are accepted, or what happens if arrays are empty.
No pagination or result limiting documented. Tools like 'extract_links', 'search_web', 'reddit_search', 'crawl_deep' can return unbounded results but no tool description mentions limits, pagination, or offset/limit parameters. This violates the 20-50 result baseline and risks context window exhaustion.
No API key injection or secret handling documented. Tools 'search_web' (tool 8) and 'reddit_search' (tool 10) state '(requires API key)' in description but do not explain how keys are configured (env vars, vault, request params). If passed as parameters, this leaks secrets into logs.