Self-healing, vendor-aware web-scraping framework for hostile sites, driveable by any LLM over MCP. Exposes Anansi's scraping capabilities as MCP tools so any LLM can fetch pages, extract structured data, run full crawls, and manage pause/resume.
Anansi MCP Server demonstrates solid definition quality with 11 well-named tools that follow verb_noun conventions. All tools have clear descriptions and complete input schemas with type definitions and parameter descriptions. The server implements comprehensive input validation (URL safety, proxy validation, action whitelisting) and appropriate risk categorization (READ_ONLY vs WRITE). However, output schemas are not explicitly documented in the source code, and some tools lack actionable error guidance. The codebase shows strong security practices (sandbox confinement for exports, action type allowlisting, permission checks via environment variables) but descriptions could be more LLM-optimized for task selection.
Fetch multiple URLs concurrently with configurable concurrency limit
Start a new crawl with a spider configuration, returning a crawl_id for tracking
Get extracted items from a completed crawl
Get the status of a running or completed crawl
Export a completed crawl's items to a file (JSON, CSV, or JSONL)
Extract structured data from a URL using CSS/XPath selectors and optional Playwright actions
Fetch a single URL with optional browser rendering, proxy support, and output format selection
List available bot profiles (crawler identities)
Output schemas not documented in source code. Tool descriptions explain what inputs are accepted but do not specify what data structure is returned or which fields LLMs should expect in responses. This forces LLMs to infer the response structure, increasing errors in downstream tool chains.
Error handling is defensive (validation before execution) but error messages returned to the LLM are not shown in the source. While the code validates URL safety, proxy URLs, action types, and impersonation targets, the structured error responses are not documented. LLMs need 'what to do next' guidance in error messages (pattern:recovery-guide).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 66 | <=2025-11-25 | v2 |
List all crawls with their current state and item counts
Pause a running crawl
Resume a paused crawl
crawl, pause_crawl, resume_crawl, and export_crawl are marked as WRITE tools but descriptions do not explicitly state that these operations have side effects or consequences. For an agent, knowing which calls are irreversible is critical for planning and retry logic. Descriptions should say 'Creates a new crawl...' or 'Marks items as exported...' to clarify state changes.
Parameters like 'fields' in extract tool lack actionable format guidance. The description says 'Field definitions with CSS or XPath selectors' but does not explain the exact structure (is it {field_name: selector} or {field_name: {type: 'css', value: selector}}?). LLMs need explicit examples of the object shape to construct valid input.
No pagination documentation for list_crawls and crawl_items. crawl_items accepts limit/offset but list_crawls does not. If a user has thousands of crawls, list_crawls will either return all (token waste) or silently truncate. The tool description should state 'Returns all crawls' or 'Paginated via limit/offset' so the LLM knows what to expect.
extract tool actions parameter is complex and underdocumented. The description mentions 'Playwright actions' and the code enforces _ALLOWED_ACTION_TYPES and _ALLOWED_PRESS_KEYS, but the tool description does not provide the action object schema or examples. An LLM calling extract with actions=[{type: 'invalid_action'}] will fail with a validation error that is not surfaced as actionable guidance.