First-run setup wizard and MCP server for dark web OSINT analysis via Tor. Provides tools for checking Tor connectivity, searching dark web engines, fetching .onion content, analyzing with LLM, and rotating circuits.
OnionClaw presents severe definition quality issues across nearly all dimensions. While tool names are action-oriented (check_*, search, fetch, ask, renew_*, etc.), most descriptions are inadequate for LLM understanding, parameter documentation is minimal or missing, and schemas are either incomplete or invisible in the provided source. The server exposes sensitive Tor/LLM operations without proper permission gates or error recovery guidance. Of 17 tools, most lack actionable descriptions and comprehensive parameter schemas. The presence of tool descriptions varies from reasonable (e.g., 'search') to dangerously vague (e.g., 'dispatch' with a generic 'name' and 'input' dict). No output schemas are documented. Error handling and security considerations are absent from the specification.
Extract entities, keywords, and structured data from content without using LLM (uses BM25 + regex).
Analyse dark web content with an LLM and get a structured OSINT report.
Ping all 12 dark web search engines via Tor and report status + latency.
Verify Tor is running and return the exit IP address.
Check GitHub for the latest OnionClaw release.
Delete all cached fetch results.
Dispatch a tool call by name with input dict. Returns result or error.
dispatch tool is a meta-tool that bypasses schema validation and safe composition. It accepts a generic 'name' string and untyped 'input' dict, making it impossible for the LLM or server to enforce constraints, validate parameters, or guide error recovery. This pattern violates the single-responsibility and schema-driven design principles.
No output schemas documented for any tool. LLMs cannot plan downstream tool chains or extract required fields (e.g., watch_register likely returns a job_id needed by watch_disable, but this is not specified). Output schema documentation is required for agentic tool composition.
No error handling or recovery guidance. Tools like renew_identity (Tor circuit rotation) and watch_register (persistent jobs) lack descriptions of failure modes, prerequisites, and actionable recovery steps. An LLM calling these without understanding failure scenarios risks incomplete operations.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 61 | <=2025-11-25 | v2 |
Fetch any URL or .onion hidden service through Tor with caching support.
LLM picks top-20 most relevant search results from a larger set.
LLM trims query to ≤5 focused keywords for optimal dark web search.
Rotate the Tor circuit and get a new exit node / identity.
Batch fetch multiple URLs via Tor and return text content dictionary.
Search dark web content via Tor using 12 active search engines simultaneously.
Run all due watch/alert jobs and return new results since last check.
Disable a watch job by its ID.
List all active watch jobs with their status and next check times.
Register a query as a persistent watch/alert job for periodic re-checking.
Sensitive operations lack permission gates or dry-run support. renew_identity and clear_cache are destructive/stateful operations (rotating Tor circuits, deleting cache) but descriptions do not indicate whether they require confirmation, permission checks, or support idempotent retries. Agents could rotate circuits or clear data unintentionally.
Parameter descriptions are often vague or incomplete. 'custom_instructions' in ask(), 'mode' enums across multiple tools, and parameter constraints (query length limits, interval ranges) are not fully documented. LLMs cannot infer that a query should be '≤5 keywords' from the description alone.
Tool dependencies are undocumented. watch_disable requires a job_id, but watch_list is not explicitly linked as the discovery tool. Similarly, fetch depends on successful Tor setup (check_tor prerequisite) but this is not stated. LLMs must guess the call order.
No documented return types or example outputs. Tools like search, scrape_all, and filter_results return complex structures (arrays of results with metadata), but the response schema is not defined. LLMs and agents cannot reliably extract data or plan follow-up actions.
LLM credentials and Tor configuration are not abstracted from tool parameters. setup.py shows LLM_PROVIDER and API keys in .env, but tools like ask() do not document how credentials are injected. If credentials were ever exposed as parameters, they would leak into logs and agent traces.
Pagination and result limits are not implemented. search() accepts max_results but no offset/cursor; scrape_all() accepts an array of URLs but no limit. Large result sets will bloat the context window and degrade LLM reasoning. Production tools enforce result caps and pagination.
Tool naming inconsistency: watch_* tools cluster under a naming pattern but lack parallel list/register/disable clarity. 'watch_check' is ambiguous, does it check a single watch or all active watches? Does it return incremental results or a full dataset? Similar ambiguity in 'analyze_nollm', LLM vs non-LLM is a distinction but unclear when each should be used vs ask().