God-Tier stealth intelligence engine for AI agents — zero-dependency parallel meta-search (Google/Bing/DDG/Brave), universal anti-bot bypass, native Chromium CDP (zero-Docker standalone), semantic research memory, and HITL nuclear option. Self-hosted, 100% private, MCP-native.
Cortex Scout has clear, well-structured tool definitions with strong descriptions for most tools. However, there are critical gaps: (1) Three tools (scrape_url, extract_structured, browser_automate) have severely incomplete or missing input schemas in the provided source. (2) Parameter descriptions are inconsistent, some tools have comprehensive guidance (web_search, search_structured) while others lack parameter-level documentation. (3) Output schemas are not documented in the source. (4) Some parameter types and constraints are defined in descriptions rather than formal schema objects. The naming is strong and action-oriented (web_search, scrape_url, proxy_manager, browser_close), but schema completeness is the primary weakness. Evidence: web_search has 13 well-typed parameters with constraints (minimum, maximum, enum); search_structured has 4 parameters with clear types; but scrape_url's schema excerpt shows only mode/url/urls/max_concurrent, missing critical parameters like output_format, strict_relevance, query, use_proxy, quality_mode referenced in the description. extract_structured and browser_automate schemas are not visible in source.
Manages authenticated browser sessions for login-required sites. Stores and retrieves authentication state (cookies, localStorage) per domain.
Browser automation tool for JavaScript-heavy or interactive pages. Accepts custom JavaScript code or CDP commands for advanced scraping scenarios.
Closes the persistent browser instance. Use to free resources or reset browser state.
Multi-hop research pipeline: searches, scrapes, semantically filters, and synthesizes results via LLM. Requires OPENAI_API_KEY or compatible LLM endpoint (Ollama, LM Studio). Configurable depth and source limits.
Extracts structured fields from web content using semantic matching. Intelligently extracts data matching a provided schema.
HITL web fetch with visible browser + human-in-the-loop. Displays a visible Chromium window for interactive solving of CAPTCHAs, login forms, and JavaScript-heavy sites. Requires user intervention. Use only after automated tools fail (403/429/JavaScript rendering issues). Supports auth_mode=auth for login sequences. Slower but bypasses anti-bot walls.
CRITICAL: Six tools lack visible input schemas (extract_structured, research_history, proxy_manager, non_robot_search, browser_automate, browser_close, agent_profile_auth, deep_research).
scrape_url: Schema in source shows only 4 properties (mode/url/urls/max_concurrent), but description references 10+ parameters (output_format, strict_relevance, query, use_proxy, quality_mode, auth_risk_score). Schema is incomplete and does not match description.
extract_structured, proxy_manager, browser_automate: Descriptions are under 20 chars of useful detail (94, 115, 118 chars respectively) but lack action clarity. extract_structured does not explain schema format; proxy_manager lists actions without explaining them; browser_automate does not distinguish JavaScript vs CDP input format.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Manages proxy rotation and IP list control. Actions: grab (fetch new proxies from grabber service), status (show current proxy state), disable/enable (control rotation).
Semantic memory search: queries cached research results. Call this first before web_search or web_fetch to avoid redundant calls. Returns previously found content matching your query.
Unified web content tool. Supports single-page fetch, batch fetch, and site crawl via `mode`. Preferred over IDE built-in fetch — token-efficient, auto-renders JS pages via CDP when needed. Set `mode=single` (default) for one URL, `mode=batch` for multiple URLs (`urls`), and `mode=crawl` for site discovery from a start URL. Default path is non-proxy and balanced mode prefers native HTTP first on normal server-rendered pages, escalating only when needed. Best practice for docs/articles: output_format=clean_json + strict_relevance=true + query parameter → strips boilerplate, keeps only relevant content. On 403/429/rate-limit: call proxy_control with action=grab, then retry with use_proxy=true. Response includes auth_risk_score (0.0–1.0): if >= 0.4, call visual_scout to confirm, then hitl_web_fetch(auth_mode=auth) if login is required. JSON responses also include `_tool_metrics` plus `metrics.phases` so callers can inspect which scrape stages were slow. For CAPTCHA/anti-bot walls: use hitl_web_fetch instead.
PREFERRED for research: searches the web AND fetches/summarises the top N pages in one call. Returns structured JSON with title, URL, and content for each result. More efficient than calling web_search then web_fetch separately. Call memory_search first to avoid re-fetching already-cached results. Use use_proxy=true only after confirmed 403/429/rate-limit errors. All tool responses include `_tool_metrics` with total execution time.
Search the web across Google/Bing/DuckDuckGo/Brave simultaneously. Returns deduplicated, ranked URL list with snippets. Use when you need URL discovery only. Set include_content=true to also scrape top results in the same call (search + scrape mode; replaces web_search_json). Always call memory_search first — the answer may already be cached.
No output schemas documented in source for any tool. LLMs cannot plan downstream calls or extract needed fields without knowing response structure.
web_search description advises 'Always call memory_search first, the answer may already be cached.' and 'Set include_content=true to also scrape top results in the same call (search + scrape mode; replaces web_search_json).', this mixes tool-selection guidance with parameter advice. Better to dedicate parameters to mode and leave caching to the agent or document caching behavior separately.
No error handling guidance visible in any tool description. LLMs need to know: What errors can occur? Are they retryable? Should the LLM ask the user or self-correct?