MCP server for web search, web scraping, URL discovery, and news ingestion with multi-engine search aggregation, content scraping, and RAG indexing capabilities
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
SearXNG MCP has 5 tools with explicit registrations in cli.py. Tool names follow verb_noun pattern (search_*, scrape_*, discover_*, index_*) which is positive. However, descriptions are vague and lack actionable details. Input schemas are present but parameter descriptions are minimal or missing critical constraints. No output schemas are documented. Error handling guidance is absent. The server exposes hardcoded paths (RAG_CLI_COLLECTIONS_ROOT) and subprocess calls without sanitization, creating security risks. Tool composition is weak, discover_urls and index_scrapes require external state (file system, subprocess tools) that agents cannot easily coordinate.
Tools (5)
discover_urlswritesource verified40/100
Discover a domain's URL set (robots/sitemap/navtree feeders); writes a pipe_scraper --url-file input.
index_scrapeswritesource verified43/100
Index previously-scraped URLs into a RAG collection (reads each URL's own sidecar).
scrape_url_chromiumread onlysource verified50/100
Scrape URL to filtered markdown (PruningContentFilter, full content, no length cap).
discover_urls accepts a file path parameter (url_file) and writes directly to disk. No validation, sanitization, or error recovery. LLM could be tricked into path traversal or overwriting critical files.
Document output schemas for all tools. Specify: return type (object/array), required fields, field types, and example structure. E.g., search_web should return {results: [{engine: string, url: string, title: string, snippet: string}], total_count: number}.
Expand parameter descriptions to 50 - 150 chars with: allowed values/format, constraints (min/max length, regex), examples of valid input, and what happens on invalid input. E.g., query: 'Search query in English, 2 - 5 keywords, no special operators. Will return results from Google, DuckDuckGo, Mojeek, StartPage, Brave, Bing, Yandex, and OpenALex.'
Convert engine list constraint in search_engine_drilldown to a JSON Schema enum: ['google', 'duckduckgo', 'mojeek', 'startpage', 'brave', 'bing', 'yandex', 'openalex']. Add note: 'Must match engine from a prior search_web response.'
Add explicit error recovery guidance. E.g., 'If no results found for query, try: (1) remove stopwords (and, the, a), (2) broaden keywords, (3) call search_web with different engines parameter.' See recovery-guide pattern.
Replace hardcoded RAG_CLI_COLLECTIONS_ROOT with an environment variable or config file. Document required setup: 'Requires WEBSEARCH_RAG_COLLECTION_ROOT env var pointing to rag-cli data/documents directory.'
Validate file paths in discover_urls to prevent path traversal: reject paths containing '..', absolute paths not under a safe base, and symlinks. Return clear error: 'Invalid url_file: must be relative path under current working directory.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 8 points across a rubric change (v1 → v2)
44/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
44
<=2025-11-25
v2
2026-03-09
F
36
-
v1
index_scrapes hardcodes RAG_CLI_COLLECTIONS_ROOT to a user's home directory (/Users/brunowinter2000/...), leaking local environment. This path is not configurable and fails on any other system.
index_scrapes runs subprocess.run(['rag-cli', ...]) without input validation, timeout, or error details. Subprocess failures return stderr/stdout as strings without structure. LLM cannot parse or recover from errors.
search_engine_drilldown requires query parameter to 'match a prior search_web call', but no mechanism enforces or tracks this. LLM has no way to know which queries are valid without trial-and-error.
discover_urls description mentions 'pipe_scraper --url-file input' without explaining what pipe_scraper is, where it lives, or how the agent should invoke it. No composition path documented.
Tool descriptions do not indicate which operations are idempotent vs destructive. search_web may be safe to retry; index_scrapes writes to disk and could duplicate entries. No explicit guidance for agent retry logic.
No input validation or enum constraints. search_engine_drilldown accepts any string for 'engine' parameter; the constraint list in description (google, duckduckgo, ...) is not enforced as an enum, so LLM can pass invalid values.
scrape_url_chromium description promises 'full content, no length cap' but real-world markdown can easily exceed context windows. No guidance on what happens when output is huge, or how to request truncation/summary.
scrape_url_chromium
Replace subprocess.run() with timeout, input sanitization, and structured error parsing. Catch non-zero returns and convert to {ok: false, error: string, stderr: string} JSON. E.g., if rag-cli times out, return 'rag-cli index timeout after 30s, ensure RAG service is running.'
Add idempotent/destructive hints to tool descriptions: 'search_web is safe to call multiple times with same query. index_scrapes may create duplicate collection entries if called twice with same URLs, consider the collection name as a version.'
Document tool chains. Add to discover_urls: 'Typical workflow: (1) discover_urls(seed_url=..., url_file=/tmp/urls.txt), (2) invoke pipe_scraper --url-file /tmp/urls.txt to scrape all URLs, (3) index_scrapes(collection=..., urls=[...]) to index results.' Link pipe_scraper docs if external.
Add per-item failure handling to index_scrapes. Return {ok: true, outcomes: [{url, status: 'indexed'|'no_sidecar'|'failed', detail: string}]} so agent knows which URLs succeeded and which can be retried.
Scrape_url_chromium: add optional max_chars parameter (e.g., 50000) to cap output size. Description: 'Scrape URL to filtered markdown. Returns up to max_chars characters; set to 0 for unlimited (may exceed context window). Default 50000.'
Add permission/scope documentation. E.g., scrape_url_chromium: 'Requires network access to target URL. May trigger rate limits or bot detection on some sites.' discover_urls/index_scrapes: 'Requires filesystem write access and rag-cli binary in PATH.'