MCP Server for OSS Search Infrastructure - Integrates SearXNG meta-search engine and Crawl4AI web crawling service. Provides tools for web searching, content extraction, and webpage crawling with caching support.
The server provides three tools (web_search, web_crawl, extract_content) with complete JSON Schema definitions and parameter descriptions. However, there are significant gaps in output schema documentation, error handling guidance, and description depth. Tool names follow verb-noun conventions but descriptions lack the specificity and dependency hints needed for optimal LLM selection. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite all tools being read-only operations. Parameter constraints are documented but scattered across description text rather than formal schema constraints. The server reads as functional but not production-grade for agent interaction.
Extract specific content from a webpage using CSS selectors or AI extraction. Uses cached crawl results when available, otherwise performs a fresh crawl.
Deep crawl and extract content from a webpage using Crawl4AI. Supports JavaScript rendering, media extraction, and intelligent content parsing.
Search the web using SearXNG meta-search engine. Aggregates results from multiple search engines including Google, DuckDuckGo, Brave, and more.
No output schemas documented for any tool. LLMs cannot predict the response structure or plan downstream tool calls without knowing what fields to expect.
Missing tool annotations. All three tools are read-only (risk: READ_ONLY) but lack the readOnlyHint annotation. LLMs cannot distinguish safe-to-retry tools from destructive ones without this metadata.
Description clarity varies. web_search description (113 chars) is adequate but generic; extract_content description (96 chars) lacks context on when to call it versus web_crawl. No dependency hints like 'Call web_search first to find URLs, then web_crawl to extract full content.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
No error handling guidance. Tool descriptions do not explain what errors users should expect (e.g., timeout, invalid URL, rate limit) or how to recover. An LLM hitting a crawl timeout has no recovery path.
Parameter enums described as free-text strings. 'categories' accepts 'general|images|videos|news|it|science' but is typed as string without enum constraint. 'extraction_strategy' and 'chunking_strategy' similarly lack formal enums. LLMs may hallucinate invalid values like 'video' or 'tree'.
Result limit enforcement unclear. web_search has 'max_results' (default 10, max 20) but no guidance on pagination. web_crawl and extract_content have no documented result limits. If a crawl returns 10,000 DOM elements, the LLM context could be exhausted.
extract_content parameter relationships undocumented. When 'selector' is omitted, does the tool default to 'AI extraction'? What if both 'content_type' and 'selector' are provided, which takes precedence? Ambiguity forces LLM guessing.