Multi-source search MCP server for AI agents. Search Reddit, X, YouTube, HN, GitHub, arXiv, Polymarket, 微博, 知乎, 豆瓣, 头条, 雪球, V2EX, B站, 小宇宙 & more — 30 sources across Chinese & English platforms with adjustable time window, source/category scoping, trending hot lists, and optional LLM synthesis.
reach-mcp demonstrates strong tool design with excellent parameter schemas and descriptions. The search tool is exceptionally well-documented with 1000+ char descriptions covering query styles, trending modes, and credential health. All 5 tools have clear naming (verb_noun pattern) and comprehensive parameter documentation. However, output schemas are not formally documented in the visible code, descriptions reference return fields (e.g., '{brief, items, sources_used, source_summary}') but there is no explicit JSON Schema definition for tool outputs. Error handling guidance is embedded in descriptions (e.g., 'NOTICE lines' for degraded sources) but not formalized as structured error categories. Tool composition is strong, each tool has a single responsibility and output from search flows naturally into fetch_content/synthesize. No security issues detected (no secrets in params, read-only operations). The main gap is formal output schema documentation and structured error recovery patterns.
Fetch the full content of ONE item found via search. Two-stage retrieval: search returns metadata + a snippet for every source; call this when an item is worth reading/hearing in full — especially after search(synthesize=false), which returns snippets only. Rich-media sources have dedicated backends — xiaoyuzhou (pass the item's audio_url → Whisper transcript of the episode), youtube (watch URL or video id → captions), bilibili (video URL → CC subtitles if the video has them, else ''); every other source falls back to Jina Reader on the item's url. Args: `source` (the item's source field), `id_or_url` (audio_url for xiaoyuzhou, url/id otherwise). Returns {source, url, content, ok}. With synthesize=true, search already backfills the top rich-media items automatically — use this for anything beyond those.
Inventory of all registered sources. Call before search when unsure what's active. Returns [{name, description, needs_auth, available, required_env, default_days, default_limit}]: available=false = gated (credential in required_env not set). No arguments.
Fetch any URL as clean markdown via Jina Reader. Use for the full text of a page found via search — a thread, article, or repo — when the item's `text` snippet isn't enough. Returns {url, content, ok}; content is '' on failure. Keyless.
Search up to 33 social & web sources in parallel, score by engagement, and synthesize a cited brief. YOU control scope. Best scoping: `category` — social: x, reddit, instagram, threads, tiktok, xiaohongshu, weibo, zhihu, douban, toutiao, linuxdo, bilibili, youtube, pinterest, bluesky, linkedin, web, quora; it: github, hackernews, v2ex, rss, arxiv, dripstack, stackoverflow, lobsters; tech: arxiv, techmeme, digg, dripstack, hackernews; polec (politics & economics): truthsocial, xueqiu, stocktwits, polymarket; podcast: xiaoyuzhou. Categories overlap (e.g. github is both it and tech) — multiple categories union. `sources` picks individual names; both together = union; both omitted = all available sources EXCEPT podcast (xiaoyuzhou is opt-in: episode transcription is slow, request it explicitly when you need podcasts). QUERY STYLE: write a short keyword query, not a full question. Literal keyword-AND sources (x, threads via Apify) match every word — 2-5 content words, no 'latest/news/how-to' filler. Keyword-slot sources (bluesky, tiktok, instagram, pinterest, linkedin, quora, xiaohongshu, weibo) take a compact phrase. Semantic sources (reddit, web, arxiv, github, hackernews, youtube, bilibili) tolerate longer natural phrasing; stackoverflow (official SE API, Q&A corpus) and lobsters (feed-filtered, so specific tech terms) take English tech keywords; douban (豆瓣 movie/TV/book/music ratings, keyless) and zhihu take Chinese titles/keywords; zhihu is hot-list browse (filtering, not search — Chinese keywords work best). The pipeline already strips question/meta words per source and retries X with shorter variants, so lead with the core subject. Match query language to platform — Chinese keywords work best for the CN sources. For WeChat 公众号 articles, use the web source with 公众号 in the query (auto-scoped to mp.weixin.qq.com) — there is no dedicated wechat source. TRENDING (热搜/热榜): set trending=true to fetch what's hot RIGHT NOW instead of searching — weibo 实时热搜 (with heat values), zhihu 热榜, toutiao 头条热榜 (with hot values), hackernews front page, lobsters hottest, linuxdo 每日热门, bilibili 综合热门 ranking, x/X trends (via the trends24 mirror — works WITHOUT the x login cookies), and github newly-hot repos (created this week, sorted by stars). `query` is IGNORED in this mode (pass ""); `sources` still scopes (e.g. sources=["weibo"]). Use it for 'what's trending on weibo', '今日热搜', or to seed a topic before a keyword search. quora and linkedin gain dated, answer-rich results when EXA_API_KEY is set (exa neural search scoped to those domains). CREDENTIAL HEALTH: if source_summary carries NOTICE lines, that source degraded (e.g. a stale login cookie fell back to a limited public path). Results are still usable, but mention the notice when it matters and suggest refreshing the named env var. The default (synthesize=true) returns a cited brief plus scored items — one call = a finished report. It auto-backfills full content for the top rich-media items (xiaoyuzhou/youtube/bilibili) before the brief. synthesize=false returns raw metadata + a per-item snippet instead, for custom post-processing; pair it with fetch_content to read any item in full. `max_chars_per_item` caps snippet length (raise for fuller CN posts, lower to save tokens). Returns {brief, items, sources_used, source_summary}. Each item: {source, title, url, author, date, score, engagement, text}. source_summary is one compact line per outcome — 'x:3; reddit:5 | EMPTY: rss, v2ex | QUOTA: tiktok(monthly limit) | ERRORS: digg(429)'; 'gated_off' means its credential env isn't set. Call list_sources if unsure what's configured.
Output schemas not formally documented as JSON Schema. Tool descriptions reference return fields (e.g., '{brief, items, sources_used, source_summary}') but the actual schema structure for complex objects is not visible in code, making it difficult for LLMs to validate response structure and plan downstream operations.
Enum constraints for category and sources parameters are documented in prose descriptions rather than as formal JSON Schema enums. LLMs cannot machine-parse the valid values from 'Best scoping: category, social: x, reddit, instagram...' and may hallucinate invalid category names.
Error handling and recovery guidance is embedded in descriptions (NOTICE lines for degraded sources, credential health) rather than formalized as structured error responses with retryable/user-fixable/fatal classifications. No explicit error taxonomy visible in tool definitions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 70 | <=2025-11-25 | v2 |
LLM-synthesize a cited brief from items returned by a prior search(synthesize=false), WITHOUT re-searching. Args: `query` (the original), `items` (the prior items list). Returns {brief}.
No pagination or result limiting guidance visible for list_sources or search results. search description mentions default limit=20 but does not document whether results can be paginated or what happens if a source returns more than limit items.