Open-source MCP server for AI agents: web search, content extraction, and library docs.
wet-mcp provides 11 tools with mixed quality. Naming is generally verb-first and clear (search, crawl, extract, sitemap, etc.). All tools have descriptions, though many are brief (10-50 chars). Input schemas are present and typed for all tools, though output schemas are not documented. Error handling is minimal, no recovery guidance, categorization, or actionable error messages visible in the source. Tool composition is reasonable (each performs a single logical action), but parameter descriptions lack constraint details (ranges, valid values, dependencies). The server follows basic definition patterns but falls short of production-grade quality in parameter validation, error guidance, and output schema documentation.
Get or set server configuration (rate limits, embedding models, providers, etc.).
Fetch and parse content from a URL using Crawl4AI browser automation.
Discover library documentation URL by name and language.
Extract structured data from HTML/text content using LLM-powered parsing.
Index a library's documentation into the vector database for semantic search.
List all media files (images, videos, PDFs) from a crawled page.
Search the web using multiple backends (SearXNG local, Tavily, Brave, Exa, Google).
Output schemas not documented for any tool. LLMs cannot infer what fields to expect or plan downstream chaining (e.g., does search return URLs? Does crawl return markdown or HTML?).
Parameter descriptions lack constraint details. 'limit' and 'offset' parameters appear without explicit ranges (min/max). 'backend' enum values are listed informally ('local', 'tavily', 'brave'...) but should be declared as formal enum constraints. 'language' parameter in discover_library and search_docs lists allowed values in description but not as JSON Schema enum.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Search library documentation with vector/keyword hybrid search and optional reranking.
Set up Google Drive or S3 sync using OAuth Device Code flow.
Fetch and parse XML sitemap from a domain.
Pre-download embedding and reranking models to avoid first-run delays.
'warmup' tool has minimal description (likely <20 chars based on definition) and empty input schema with no explanation of purpose, prerequisites, or expected behavior. This violates the 10-1024 character and 'every parameter needs a description' rules.
'config' tool conflates two responsibilities: get status AND set configuration via an 'action' enum. Should split into 'get_config' (read-only, safe to call) and 'set_config' (write, requires permission/confirmation). Current design forces LLM to reason about action branching.
No error handling guidance visible. Tools do not document what errors are possible, how they are classified (retryable vs. fatal), or what the agent should do next. E.g., what happens if a URL is unreachable in 'crawl'? If API key is invalid in 'search'? If vector DB is unavailable in 'search_docs'?
'setup_sync' accepts 'client_id' and 'client_secret' as tool parameters with a note that they 'override' environment variables. This is a security anti-pattern, credentials must never be tool parameters. Secrets in parameters leak into agent logs and trace history. Use server-side secret injection only.
Parameter descriptions for 'search_docs' mention 'rerank' as optional but do not explain what reranking does, when to enable/disable it, or the performance/latency tradeoff. 'extract' accepts a 'schema' object but does not document expected JSON Schema format or validation rules.
'ingest_docs' and 'warmup' are write operations but descriptions do not explicitly state that they modify state. Agents need to know which calls have irreversible consequences or side effects.
'list_media' accepts 'types' as an array but does not specify valid enum values (image, video, pdf, audio are mentioned informally). LLMs may pass invalid MIME types or capitalization variants, causing failures.
'crawl' accepts 'exclude_selectors' and 'wait_for_selector' but does not explain CSS selector syntax, escaping rules, or common pitfalls. LLMs may pass invalid selectors without guidance on what went wrong.