Open-source MCP server for AI agents: web search, content extraction, and library docs.
wet-mcp has 12 tools with generally present but inconsistent schema definitions and descriptions. All tools are registered with names and input schemas are visible in src/wet_mcp/server.py. However, tool descriptions are terse (most 10 - 80 chars), parameter descriptions are minimal or absent, and output schemas are not documented. Error handling guidance is not evident from the visible code. Tool naming follows verb_noun patterns (search_web, crawl, extract) which is good, but descriptions lack LLM-optimized guidance about WHEN to use each tool vs. alternatives, and parameter constraints are sparse. The server implements fastmcp (a Python framework) but the actual tool handler implementations are not fully visible in the provided code sample, so composition quality and error recovery patterns cannot be fully assessed. STDIO-only transport caps protocol readiness at 50, pulling the overall evaluation down.
Get server configuration and status.
Crawl and extract content from a URL using Playwright and Crawl4AI.
Discover and retrieve documentation for programming libraries across multiple ecosystems.
Download and cache files from URLs.
Extract structured content from a URL (text, images, links, metadata).
Index a library's documentation for fast semantic search.
List all media (images, videos) from a webpage.
Run Google Drive sync setup using OAuth Device Code flow.
Tool descriptions are uniformly terse (10 - 80 characters). E.g. 'Crawl and extract content from a URL using Playwright and Crawl4AI' tells the LLM WHAT but not WHY or WHEN.
Parameter descriptions are missing or extremely sparse. E.g., 'crawl' tool has action='string, enum=[extract, markdown, html]' but no description of what each action does or when to choose one. LLMs cannot infer semantics from enum values alone.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 52 | <=2025-11-25 | v2 |
Pre-download models and run setup to avoid first-run delays.
Search and retrieve relevant documentation for libraries from an indexed database.
Search the web using SearXNG or external backends (Tavily, Brave, Exa).
Fetch and parse sitemap.xml from a domain.
Output schemas are not documented. Visible definitions show input shapes but do not describe what fields are returned, their types, or how to chain results into downstream tools. E.g., search_web returns results, but result structure (count, pagination, field names) is not specified.
Error handling and recovery guidance are not visible in tool definitions. No descriptions mention how failures are signaled, when to retry, or what the user/agent should do next. E.g., 'sitemap' may fail if domain has no sitemap.xml, but no recovery hint is provided.
Tool names 'run_warmup' and 'run_setup_sync' are vague. 'run_warmup' could mean many things (cache warming, model download, system prep). 'run_setup_sync' conflates setup with sync. Better names: 'download_models' and 'setup_drive_sync' or 'authorize_drive'.
Parameter 'region' in search_web is described as '2-letter ISO 3166-1 alpha-2 geo code (optional)' but lacks examples or a constraint pattern. LLMs may guess wrong (e.g. passing full country name instead of code).
Tool 'config' uses a generic 'action' parameter with enum [status, list_backends, warmup_status]. This violates the single-responsibility principle, three separate tools (get_config_status, list_backends, get_warmup_status) would be clearer and individually composable.
Tools 'crawl' and 'extract' both retrieve content from URLs. The distinction is unclear from descriptions alone. crawl mentions Playwright+Crawl4AI and action parameter; extract mentions 'structured content'. LLMs may conflate them without clear guidance on when to use which.