MCP Server for Crawl4AI - provides web crawling and search capabilities
This MCP server has significant definition quality gaps. While both tools have basic descriptions and input schemas, the descriptions lack depth and clarity needed for LLM selection. The 'search' tool description is entirely in Chinese, making it inaccessible to English-speaking agents. Schemas are present but lack proper enum constraints for the 'format' and 'engine' parameters despite listing discrete valid options in descriptions. Parameter descriptions are generic and lack actionable guidance on constraints, formats, or when to use each parameter. Error handling is minimal, tools return JSON error objects without recovery guidance. No output schemas are documented. Tool naming is acceptable but descriptions fall well below the 10-1024 character baseline for optimal LLM comprehension.
Crawl a webpage and return its content in a specified format. Args: url: The URL to crawl format: The format of the content to return. Options: - raw_markdown: The basic HTML→Markdown conversion - markdown_with_citations: Markdown including inline citations that reference links at the end - references_markdown: The references/citations themselves (if citations=True) - fit_markdown: The filtered/"fit" markdown if a content filter was used - fit_html: The filtered HTML that generated fit_markdown - markdown: The default markdown format
执行网络搜索并返回结果。 Args: query: 搜索查询字符串 num_results: 返回结果的数量,默认为10 engine: 使用的搜索引擎,可选值: - "duckduckgo": 使用DuckDuckGo搜索(默认) - "google": 使用Google搜索(需要配置API密钥)
search tool description is entirely in Chinese, making it inaccessible to English-speaking LLM agents and violating the tool-description pattern requirement for clarity.
'format' and 'engine' parameters list valid discrete options in descriptions but fail to declare them as enums in the schema. Free-form strings invite hallucinated values like 'custom_format' or 'bing'. Must use JSON Schema enum constraints.
No output schemas documented. LLMs cannot plan downstream tool calls or extract structured data. read_url returns a string; search returns JSON, both lack explicit schema declarations of what fields/structure the agent should expect.
Error handling is minimal and non-actionable. Both tools return generic error JSON (e.g., '{"error": "..."}') without recovery guidance. No indication of whether errors are retryable, require user input, or are fatal.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter descriptions are generic and lack actionable constraints. The 'url' parameter has no format guidance (http/https?). The 'query' parameter has no length limits or special character guidance. Numeric 'num_results' lacks min/max bounds.
Tool descriptions lack 'WHEN to use' context. The read_url description does not explain when to prefer raw_markdown vs markdown_with_citations. The search description omits guidance on engine differences (DuckDuckGo vs Google) or when each is appropriate.