MCP server to integrate Katana web crawler with Claude
The server has 4 tools with complete schemas and descriptions, but quality is uneven. Schemas are fully defined with proper types, enums, and defaults, which is strong. However, descriptions are in Vietnamese and lack the LLM-optimized clarity expected for production. Parameter descriptions are present but generic, many lack guidance on when to use them or what values are valid beyond defaults. Error handling is absent from the code sample, no recovery guidance, no input validation logic visible, no categorization of errors. Tool names follow verb_noun pattern (katana_crawl, katana_check_version) which is good. Overall composition is reasonable but the first tool (katana_crawl) is overloaded with 24 parameters, many of which could be grouped or made optional. No output schemas documented in the code.
Kiểm tra phiên bản Katana đã cài đặt
Chạy Katana để crawl một URL hoặc danh sách URLs. Trả về các URLs, endpoints, và thông tin được phát hiện.
Chạy Katana crawl từ file chứa danh sách URLs
Chạy Katana để crawl và lấy toàn bộ các URL file JavaScript (.js), sau đó tải toàn bộ file JS đó về một thư mục trên máy.
Overloaded tool signature: katana_crawl has 24 parameters. LLM reasoning degrades significantly when tools have >8 parameters. Many options (headless, js_crawl, xhr, form_fill, form_extract, form_submit, keep_form_data, filter_status, filter_regex, exclude_regex) should be grouped into a config object or split into separate specialized tools.
Descriptions are in Vietnamese and lack LLM-optimized structure. None answer the three critical questions: What does it do? When should the LLM call it instead of a similar tool? What does it return? For example, katana_crawl description does not explain what 'endpoints' means, when to use headless mode vs js_crawl, or what the output structure is.
No output schema documented. Tools return data but the code sample does not show what structure the LLM receives. This violates the pattern requirement that 'LLMs need to know what fields to expect so they can plan downstream tool calls.' Without documented output schemas, LLMs cannot reliably chain tools or extract results.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Error handling is not visible in the code sample. No try-catch blocks, no validation logic, no recovery guidance visible. The katana_download_js tool writes to disk (WRITE risk) but there is no error classification, no dry-run option, and no confirmation step as recommended by pattern:confirmation-request.
Parameter dependency documentation is missing. For example, store_response and store_response_dir are mutually dependent, if store_response=false, store_response_dir is ignored. Similarly, url vs urls vs file_path in katana_download_js are mutually exclusive but this is not documented. LLMs will pass multiple conflicting values unless explicitly told not to.
Default values for boolean flags are confusing. headless defaults to True (correct) but description says 'default: false', contradiction. This will confuse both developers and LLMs. All defaults must be verified against actual behavior.
No input validation visible. Parameters like depth, concurrency, and timeout accept integers with no bounds checking shown. An LLM could pass depth=999999 or concurrency=-1 and crash the tool. Validation should be explicit with clear error messages like 'depth must be 1 - 100'.
katana_download_js accepts three mutually exclusive input methods (url, urls, file_path) but does not clarify priority if multiple are provided. Which takes precedence? Should the tool error if more than one is set? This ambiguity will cause unpredictable behavior.