An adaptive penetration testing system using LangGraph and MCP tools for reconnaissance and web analysis with human-in-the-loop decision making
Three tools with basic schemas and descriptions, but significant gaps in parameter documentation, output schema clarity, and error handling. Tool names are action-oriented (port_scanner, banner_grabber, web_search) but descriptions are terse and lack actionable context. All three tools are READ_ONLY reconnaissance/analysis operations. Parameters lack constraint descriptions (e.g., valid port ranges, expected query formats). Output schemas are inferred from return type hints but not formally documented in descriptions. Error handling is minimal, tools silently return empty results on failure rather than guiding recovery.
Simple TCP banner grabber for PoC use.
Scans the specified ports on the target IP and returns a list of open ports.
Performs a web search using the Google Custom Search API and returns combined query and search results.
Descriptions critically short and under-specified. port_scanner description is 8 words. banner_grabber admits PoC status without production guidance. web_search lacks output format documentation.
Output schemas not documented. Tools return List[int], Dict[int, str], and str without describing the structure in parameter or return documentation. LLMs cannot infer correct downstream tool chaining.
task_id parameter appears in all three tools but its purpose is never explained. Is it for tracking? Logging? Mandatory or optional? All three mark it as REQUIRED but do not justify why or how it is used.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | 2026-07-28+ | v2 |
No error handling guidance. Port scanner and banner_grabber silently return empty results on socket errors. LLM receives no signal of why the call failed (timeout, unreachable, permission denied, invalid IP) and cannot self-correct or retry intelligently.
banner_grabber uses 'banner_grabber' as name instead of verb_noun pattern (e.g., 'grab_banners' or 'get_banners'). This is stylistically inconsistent with port_scanner and web_search.
No parameter constraints. 'ports' parameter accepts any list of integers with no validation of port range (0-65535). web_search's num_results has no documented upper bound despite API likely limiting to 100.
No tool-chaining guidance. Docstrings do not explain the relationship between tools. Should an agent call port_scanner first, then banner_grabber on open ports? Or web_search on discovered services? No dependencies documented.