AI-Driven Security Assessment - Model Context Protocol Server for pentest assessment management, findings/observations tracking, reconnaissance data, and command execution in containerized pentest environments
AIDA provides 6 pentest-focused tools with reasonable structure but suffers from inconsistent documentation quality, missing output schemas, and unclear error handling patterns. Tool names follow verb_noun convention (load_, create_, list_, add_, update_), which is good for LLM parsing. However, parameter descriptions lack specificity around constraints and acceptable ranges. Most critically, OUTPUT schemas are completely undocumented, LLMs cannot predict what fields will be returned or plan downstream tool composition. The add_card tool demonstrates good enum use (card_type, severity), but update_card lacks documentation about what fields are actually modifiable and what the return structure looks like. No tool includes error recovery guidance.
Add a card (finding, observation, or info) - returns the created card ID
Create a new pentest assessment and auto-load it. Gather scope, targets, and constraints from the user BEFORE calling this tool.
List existing assessments. Useful to check what exists before creating a new one or to find an assessment to load.
List all cards with optional type filter
Load an existing assessment to begin work. Returns full state: scope, phases, cards (findings/observations/info), recon data, credentials, and workspace structure.
Update an existing card by ID
Output schemas completely undocumented. No tool describes what fields, types, or structure it returns. LLMs cannot plan multi-tool workflows without seeing what data flows from one tool to the next.
Parameter descriptions lack actionable constraints. Example: 'name' in load_assessment has no format/length guidance; 'limit' in list_assessments has no min/max bounds. CVSS vector in add_card provides example format but no validation rules or error handling if invalid.
No error recovery guidance. Tools provide no documentation of failure modes, error classification (retryable vs user-fixable vs fatal), or what the LLM should do if a call fails. Example: What happens if create_assessment is called with a duplicate name? What should the agent do next?
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Tool composition unclear. load_assessment returns 'full state' but update_card only accepts card_id, title, target_service, cvss_vector. Unclear what fields are actually updatable, what immutable fields exist, and whether update_card returns the updated card or just a success confirmation. This breaks tool chaining.
Missing parameter descriptions for critical fields. 'target_domains' and 'ip_scopes' in create_assessment are documented as arrays of strings but lack format guidance (CIDR notation? wildcards?). 'credentials' and 'access_info' are free-form strings with no validation rules, invite LLM to pass unsafe or malformed data.
Enum value 'non_specifie' in create_assessment environment parameter is a typo ('non_specified'?). This forces LLMs to hallucinate or use incorrect casing, causing API rejections.
add_card description says 'returns the created card ID' but no output schema shown. Unclear if return is {id: int}, {card_id: int}, or {success: true, id: int}. LLM cannot reliably extract the ID for downstream update_card calls.
list_assessments and list_cards lack pagination guidance. No mention of what happens if result count exceeds the 'limit' parameter, whether a next_cursor or total_count is returned, or how LLM should handle incomplete result sets.