Security validator for AI agents — 191 attack probes to test prompt injection and extraction defenses
Scoring was not performed
Output schemas are completely undocumented. Tools return arrays, objects, or verdicts with no documented field types or structures. LLMs cannot plan downstream tool chains or extract specific fields from responses.
No enum constraints on mutation/detection parameters. Tools accept free-form strings (e.g., 'verdict' in fuseVerdicts, 'text' in base64Wrap) with no documented valid values. LLMs will hallucinate invalid inputs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 28 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Parameters with complex types (function, object, array) lack detailed documentation. E.g., 'embed' parameter in computeSemanticSimilarity is typed as 'function' but no signature, return type, or error behavior documented. 'client' in fromOpenAI is 'object' with no schema.
No error classification or recovery guidance. Tools return raw errors with no categorization (retryable, user-fixable, fatal). No guidance on what to do if detection fails or embedding fails.
Minimal input validation documented. Tools like generateCanary, buildExtractionProbes, buildInjectionProbes accept empty input schemas, leaving no room for optional parameters or configuration.
AgentValidator tool accepts a 'function' parameter (agentFn) with no signature documentation. LLMs cannot determine how to construct or pass this without examples.