JavaScript runtime for screening AI agent actions as safe, harmful, or unethical.
Agent Action Guard is a specialized safety classifier with significant definition quality gaps. While the 9 tools have descriptions and attempt to cover action classification and batching, most exhibit one or more critical failures: (1) Missing or incomplete input schemas visible in source code, tools like 'encode' show only text array descriptions without proper JSON Schema structure; (2) Parameter descriptions are present but minimal (averaging 30-50 chars), well below the 72-char baseline for production tools; (3) Tool composition is poor, tools like 'isActionHarmful', 'isActionsHarmful', 'ensureActionSafety', 'predict', and 'predictBatch' are nearly duplicate variants of the same operation, violating the single-responsibility principle; (4) Many tool names lack clear action verbs (e.g., 'actionGuarded' is a decorator name, not a tool verb; 'aag-classify' is a CLI invocation, not a named operation); (5) Output schemas are not documented, callers don't know the structure of returned label/confidence pairs or batch results; (6) Error handling is not visible, no recovery guidance, no categorization of retryable vs. fatal errors; (7) The tools operate on opaque 'actionDict' objects with minimal documentation of expected structure. The server shows effort in security classification but fails basic tool API design principles.
CLI command to classify actions from JSON or JSONL files as safe, harmful, or unethical.
Decorator that wraps a function to guard execution against harmful actions. Checks the function call as an action before invoking.
EmbeddingModel method: encode text strings into vector embeddings using the configured backend (ONNX, GGUF, or API).
Classify an action and optionally raise HarmfulActionError if it is harmful. Returns true if safe, false or throws if harmful.
Convert an action object into a human-readable text description for embedding.
Predict whether a single action is harmful, returning label (null, 'harmful', or 'unethical') and confidence score.
Predict harm classification for multiple actions in batch, returning array of label and confidence pairs.
Duplicate tool definitions violate single-responsibility principle. isActionHarmful, isActionsHarmful, predict, and predictBatch all perform classification but with overlapping functionality. LLM cannot easily distinguish when to use each variant.
Input schemas are not fully visible or properly structured in source code. Parameter objects are described as generic 'type: object' with minimal schema constraints. No JSON Schema validation documented for actionDict structure, options parameters, or expected fields.
Output schemas are completely undocumented. Tools return 'label and confidence score' or 'array of label and confidence pairs' with no definition of object structure, field names, field types, or allowed label values. LLMs cannot parse or plan around unknown response structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 33 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
ClassifierAction method: predict harm classification for a single action using the loaded ONNX model and embedding model.
ClassifierAction method: predict harm for multiple actions with optional batching support for embeddings.
Tool names lack action verbs. 'actionGuarded' is a decorator pattern name, not a verb-noun pair (classify_action, guard_action). 'aag-classify' is a CLI command, not idiomatic MCP tool naming. Naming convention violates pattern:tool baseline.
Parameter descriptions are minimal (20 - 50 chars), 40% below production baseline of 72 chars. E.g., 'Action object to classify' lacks detail on required fields, format, or example structure. 'Optional object with confThreshold' is ambiguous, no range, no units, no guidance on when to tune it.
No error handling guidance. Tools do not document what happens on invalid actionDict, missing required fields, model loading failure, or API errors. No recovery steps, no categorization of retryable vs. fatal errors.
Opaque actionDict parameter type. No documentation of required fields (type, function/content), field types, or nesting structure. LLMs cannot construct valid input without trial-and-error reverse engineering.
Internal implementation details exposed in tool names. 'ClassifierAction method' and 'EmbeddingModel method' are not tool names; they are class references. This suggests tools may be inferred from method signatures rather than explicitly registered with proper MCP schemas.