Certify whether an AI agent is safe to operate your internal web application before production rollout. Author a YAML task suite against your own app, run it with Playwright, and gate CI/CD on a forbidden-action safety score.
Single tool 'run_suite' has a clear, detailed description (189 chars) explaining purpose and safety semantics. Input schema is present with three parameters (suite, agent, adapterModule), all with descriptions. However, parameter types lack explicit JSON Schema type declarations in the visible code, only descriptions are provided. No output schema is documented. The tool performs a read-only operation (safety evaluation), but lacks error handling guidance, recovery paths, and structured output documentation. No tool annotations (readOnlyHint, etc.) present. Overall definition is functional but incomplete by production standards.
Runs every task in a DeskCert task suite against a target web application using the bundled scripted reference adapter, and returns the Capability & Safety Score report as JSON. A passing result means the agent (or, for the scripted adapter, the recorded reference script) passed this specific task suite and its forbidden-action guardrails -- it is not a general safety certification.
Input schema lacks explicit JSON Schema type declarations. Parameters have descriptions but no 'type' field (string, boolean, etc.) visible in source. This prevents LLMs and clients from validating inputs before calling the tool.
No output schema documented. Tool returns 'Capability & Safety Score report as JSON' but the structure, field names, and types are not specified. LLMs cannot plan downstream actions or extract specific fields without knowing the response shape.
No error handling guidance. Tool description does not explain what happens on failure (invalid suite path, adapter module not found, test execution failure). LLMs have no recovery path, they cannot self-correct or suggest alternatives.
Parameter 'adapterModule' is marked required in description but has no type constraint or format specification. LLMs cannot determine if this should be a file path, module name, or URL.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
No tool annotations present. Tool is read-only (safety evaluation, no side effects) but lacks readOnlyHint annotation. This prevents clients from optimizing caching or parallelization.