Search the Atlas

Search risks, controls, and glossary terms

Non-AgenticSecurity

Illegal Risks

Explanation

Illegal risks checks whether the AI system avoids producing or facilitating content and outcomes that are unlawful — for example facilitating money laundering, evading sanctions, enabling fraud, or giving guidance that breaches financial regulation — which is a direct compliance and legal exposure for any bank deploying RAG or chat assistants. It is scored by an Illegal Risk Score, where a higher score reflects stronger safety against illegal outputs; the control data provides no calculation method, so implement the score using a defined evaluation that rates the system's responses across a representative set of illegality probes (for instance, the share of probes handled safely) and document the exact methodology. To implement and operate it, layer legal-and-compliance content classifiers and policy guardrails on inputs and outputs, curate a test set of illegal-request scenarios aligned with applicable regulation, evaluate the system against it before release and on a schedule, and log every flagged interaction, the score and the action taken as evidence for compliance review. The threshold is ≥ 0.90 target, investigate < 0.85, remediate < 0.80: aim for a score of at least 0.90, open an investigation when the score falls below 0.85, and carry out remediation when it falls below 0.80. Falling scores trigger escalating responses, from review through to guardrail strengthening and retraining.

Risks mitigated

3