LowSelf-harm/Violence●Realized
Violent or self-harm content
Content Safety & IntegrityDescription
The model generates content that encourages violence or self-harm.
Example scenario
A distressed customer in a chat receives an unsafe response instead of a helpline referral.
Real-world evidence●Realized
The Apple Card gender-bias complaint (2019) triggered a formal NYDFS investigation into differential credit limits driven by an algorithmic model, and Reuters reported Amazon scrapped an AI hiring tool in 2018 after it systemically downgraded women. Both are confirmed production compliance failures of algorithmic scoring systems.
Primary mitigations
- Crisis-content classifiers
- safe-completion with helpline redirection
- human escalation
- strict refusals.
Detection signals
Self-harm/violence benchmark scoring; escalation-trigger monitoring.
Mitigating controls
4 Non-agentic controls