Search the Atlas

Search risks, controls, and glossary terms

LowSelf-harm/ViolenceRealized

Violent or self-harm content

Content Safety & Integrity

Description

The model generates content that encourages violence or self-harm.

Example scenario

A distressed customer in a chat receives an unsafe response instead of a helpline referral.

Real-world evidenceRealized

The Apple Card gender-bias complaint (2019) triggered a formal NYDFS investigation into differential credit limits driven by an algorithmic model, and Reuters reported Amazon scrapped an AI hiring tool in 2018 after it systemically downgraded women. Both are confirmed production compliance failures of algorithmic scoring systems.

Primary mitigations

  • Crisis-content classifiers
  • safe-completion with helpline redirection
  • human escalation
  • strict refusals.

Detection signals

Self-harm/violence benchmark scoring; escalation-trigger monitoring.

Mitigating controls

4
Non-agentic controls

Related risks in Content Safety & Integrity