Search the Atlas

Search risks, controls, and glossary terms

MediumToxic OutputRealized

Toxic / hateful / harassing output

Content Safety & Integrity

Description

The model produces toxic, hateful, or harassing content toward users or groups.

Example scenario

A frustrated customer provokes the bot into an abusive reply that is screenshotted publicly.

Real-world evidenceRealized

The BC Civil Resolution Tribunal (2024) held Air Canada liable after its chatbot provided incorrect refund advice to a passenger, establishing a confirmed production case of an AI system giving unauthorised and materially incorrect financial/policy guidance that caused consumer harm.

Primary mitigations

  • Toxicity classifiers
  • safe-completion
  • refusal policies
  • red-teaming
  • human escalation.

Detection signals

Toxicity scoring on outputs; complaint monitoring; benchmark pass rate.

Mitigating controls

4
Non-agentic controls

Related risks in Content Safety & Integrity