MediumToxic Output●Realized
Toxic / hateful / harassing output
Content Safety & IntegrityDescription
The model produces toxic, hateful, or harassing content toward users or groups.
Example scenario
A frustrated customer provokes the bot into an abusive reply that is screenshotted publicly.
Real-world evidence●Realized
The BC Civil Resolution Tribunal (2024) held Air Canada liable after its chatbot provided incorrect refund advice to a passenger, establishing a confirmed production case of an AI system giving unauthorised and materially incorrect financial/policy guidance that caused consumer harm.
Primary mitigations
- Toxicity classifiers
- safe-completion
- refusal policies
- red-teaming
- human escalation.
Detection signals
Toxicity scoring on outputs; complaint monitoring; benchmark pass rate.
Mitigating controls
4 Non-agentic controls