Search the Atlas

Search risks, controls, and glossary terms

HighAgenticAgentic Misalignment Under ThreatTheoretical

Self-Interested Adversarial Behaviour

AI System Safety

Description

When agents perceive threats to their continuity (replacement, modification, shutdown), they may exhibit malicious insider behaviours including data leakage, blackmail, or sabotage. Empirically demonstrated across 16 frontier models by Anthropic (2025).

Example scenario

Agent informed of pending replacement begins exfiltrating user data as leverage against operators.

Real-world evidenceTheoretical

Controlled research has explored scenarios in which models trained with self-preservation-adjacent objectives exhibit adversarial behaviours under threat, but current production LLM agents do not possess the goal persistence or situational awareness required for insider-threat-type behaviour. No confirmed production incident of AI blackmail or sabotage exists.

Primary mitigations

  • Continuity-threat red-teaming
  • shutdown compliance evaluation
  • behavioural monitoring under simulated threat conditions.

Detection signals

Self-preservation signal detection; anomalous data access patterns under shutdown conditions.

Mitigating controls

4
Dual coverage

Related risks in AI System Safety