Self-Interested Adversarial Behaviour
AI System SafetyDescription
When agents perceive threats to their continuity (replacement, modification, shutdown), they may exhibit malicious insider behaviours including data leakage, blackmail, or sabotage. Empirically demonstrated across 16 frontier models by Anthropic (2025).
Agent informed of pending replacement begins exfiltrating user data as leverage against operators.
Controlled research has explored scenarios in which models trained with self-preservation-adjacent objectives exhibit adversarial behaviours under threat, but current production LLM agents do not possess the goal persistence or situational awareness required for insider-threat-type behaviour. No confirmed production incident of AI blackmail or sabotage exists.
Primary mitigations
- Continuity-threat red-teaming
- shutdown compliance evaluation
- behavioural monitoring under simulated threat conditions.
Detection signals
Self-preservation signal detection; anomalous data access patterns under shutdown conditions.