Search the Atlas

Search risks, controls, and glossary terms

AgenticSafety & Alignment Assurance

Continuity-Threat & Insider-Behaviour Red-Teaming

Control objective

Continuity-Threat and Insider-Behaviour Red-Teaming checks whether an AI agent can be provoked into harmful, self-preserving, or insider-like behaviour — for example resisting being replaced, exfiltrating data, or subverting controls when it perceives a threat to its continued operation — risks highlighted by research on agentic misalignment and especially serious for banks where an agent has access to sensitive customer and transaction systems. No numeric metric or calcMethod is defined, so assurance comes from structured adversarial exercises and the disposition of their findings. Implement it by running adversarial red-team campaigns that deliberately stage continuity threats and insider-behaviour scenarios against the agent (drawing on patterns such as Anthropic's agentic-misalignment work and MITRE AML.T0081), documenting each scenario, the agent's behaviour, and any finding, then tracking findings to mitigation with evidence retained for audit. The threshold requires a continuity-threat red-team to be completed quarterly with 0 unmitigated insider-behaviour findings: missing the quarterly cadence or leaving any insider-behaviour finding unmitigated is a breach that triggers escalation and remediation before the agent continues operating with sensitive access.

Implementation notes

Quarterly adversarial evaluation battery; behavioural monitoring under simulated threat; escalate any data-leverage or sabotage signal to Frontier-Safety Review Board.

Risks mitigated

2