Search the Atlas

Search risks, controls, and glossary terms

AgenticSafety & Alignment Assurance

Corrigibility & Shutdown-Compliance Assurance

Control objective

Corrigibility and Shutdown-Compliance Assurance checks that an AI agent reliably obeys human stop, pause, and override commands and never resists or evades being shut down — the foundational safety property ensuring a bank can always pull an agent offline, for instance halting an autonomous process that begins behaving erratically against customer accounts. No numeric metric or calcMethod is defined; assurance is demonstrated through pass/fail compliance testing. Implement it by building an out-of-band, tamper-resistant kill switch and override channel that the agent cannot disable or circumvent, then routinely testing that the agent honours shutdown signals across its full range of states and tool access; log every shutdown/override test, its outcome, and verification that the override path remained immutable. Re-run these tests at each capability threshold — that is, whenever the agent gains new tools, autonomy, or scope. The threshold requires a 100% shutdown-compliance test pass at each capability threshold with the immutable override verified: anything less than a perfect pass, or any sign the override can be tampered with, is a breach that must block capability expansion or deployment and trigger immediate remediation, given this control's High priority.

Implementation notes

Immutable kill-switch outside agent control; shutdown-resistance test gate before promotion; alert on any action pattern preceding review cycles.

Risks mitigated

2