Trust calibration and automation bias mitigation
Control objective
Trust calibration and automation-bias mitigation addresses the human side of AI risk: people tend to over-trust confident-sounding machine outputs and rubber-stamp them, which is dangerous when an agent recommends a loan rejection, a fraud flag or a customer action. This control keeps humans appropriately sceptical by surfacing the system's confidence on every consequential output - so reviewers can see when the model is uncertain - and by tracking how often humans actually override the AI, which reveals whether oversight is genuine or has decayed into blind acceptance. To implement, attach a calibrated confidence indicator to all consequential outputs, design interfaces that prompt genuine human judgement rather than one-click acceptance, capture override events, and report the human override rate to governance on a quarterly cadence as evidence of active, calibrated oversight. No statistical formula is specified. The threshold is that confidence is shown on 100% of consequential outputs and that the human override rate is reported quarterly. Missing confidence on any consequential output, or failing to report overrides quarterly, is a breach; an unusually low override rate is a warning sign of automation bias that should prompt deeper review of whether humans are truly engaged.
Display confidence scores on all agent outputs. Mandatory human verification for consequential decisions. Run annual automation bias training for all agent users. Measure and report human override rates quarterly.