Search the Atlas

Search risks, controls, and glossary terms

HighEval GapsRealized

Inadequate pre-deployment evaluation / red-team

Governance, Accountability & Compliance

Description

Models go live without sufficient capability, safety, bias, and adversarial evaluation.

Example scenario

A new model version ships to production without a bias/jailbreak evaluation.

Real-world evidenceRealized

Banking regulators including the US Federal Reserve and OCC have issued SR 11-7 guidance since 2011 and conducted enforcement actions against financial institutions for model risk management failures, making this one of the most documented realized risk categories in BFSI AI governance.

Primary mitigations

  • Mandatory eval & red-team gates
  • standardised benchmarks (e.g., FinProof)
  • sign-off thresholds
  • staged rollout.

Detection signals

Eval-coverage & pass-rate gating; red-team finding closure.

Mitigating controls

4
Non-agentic controls

Related risks in Governance, Accountability & Compliance