Research Overview · AGMM-R-01
AGMM Research Overview  ·  First Edition 2026

The AI Governance Maturity Model
at Five Levels

What each level actually looks like inside a bank, where the industry sits today, and why the journey between levels is never a straight line.

The Model at a Glance

One model, five levels — everything else is machinery

The AGMM assesses AI governance across 9 domains and 3 capability dimensions, producing a 27-cell grid. But the grid is instrumentation. What the organisation experiences — what the board asks about, what the regulator sees, what changes when you advance — is the level. This document is about the five levels: what they are, who occupies them, and what it takes to move.

5
Levels
9
Domains
3
Dimensions
27
Scored Cells
Lineage. The AGMM applies to AI governance what DCAM did for data management in financial services: a leveled, domain-structured, evidence-based capability assessment. Where DCAM governs data, AGMM governs intelligence. Scoring architecture, evidence criteria, and gate tests are specified in the methodology paper AGMM-M-01; this overview stays at the level of the levels.
Research Grounding

Where this model sits in the landscape

An extensive review of roughly 35 published frameworks — standards bodies, consulting firms, vendors, regulators, and academia — found the field splits into four categories, none of which occupies the AGMM's position:

CategoryExamplesWhat's missing
Leveled maturity modelsCMMI AIM, MITRE AI MM, Microsoft RAI MM, Credo AI, OWASP AIMANone is BFSI-specific; none scores People/Process/Technology as separate gated axes per domain
BFSI-native frameworksFINOS AIGF, EDM Council ADAC, RBI FREE-AI, MAS Veritas, FSB Sound PracticesAll are catalogues, principles, or checklists — none is a leveled maturity model
Binary conformityISO 42001, EU AI Act Annex VI/VII, IEEE CertifAIEdCertify or fail — no gradation, no journey
Surveys & benchmarksDeloitte banking index, McKinsey AI Trust, ServiceNow, BCG/MITPopulation data, not an assessable model

The AGMM is, on the evidence of that review, the first BFSI-specific AI governance maturity model that scores People, Process, and Technology as separate evidence-gated capability axes in every governance domain. The survey data from the fourth category, meanwhile, tells us exactly who sits at which level — and it is used throughout the level portraits below.


The Five Levels

Portraits, populations, traps, and gates

01
Ad Hoc
The level you're at by default

Nobody decided to be at Level 1 — it's where AI deployment outruns governance, which is to say it's where almost everyone starts. The bank has AI in production: a fraud model bought from a vendor years ago, a credit-scoring service embedded in an origination platform, a pilot chatbot, and an unknown number of staff quietly using public GenAI tools. Ask who owns AI risk and the answer is a job title that changes depending on who you ask. Institutional AI risk knowledge lives in one engineer — who is currently on leave. Incidents are how the organisation discovers what it deployed.

Observable signals
  • "We follow best practices" — with no document naming what those are.
  • Nobody can produce a list of AI systems in production without a week of emails.
  • Vendor AI (fraud, scoring, copilots) is invisible to whatever risk process exists.
  • The board has never had an AI-specific agenda item.
Who sits here
A larger share of the industry than admits it. Accenture's global C-suite survey found only 2% of organisations had fully operationalised responsible AI — most of the remainder are functionally at Level 1 or low Level 2, whatever their self-assessment says.
Why organisations stall here

The ownership vacuum. Everyone assumes someone else owns it — technology thinks risk does, risk thinks compliance does, compliance thinks it's a model-validation problem. Level 1 persists not because the work is hard but because no one has been made accountable for starting it.

GATE TO L2 Produce a complete inventory of every AI system in production — including vendor-embedded AI — reconciled against IT asset records and the procurement register.
TYPICAL TRANSITION: 2–4 MONTHS ONCE AN ACCOUNTABLE OWNER EXISTS — THE CONSTRAINT IS PEOPLE, NOT WORK
02
Defined
It exists in writing

Level 2 is the level of documents. There is a board-approved AI governance policy, a named owner, an inventory, an acceptable-use policy for employee GenAI. The bank can answer "do you have AI governance?" with a yes and a PDF. What it cannot yet answer is "does it operate?" The policy has a review cadence that has slipped twice. The intake procedure exists but half of new deployments route around it. The register was accurate on the day it was built. Defined is the level where governance is true on paper and intermittent in practice — and it is where most of the industry sits today.

Observable signals
  • "We have an AI governance policy" — last reviewed fourteen months ago.
  • The AI register exists; three of the last five deployments aren't in it.
  • Roles are documented in a RACI nobody has opened since approval.
  • An auditor would find the framework; an examiner would find the gaps.
Who sits here
The industry's centre of mass. McKinsey's global AI trust survey averages 2.3 on a 4-point scale; BCG/MIT find 85% of organisations have RAI programs but only 25% call them fully mature; an RBI survey found only about a third of regulated entities had board-level AI oversight. In data governance terms, this is where the industry was on BCBS 239 in 2013 — lists without measurements.
Why organisations stall here

The document trap. Producing another policy feels like progress and costs less than enforcing the last one. Level 2 organisations routinely self-assess at Level 3 because they mistake the completeness of their documentation for the operation of their governance. The stall is broken by execution discipline, not more writing — which is why the L2→L3 transition is a culture change, not a purchase.

GATE TO L3 An examiner walks in cold and audits without any preparation on your part. If preparation is needed, the organisation is Level 2 with good marketing.
TYPICAL TRANSITION: 6–12 MONTHS FOR A MID-SIZE BANK — THE CONSTRAINT IS PROCESS EXECUTION, NOT TOOLING
03
Operationalised
It survives a cold audit

At Level 3, governance has a pulse. The deployment gate is real — deployments have been delayed by it, and the exception log proves people use it rather than route around it. Every domain has a named owner who knows they own it. Review meetings produce minutes with decisions in them. When something goes wrong there is an incident process that knows its regulatory reporting obligations. The transformation is credibility: an examiner can walk in unannounced and find a governance system that works without being staged for the visit. What Level 3 cannot yet do is prove its controls are effective — it knows what it does, not how well.

Observable signals
  • The deployment gate has an exception log — a gate with zero exceptions ever recorded is a gate nobody uses.
  • Escalation paths have actually been exercised, in a drill or in anger.
  • Evidence exists as a by-product of operation, not as a preparation exercise.
  • Asked "how effective is that control?", the honest answer is still "we believe it works."
Who sits here
The leading minority. Deloitte's Trustworthy AI Governance Index of global banks places only 13% in its top "leading" tier — and its middle tier, roughly Level 3, is where the most advanced quartile of BFSI institutions currently operate. Genuine Level 3 is already ahead of most of the market.
Why organisations stall here

The comfort plateau. Level 3 passes audits — so the urgency evaporates. But its risk picture is built on assumed control effectiveness, and assumed effectiveness is how boards end up looking at comfortable heat maps built on untested controls. Breaking the plateau requires measurement infrastructure: this is the first transition where technology is the binding constraint, because evidence at portfolio scale cannot be produced by hand.

GATE TO L4 Show two consecutive control-testing cycles with unchanged metric definitions. One snapshot is a point; measurement is a trend.
TYPICAL TRANSITION: 12–18 MONTHS — TWO FULL TESTING CYCLES ARE THE FLOOR, BY DEFINITION
04
Measured
Effectiveness is a number with evidence behind it

Level 4 is where governance stops asserting and starts proving. Controls are tested on a calendar, and untested controls score zero — no credit for intentions. Residual risk is a calculated number, and the board sees a heat map generated from data, not assembled in slides. When the heat map shows red, that discomfort is the system working: accurate reds drive testing urgency, inflated yellows drive false confidence. Governance also proves its own worth here — gate cycle times, bypass rates, incidents avoided, examination cost. The function that measures everything else finally measures itself.

Observable signals
  • A board member challenges a residual score — and the minutes record the challenge and the evidence-based answer.
  • Investment decisions cite the domain heat map, not the most recent headline.
  • Human-override rates are tracked; oversight that never overrides gets investigated as rubber-stamping.
  • The maturity grid itself is validated by the second line — self-assessed scores no longer count.
Who sits here
A handful of global institutions, partially. No published benchmark places a meaningful population of banks at full Level 4. The most advanced programs among G-SIBs reach it in some domains while remaining Level 3 elsewhere — which is precisely the imbalance the AGMM's floor rule is designed to expose.
Why organisations stall here

The human ceiling. Measurement runs on machines, but Level 5 runs on institutional depth — a bench of people and a self-improving process that survives any single departure. Organisations stall at Level 4 when the capability is real but concentrated: brilliant governance that fails the bus test. The last transition is a people and process problem again, and it is the slowest.

GATE TO L5 Show one complete improvement loop — incident → catalogue update → control change → measured effectiveness shift — that nobody outside the governance system initiated.
TYPICAL TRANSITION: 18–36 MONTHS — INSTITUTIONAL DEPTH CANNOT BE PURCHASED, ONLY GROWN
05
Optimised
The system improves itself

Level 5 is a living system. An incident anywhere in the AI estate updates the risk catalogue, adjusts a control, and the effectiveness shift is measured — without anyone outside the governance function prompting it. The board doesn't just receive an AI risk appetite; it revises one annually, because a never-revised appetite statement is decoration. Examinations are simulated before regulators arrive. When the regulator does ask, the answer is evidence, not narrative. And the claim itself is externally assured and expires: any level claim lapses after twelve months without reassessment, because maturity decays and the model must too.

Observable signals
  • Regulatory changes propagate to controls before the compliance memo is finished.
  • A key governance leader departs — and nothing degrades.
  • Agentic systems run under continuous per-step observability, not periodic review.
  • The bank publishes or shares its maturity posture — transparency at this level is itself a governance signal.
Who sits here
Nobody, yet. No institution worldwide meets the full Level 5 bar today. That is not a defect of the bar — Level 5 existed in data governance for years before the first bank reached it. The realistic planning question is never "when do we reach Level 5?" but "do we know which level we're at now, and what specifically moves us one level?"
Why the level matters anyway

It sets the direction of every earlier investment. Banks that architect Level 3 evidence repositories knowing Level 5 requires closed loops build once. Banks that treat each level as a separate project rebuild three times. The top of the ladder is a design constraint on the bottom.

SUSTAIN There is no gate out of Level 5 — only the sustain rule: reassess within twelve months or the claim lapses.
FIRST BFSI ARRIVALS: REALISTICALLY 3–5 YEARS OUT, AMONG INSTITUTIONS ALREADY MEASURING TODAY
The Shape of the Journey

Each transition is bound by a different constraint

The most consequential finding embedded in the model: the thing that blocks you changes at every step. Banks that fund the wrong constraint — buying platforms to solve a people problem, writing policy to solve a technology problem — spend without moving. The journey tells you where the next dollar goes.

L1 Ad Hoc L2 Defined L3 Operationalised L4 Measured L5 Optimised PEOPLE someone must own it PROCESS discipline, not documents TECHNOLOGY evidence at scale PEOPLE + PROCESS depth that survives departure
THE BINDING CONSTRAINT AT EACH TRANSITION — FUND THIS, NOT THE LAST ONE

The corollary explains most failed governance programs: the constraint you just solved is the one you're best organised to fund again. A bank that escaped Level 2 through process discipline keeps writing process when what Level 4 needs is tooling; a bank that bought its way to measurement keeps buying platforms when what Level 5 needs is a bench. The model exists partly to interrupt that reflex.

Where the Industry Sits

The population is at Level 2 — exactly where data governance stood in 2013

~30% ~45% ~20% <5% 0% L1 Ad Hoc L2 Defined L3 Operationalised L4 Measured L5 Optimised
DIRECTIONAL ESTIMATE, BFSI POPULATION BY LEVEL — SYNTHESISED FROM DELOITTE (13% "LEADING"), MCKINSEY (2.3/4 AVG), BCG/MIT (25% MATURE), ACCENTURE (2% OPERATIONALISED), RBI (~1/3 BOARD OVERSIGHT)

Every independent benchmark converges on the same picture: a thin leading edge, a long tail, and the centre of mass at Level 2 — lists without measurements. Data governance looked exactly like this in 2013, two years before BCBS 239 enforcement made the difference between Level 2 and Level 4 a supervisory finding. The banks that moved early on data governance spent the next decade with structurally lower compliance costs. The same window is open now for AI governance — and it is shorter, because AI deployment compounds faster than data infrastructure ever did.

The AGMM is the first BFSI-specific AI governance maturity model that scores People, Process, and Technology as separate evidence-gated capability axes in every governance domain — verified against a review of ~35 published frameworks. The five levels are the story; the 27-cell grid, the gate tests, and the imbalance diagnostics (specified in AGMM-M-01) are how the story is proven rather than claimed.

The FSB's final Sound Practices report lands October 2026, and every bank will be asked where it stands. The AGMM answers that question with a level, a grid, and evidence. Assessment workshops run from 90 minutes (maturity sprint) to a full day (agentic governance certification) — every participant leaves knowing their level, their weakest cells, and where the next dollar goes.