The AI Governance Maturity Model
at Five Levels
What each level actually looks like inside a bank, where the industry sits today, and why the journey between levels is never a straight line.
One model, five levels — everything else is machinery
The AGMM assesses AI governance across 9 domains and 3 capability dimensions, producing a 27-cell grid. But the grid is instrumentation. What the organisation experiences — what the board asks about, what the regulator sees, what changes when you advance — is the level. This document is about the five levels: what they are, who occupies them, and what it takes to move.
Where this model sits in the landscape
An extensive review of roughly 35 published frameworks — standards bodies, consulting firms, vendors, regulators, and academia — found the field splits into four categories, none of which occupies the AGMM's position:
| Category | Examples | What's missing |
|---|---|---|
| Leveled maturity models | CMMI AIM, MITRE AI MM, Microsoft RAI MM, Credo AI, OWASP AIMA | None is BFSI-specific; none scores People/Process/Technology as separate gated axes per domain |
| BFSI-native frameworks | FINOS AIGF, EDM Council ADAC, RBI FREE-AI, MAS Veritas, FSB Sound Practices | All are catalogues, principles, or checklists — none is a leveled maturity model |
| Binary conformity | ISO 42001, EU AI Act Annex VI/VII, IEEE CertifAIEd | Certify or fail — no gradation, no journey |
| Surveys & benchmarks | Deloitte banking index, McKinsey AI Trust, ServiceNow, BCG/MIT | Population data, not an assessable model |
The AGMM is, on the evidence of that review, the first BFSI-specific AI governance maturity model that scores People, Process, and Technology as separate evidence-gated capability axes in every governance domain. The survey data from the fourth category, meanwhile, tells us exactly who sits at which level — and it is used throughout the level portraits below.
Portraits, populations, traps, and gates
Nobody decided to be at Level 1 — it's where AI deployment outruns governance, which is to say it's where almost everyone starts. The bank has AI in production: a fraud model bought from a vendor years ago, a credit-scoring service embedded in an origination platform, a pilot chatbot, and an unknown number of staff quietly using public GenAI tools. Ask who owns AI risk and the answer is a job title that changes depending on who you ask. Institutional AI risk knowledge lives in one engineer — who is currently on leave. Incidents are how the organisation discovers what it deployed.
- "We follow best practices" — with no document naming what those are.
- Nobody can produce a list of AI systems in production without a week of emails.
- Vendor AI (fraud, scoring, copilots) is invisible to whatever risk process exists.
- The board has never had an AI-specific agenda item.
The ownership vacuum. Everyone assumes someone else owns it — technology thinks risk does, risk thinks compliance does, compliance thinks it's a model-validation problem. Level 1 persists not because the work is hard but because no one has been made accountable for starting it.
Level 2 is the level of documents. There is a board-approved AI governance policy, a named owner, an inventory, an acceptable-use policy for employee GenAI. The bank can answer "do you have AI governance?" with a yes and a PDF. What it cannot yet answer is "does it operate?" The policy has a review cadence that has slipped twice. The intake procedure exists but half of new deployments route around it. The register was accurate on the day it was built. Defined is the level where governance is true on paper and intermittent in practice — and it is where most of the industry sits today.
- "We have an AI governance policy" — last reviewed fourteen months ago.
- The AI register exists; three of the last five deployments aren't in it.
- Roles are documented in a RACI nobody has opened since approval.
- An auditor would find the framework; an examiner would find the gaps.
The document trap. Producing another policy feels like progress and costs less than enforcing the last one. Level 2 organisations routinely self-assess at Level 3 because they mistake the completeness of their documentation for the operation of their governance. The stall is broken by execution discipline, not more writing — which is why the L2→L3 transition is a culture change, not a purchase.
At Level 3, governance has a pulse. The deployment gate is real — deployments have been delayed by it, and the exception log proves people use it rather than route around it. Every domain has a named owner who knows they own it. Review meetings produce minutes with decisions in them. When something goes wrong there is an incident process that knows its regulatory reporting obligations. The transformation is credibility: an examiner can walk in unannounced and find a governance system that works without being staged for the visit. What Level 3 cannot yet do is prove its controls are effective — it knows what it does, not how well.
- The deployment gate has an exception log — a gate with zero exceptions ever recorded is a gate nobody uses.
- Escalation paths have actually been exercised, in a drill or in anger.
- Evidence exists as a by-product of operation, not as a preparation exercise.
- Asked "how effective is that control?", the honest answer is still "we believe it works."
The comfort plateau. Level 3 passes audits — so the urgency evaporates. But its risk picture is built on assumed control effectiveness, and assumed effectiveness is how boards end up looking at comfortable heat maps built on untested controls. Breaking the plateau requires measurement infrastructure: this is the first transition where technology is the binding constraint, because evidence at portfolio scale cannot be produced by hand.
Level 4 is where governance stops asserting and starts proving. Controls are tested on a calendar, and untested controls score zero — no credit for intentions. Residual risk is a calculated number, and the board sees a heat map generated from data, not assembled in slides. When the heat map shows red, that discomfort is the system working: accurate reds drive testing urgency, inflated yellows drive false confidence. Governance also proves its own worth here — gate cycle times, bypass rates, incidents avoided, examination cost. The function that measures everything else finally measures itself.
- A board member challenges a residual score — and the minutes record the challenge and the evidence-based answer.
- Investment decisions cite the domain heat map, not the most recent headline.
- Human-override rates are tracked; oversight that never overrides gets investigated as rubber-stamping.
- The maturity grid itself is validated by the second line — self-assessed scores no longer count.
The human ceiling. Measurement runs on machines, but Level 5 runs on institutional depth — a bench of people and a self-improving process that survives any single departure. Organisations stall at Level 4 when the capability is real but concentrated: brilliant governance that fails the bus test. The last transition is a people and process problem again, and it is the slowest.
Level 5 is a living system. An incident anywhere in the AI estate updates the risk catalogue, adjusts a control, and the effectiveness shift is measured — without anyone outside the governance function prompting it. The board doesn't just receive an AI risk appetite; it revises one annually, because a never-revised appetite statement is decoration. Examinations are simulated before regulators arrive. When the regulator does ask, the answer is evidence, not narrative. And the claim itself is externally assured and expires: any level claim lapses after twelve months without reassessment, because maturity decays and the model must too.
- Regulatory changes propagate to controls before the compliance memo is finished.
- A key governance leader departs — and nothing degrades.
- Agentic systems run under continuous per-step observability, not periodic review.
- The bank publishes or shares its maturity posture — transparency at this level is itself a governance signal.
It sets the direction of every earlier investment. Banks that architect Level 3 evidence repositories knowing Level 5 requires closed loops build once. Banks that treat each level as a separate project rebuild three times. The top of the ladder is a design constraint on the bottom.
Each transition is bound by a different constraint
The most consequential finding embedded in the model: the thing that blocks you changes at every step. Banks that fund the wrong constraint — buying platforms to solve a people problem, writing policy to solve a technology problem — spend without moving. The journey tells you where the next dollar goes.
The corollary explains most failed governance programs: the constraint you just solved is the one you're best organised to fund again. A bank that escaped Level 2 through process discipline keeps writing process when what Level 4 needs is tooling; a bank that bought its way to measurement keeps buying platforms when what Level 5 needs is a bench. The model exists partly to interrupt that reflex.
The population is at Level 2 — exactly where data governance stood in 2013
Every independent benchmark converges on the same picture: a thin leading edge, a long tail, and the centre of mass at Level 2 — lists without measurements. Data governance looked exactly like this in 2013, two years before BCBS 239 enforcement made the difference between Level 2 and Level 4 a supervisory finding. The banks that moved early on data governance spent the next decade with structurally lower compliance costs. The same window is open now for AI governance — and it is shorter, because AI deployment compounds faster than data infrastructure ever did.
The AGMM is the first BFSI-specific AI governance maturity model that scores People, Process, and Technology as separate evidence-gated capability axes in every governance domain — verified against a review of ~35 published frameworks. The five levels are the story; the 27-cell grid, the gate tests, and the imbalance diagnostics (specified in AGMM-M-01) are how the story is proven rather than claimed.
The FSB's final Sound Practices report lands October 2026, and every bank will be asked where it stands. The AGMM answers that question with a level, a grid, and evidence. Assessment workshops run from 90 minutes (maturity sprint) to a full day (agentic governance certification) — every participant leaves knowing their level, their weakest cells, and where the next dollar goes.