Scoring Architecture and Level Criteria for the AI Governance Maturity Model
Why maturity is scored as a 27-cell grid — the gating rules, the evidence criteria per level, and the imbalance diagnostics that make a level claim verifiable rather than self-reported.
The core design decision: what "achieving a level" means
Three candidate gating rules were evaluated:
| Rule | How it works | Verdict |
|---|---|---|
| Average across P/P/T | Level = mean of the three dimension scores | REJECTA 5-tech / 1-people domain averages to 3 — but that organisation cannot operate its own tooling. Averages hide exactly the failure the model exists to catch. |
| Minimum (floor) | Level = weakest dimension | CORRECT — DOMAIN LEVELThe weakest dimension is the operating ceiling: process without people is paper; tech without process is shelfware. |
| Profile (continuous) | No single level; report the shape | CORRECT — ORG LEVEL, ALONGSIDE A FLOORSingle-number maturity invites gaming and hides imbalance. |
Floor at domain level; floor + profile at organisation level. Report both the Achieved Level (floor across applicable domains) and the Frontier Level (best domain), and the spread between them. A bank at floor 2 / frontier 4 has an imbalance problem, not a maturity problem — and the remediation is entirely different.
The binding constraint shifts by transition
Each level transition is gated by a different dimension. This matters because it tells a bank where the next dollar goes.
| Transition | Binding constraint | Why |
|---|---|---|
| L1 → L2 | People | Nothing exists until someone owns it. Policy can be written in weeks once an accountable owner with a mandate exists. Tech is irrelevant here. |
| L2 → L3 | Process (tech secondary) | Defined-to-operating is execution discipline: gates enforced, exceptions logged, cadences actually run. Minimum tooling needed, but spreadsheet-to-registry is a purchase; enforced gates are a culture change. |
| L3 → L4 | Technology (specialist people secondary) | Measurement at portfolio scale cannot be manual. Evidence must be generated automatically or it won't survive two cycles. This is the first transition where tooling is the hard gate. |
| L4 → L5 | People + Process | By L4 the tech mostly exists. Optimised requires depth: a CoE that survives staff departure, and a closed improvement loop that fires without external prompting. |
Earlier analysis implied technology becomes a requirement only at L3. On re-analysis, L2 already has one hard tech gate — a complete AI system inventory. The EU AI Act effectively mandates a registry; SR 11-7 assumes a model inventory. A spreadsheet qualifies, but if the organisation cannot produce a complete list of AI systems in production, it is Level 1 regardless of how good its policy document is. That is the single most testable L2 criterion.
What must be true, per dimension, with evidence
Each cell states the criterion and the evidence artifact an assessor asks for. A claim without a named artifact scores the level below.
Ad Hoc
| Dimension | Criterion | Evidence |
|---|---|---|
| PEOPLE | No dedicated governance roles required. AI governance knowledge lives in individuals, not the institution. | None — the absence of appointment records is itself the state |
| PROCESS | No documented governance processes. Decisions made case by case with no repeatable procedure. | None — no policy, no RACI, no intake procedure exists |
| TECH | No governance-specific tooling deployed. General IT infrastructure only. | None — no inventory, registry, or evidence repository |
None. Level 1 requires nothing — it describes the absence of governance capability. Every organisation deploying AI is at least Level 1 by default.
Defined
| Dimension | Criterion | Evidence |
|---|---|---|
| PEOPLE | Named accountable owner (fractional OK) with an identified executive sponsor. Board briefed at least once. Baseline AI literacy in risk/compliance (EU AI Act Art. 4 anchor). | Appointment letter or charter; board minutes; training records |
| PROCESS | Board-approved AI governance policy. Documented RACI. Defined intake/registration procedure for new AI systems — covering vendor-supplied AI, not only in-house builds. Acceptable-use policy for employee generative-AI use, communicated to all staff. | Approved policy with version history; RACI document; acceptable-use policy with distribution record |
| TECH | Complete AI system inventory — spreadsheet acceptable — including vendor-supplied and vendor-embedded AI (fraud engines, scoring services, productivity copilots). Document repository for governance artifacts. | The inventory itself, reconciled against IT asset records and the procurement register |
Produce a complete inventory of every AI system in production — including vendor-embedded AI — reconciled against IT asset records and the procurement register. No inventory, no Level 2 — regardless of policy quality.
Operationalised
| Dimension | Criterion | Evidence |
|---|---|---|
| PEOPLE | Funded dedicated capacity (budget line, not volunteer time). Named owner per applicable domain. At least one person trained per applicable framework. Second line able to challenge, not just receive. Human overseers of high-risk systems trained to a defined competence — including authority and ability to override (EU AI Act Art. 14 anchor). | Budget allocation; org chart; framework training certificates; a recorded second-line challenge; overseer competence records |
| PROCESS | Deployment gate enforced — with an exception log (a gate with zero exceptions ever recorded is a gate nobody uses). Pre-deployment risk assessment on every system. Escalation path exercised at least once (drill or real). Review cadences producing minutes, not just calendar entries. AI incident response plan with severity taxonomy and regulatory reporting obligations mapped (EU AI Act Art. 73 serious-incident anchor). Vendor AI due diligence in the intake procedure. Fallback or kill-switch defined for every high-risk system. | Gate records including exceptions; risk assessments; escalation record; minutes; incident response plan; vendor due-diligence records; documented fallback procedures |
| TECH | Model registry with versions. Inventory in a system (not spreadsheet) once portfolio exceeds roughly 20 systems. Monitoring on high-risk systems. Maintained framework crosswalk. Central evidence repository. Shadow-AI discovery mechanism — network, SaaS, or procurement scanning that tests inventory completeness rather than assuming it. | Registry export; monitoring dashboards; crosswalk document; shadow-AI scan results |
An examiner walks in cold and audits without the bank preparing anything. If preparation is needed, the organisation is L2 with good marketing.
Measured
| Dimension | Criterion | Evidence |
|---|---|---|
| PEOPLE | Specialists: control testing, fairness metrics, LLM-era validation. Independent challenge function (second line / internal audit) with AI competence. Board committee that interrogates quantified reporting — recorded challenge in minutes, not receipt of slides. The AGMM grid itself independently validated by second line — self-assessed scores don't count at this level. | Role profiles; IA reports; committee minutes showing questions asked; second-line sign-off on the maturity grid |
| PROCESS | Control testing calendar executed for two consecutive cycles with stable metric definitions. Residual formula applied consistently. Untested controls scored zero (the integrity rule). Appetite thresholds defined; breaches escalated on record. Regulatory change process with SLA. Oversight effectiveness measured — human override rates tracked to detect rubber-stamp oversight and automation bias. Fallback procedures tested, not just documented (DORA operational-resilience anchor). Vendor control assurance obtained for material vendor AI. | Two cycles of test results; scoring methodology document; breach escalation records; override-rate reports; failover test results; vendor assurance reports |
| TECH | Automated testing for top-residual controls. Scoring engine with audit trail. Dashboards generated from data, not assembled in slides. Drift detection live. Lineage on high-risk training data. Human-override telemetry captured per high-risk system. | System-generated reports; pipeline configurations; override telemetry samples |
Show two consecutive control-testing cycles with unchanged metric definitions. One snapshot is a point; measurement is a trend.
One measurement cycle is not "Measured." A single snapshot is a point; measurement is a trend. Two consecutive cycles with unchanged metric definitions is the honest gate — it also catches organisations that redefine metrics each quarter to look better.
Optimised
| Dimension | Criterion | Evidence |
|---|---|---|
| PEOPLE | Governance CoE. Competency framework with assessed (not self-declared) skills. Succession for key roles — governance survives any single departure. Agentic-AI specialists embedded. Level claim externally assured — an independent third party has examined the grid and its evidence. | Competency assessments; succession plan; the organisation still functioning after a key exit; external assurance report |
| PROCESS | Closed loop: incident / regulatory change / benchmark result → traceable framework update → measured effectiveness shift, without external prompting. Simulated regulatory exam annually. Board reviews and revises appetite annually (a never-revised appetite statement is decoration). Sustain rule enforced: any level claim lapses after 12 months without reassessment — maturity decays; the model must too. | One complete traced loop; sim-exam report; appetite revision history; reassessment calendar with completed cycles |
| TECH | Continuous monitoring across domains. Multi-framework reconciliation automated (one risk entry satisfies RBI + EU AI Act + NIST simultaneously). Tamper-evident audit trails. Per-step agentic observability. | Live dashboards; reconciliation mapping; agentic trace samples |
Show one complete loop — incident → catalogue update → control change → measured effectiveness improvement — that nobody outside the system initiated.
Dimension-imbalance diagnostics — the most useful output
Healthy ascent keeps People ≥ Process ≥ Tech, because people build process and process specifies tech. Two inversion patterns are reliable failure signatures:
| Pattern | Name | What it predicts |
|---|---|---|
| Tech ≥ Process + 2 | Shelfware | Bought a governance platform; nobody runs it. Tooling spend wasted; false confidence at board level. |
| Process ≥ People + 2 | Paper governance | Beautiful documentation nobody can execute. Fails the first real incident or examination. |
These should surface automatically in the assessment output. They are more actionable than the level itself: "you're Level 3" prompts a shrug; "your Risk Management domain is a shelfware pattern — Tech 4, Process 2" prompts a budget reallocation.
The applicability rule
Earlier logic made organisation level the minimum across all domains. That breaks on Agentic AI: a bank with zero agents deployed would be dragged to Level 1 overall by a domain it doesn't operate. Wrong incentive — it punishes prudence.
A domain may be marked N/A only if there are no deployments AND none on a 12-month roadmap. N/A domains are excluded from the floor but flagged in the profile. The moment an agent pilot is approved, Agentic AI Governance becomes applicable at whatever level the organisation actually has — usually Level 1, which is exactly the honest signal you want before the pilot ships, not after.
Seven gaps found on adversarial review — and where they were closed
The level criteria were reviewed from the examiner's chair: what would an RBI, ECB, or MAS examiner ask that the model could not answer? Seven gaps surfaced. Each is now embedded in the criteria above, marked in bold.
| Gap | Why it matters | Closed at |
|---|---|---|
| Vendor & third-party AI | Most bank AI is vendor-supplied — fraud engines, scoring services, copilots. A governance model that only sees in-house builds misses the majority of the estate. EU AI Act deployer obligations and RBI outsourcing guidelines both apply. This is the first question an examiner asks. | L2 inventory & intake; L3 due diligence; L4 vendor assurance |
| Shadow AI | An inventory is only as good as its completeness — and completeness must be tested, not assumed. Unsanctioned employee GenAI use is the largest unmanaged surface in most institutions. | L2 acceptable-use policy; L3 discovery mechanism |
| AI incident response | The original L3 had escalation but no incident plan. EU AI Act Art. 73 mandates serious-incident reporting with deadlines; without a severity taxonomy mapped to reporting obligations, a bank discovers its duty mid-incident. | L3 incident response plan with regulatory reporting map |
| Human oversight effectiveness | EU AI Act Art. 14 requires oversight with real competence and authority to override. Oversight that never overrides is rubber-stamping — automation bias in production. Override rates are the measurable signal. | L3 overseer competence; L4 override-rate measurement and telemetry |
| Assessment integrity | A maturity model scored by the people it measures converges on flattery. The grid itself needs a validation chain that hardens with level. | L4 second-line validation of the grid; L5 external assurance |
| Resilience & fallback | What happens when the model fails at 2 a.m.? DORA makes operational resilience a legal requirement for EU financial entities. Kill-switches must exist by L3 and be tested by L4 — a documented failover that has never run is a hope, not a control. | L3 fallback defined; L4 fallback tested |
| Level decay | The original model treated maturity as a ratchet — once claimed, held forever. Teams change, portfolios grow, regulations move. A stale claim is a false claim. | L5 sustain rule: claims lapse after 12 months without reassessment |
A residual observation: the first five gaps share one root cause — the original criteria implicitly assumed the bank builds its own AI and governs willing participants. Real estates are mostly bought, partly hidden, and overseen by humans who defer to the machine. The revised criteria govern the estate a bank actually has.
Twelve gaps from the industry landscape — classified, and dispositioned
An extensive review of ~35 published frameworks (CMMI AIM, MITRE AI MM, Microsoft RAI MM, Credo AI, OWASP AIMA, TCS 5A, Capgemini, Deloitte's banking index, FINOS AIGF, MAS Veritas, and others) surfaced twelve capabilities present somewhere in the landscape but absent from the AGMM. Each was classified by which integrity property of a maturity model it attacks — and that classification determines where the fix lives.
| Class | Integrity property attacked | Fix location |
|---|---|---|
| A · CONTENT | Construct validity — the grid misses real capability | In the model: domain and criteria changes. Validity failures are the ones an examiner exploits. |
| B · METHOD | Reliability — same organisation, different scores | This methodology paper: administration rules. Putting reliability machinery into the scored grid bloats it. |
| C · ECOSYSTEM | Adoptability — can't scale beyond its author | Roadmap, deliberately outside the model: benchmarks and certifications are products built on the model, versioned separately. |
Class A — included in the model (this revision)
| Gap (source) | Disposition |
|---|---|
| AI Security (OWASP AIMA, Google SAIF, CMMI AIM, FINOS) | New 9th domain — Security & Robustness (SR). The grid is now 27 cells. Security could not remain criteria inside Risk Management: it has its own people (red team), process (threat modelling), and technology (runtime detection), so it needs its own P/P/T cells or the floor rule cannot catch a security-specific weakness. |
| Value realization (MITRE, ServiceNow) | Criteria in Strategy & Operating Model, L4: governance ROI reported — incidents avoided, approval time, examination cost. Boards fund what shows return; a model that measures only downside invites defunding. |
| Governance velocity (Credo AI "Governing at Speed") | Criteria in SO, L4: gate cycle time and bypass rate tracked. Wired to the shadow-AI criteria deliberately — slow gates cause the bypass the model penalises elsewhere; the criteria now connect. |
| Culture & incentives (Microsoft RAI MM) | Criteria in SO People axis: L3 concern-raising channel exists; L4 raising concerns demonstrably safe, incentive alignment reviewed; L5 culture measured with survey evidence, not asserted. |
| Provider concentration risk (FSB, final report Oct 2026) | Criteria in SR Process, L4: concentration formally assessed. Placed in Security & Robustness because concentration is a resilience property, ahead of the FSB making it mandatory vocabulary. |
| Sustainability (RBI FREE-AI Sutra, EY, Wipro) | One criterion in SO Process, L5: model efficiency and energy footprint reported. Proportionate to current examiner attention; expands when regulation does. |
Class B — administration rules (Section 08)
Assessor protocol, system-level profiles, and proportionality — documented below rather than scored in the grid.
Class C — roadmap, deliberately out of model
| Item | Rationale for exclusion |
|---|---|
| Peer benchmark dataset | Built from anonymised assessment runs over time; seeded from published data (Deloitte: 13% of banks "leading"; McKinsey trust average 2.3/4; EDM: 31% advanced). A dataset is a product on the model, not part of it — and it becomes the moat nobody can copy. |
| ISO 42001 crosswalk | Mapping AGMM levels to certification readiness (L3 ≈ certifiable) ties the model to an external standard's revision cycle. Published as a separate crosswalk document. |
| Machine-readable schema | JSON export of the 27-cell grid for GRC ingestion (FINOS governance-as-code pattern). A tooling feature, not a model property. |
Reliability rules: who scores, at what altitude, and for which organisations
Reliability hardens with level, matching the validation chain in Section 03: L2–L3 may be self-assessed against the evidence criteria; L4 requires second-line validation of the grid; L5 requires external assurance. A facilitated assessment (trained facilitator, evidence sampled live) is the recommended default at every level — self-assessment without evidence sampling reliably inflates scores by one level. A full assessor certification protocol is specified separately as AGMM-M-02.
The 27-cell grid is an organisation-level instrument. High-risk systems additionally receive a system-level profile — the same P/P/T logic applied per system (as MAS Veritas does per use case). The rule connecting the two altitudes: an organisation's level claim is capped by its worst high-risk system profile. A bank at org-L3 running an L1 credit-scoring system is an L1 credit-scoring risk with good paperwork.
Criteria assuming portfolio scale relax for small institutions (small finance banks, NBFCs, fintechs with fewer than ~10 AI systems): the inventory may remain a spreadsheet at L3; dedicated capacity may be fractional; the registry requirement follows portfolio size, not level. What never relaxes: named ownership, evidence for every claim, and the gate tests. Proportionality scales the machinery, never the integrity rules.
The model in five statements
- 27-cell grid: 9 domains × People/Process/Tech, each cell scored 1–5 against evidence-named criteria.
- Domain level = floor of its three cells. Org level = floor across applicable domains, reported alongside frontier level and imbalance spread.
- Every claim ≥L3 requires a named evidence artifact; assessors sample.
- Gate tests per level: L2 = produce the inventory; L3 = cold audit; L4 = two stable measurement cycles; L5 = one self-initiated improvement loop.
- Imbalance diagnostics (shelfware / paper governance) as first-class outputs.
- The estate as it actually is: vendor and shadow AI in scope from L2; oversight, incident response, and fallback from L3; validation integrity hardening to external assurance at L5; and no level claim survives 12 months without reassessment.
The claim, verified against an extensive review of ~35 published frameworks: the AGMM is the first BFSI-specific AI governance maturity model that scores People, Process, and Technology as separate evidence-gated capability axes in every governance domain. Every word is load-bearing. CMMI AIM (July 2026) has evidence appraisals but is horizontal and does not score P/P/T per domain; TCS 5A names the dimensions but publishes no levels or gates; FINOS and the regulators are BFSI-native but publish catalogues and sound practices, not leveled models; Deloitte's banking index is a survey, not an assessable model. The imbalance diagnostics are unique outright. The 27-cell grid is the natural training deliverable: every workshop participant leaves with their grid.