Methodology Paper · AGMM-M-01
AGMM Methodology  ·  First Edition 2026

Scoring Architecture and Level Criteria for the AI Governance Maturity Model

Why maturity is scored as a 27-cell grid — the gating rules, the evidence criteria per level, and the imbalance diagnostics that make a level claim verifiable rather than self-reported.

Derivation. This methodology was derived from first principles — CMMI's institutionalization logic, DCAM's evidence-based scoring, and the operative regulatory anchors (EU AI Act Articles 4, 9 and 17; SR 11-7; RBI Advisory 6; NIST AI RMF GOVERN) — rather than by extending prior drafts. Three places where the analysis reversed earlier conclusions are flagged as divergences in the text.
01 · Gating Rules

The core design decision: what "achieving a level" means

Three candidate gating rules were evaluated:

RuleHow it worksVerdict
Average across P/P/T Level = mean of the three dimension scores REJECTA 5-tech / 1-people domain averages to 3 — but that organisation cannot operate its own tooling. Averages hide exactly the failure the model exists to catch.
Minimum (floor) Level = weakest dimension CORRECT — DOMAIN LEVELThe weakest dimension is the operating ceiling: process without people is paper; tech without process is shelfware.
Profile (continuous) No single level; report the shape CORRECT — ORG LEVEL, ALONGSIDE A FLOORSingle-number maturity invites gaming and hides imbalance.
Recommendation

Floor at domain level; floor + profile at organisation level. Report both the Achieved Level (floor across applicable domains) and the Frontier Level (best domain), and the spread between them. A bank at floor 2 / frontier 4 has an imbalance problem, not a maturity problem — and the remediation is entirely different.

02 · Binding Constraints

The binding constraint shifts by transition

Each level transition is gated by a different dimension. This matters because it tells a bank where the next dollar goes.

TransitionBinding constraintWhy
L1 → L2 People Nothing exists until someone owns it. Policy can be written in weeks once an accountable owner with a mandate exists. Tech is irrelevant here.
L2 → L3 Process (tech secondary) Defined-to-operating is execution discipline: gates enforced, exceptions logged, cadences actually run. Minimum tooling needed, but spreadsheet-to-registry is a purchase; enforced gates are a culture change.
L3 → L4 Technology (specialist people secondary) Measurement at portfolio scale cannot be manual. Evidence must be generated automatically or it won't survive two cycles. This is the first transition where tooling is the hard gate.
L4 → L5 People + Process By L4 the tech mostly exists. Optimised requires depth: a CoE that survives staff departure, and a closed improvement loop that fires without external prompting.
Divergence 1 — of 3

Earlier analysis implied technology becomes a requirement only at L3. On re-analysis, L2 already has one hard tech gate — a complete AI system inventory. The EU AI Act effectively mandates a registry; SR 11-7 assumes a model inventory. A spreadsheet qualifies, but if the organisation cannot produce a complete list of AI systems in production, it is Level 1 regardless of how good its policy document is. That is the single most testable L2 criterion.

03 · Level Criteria

What must be true, per dimension, with evidence

Each cell states the criterion and the evidence artifact an assessor asks for. A claim without a named artifact scores the level below.

LEVEL 1

Ad Hoc

DimensionCriterionEvidence
PEOPLENo dedicated governance roles required. AI governance knowledge lives in individuals, not the institution.None — the absence of appointment records is itself the state
PROCESSNo documented governance processes. Decisions made case by case with no repeatable procedure.None — no policy, no RACI, no intake procedure exists
TECHNo governance-specific tooling deployed. General IT infrastructure only.None — no inventory, registry, or evidence repository
L1 Gate Test

None. Level 1 requires nothing — it describes the absence of governance capability. Every organisation deploying AI is at least Level 1 by default.

LEVEL 2

Defined

DimensionCriterionEvidence
PEOPLENamed accountable owner (fractional OK) with an identified executive sponsor. Board briefed at least once. Baseline AI literacy in risk/compliance (EU AI Act Art. 4 anchor).Appointment letter or charter; board minutes; training records
PROCESSBoard-approved AI governance policy. Documented RACI. Defined intake/registration procedure for new AI systems — covering vendor-supplied AI, not only in-house builds. Acceptable-use policy for employee generative-AI use, communicated to all staff.Approved policy with version history; RACI document; acceptable-use policy with distribution record
TECHComplete AI system inventory — spreadsheet acceptable — including vendor-supplied and vendor-embedded AI (fraud engines, scoring services, productivity copilots). Document repository for governance artifacts.The inventory itself, reconciled against IT asset records and the procurement register
L2 Gate Test

Produce a complete inventory of every AI system in production — including vendor-embedded AI — reconciled against IT asset records and the procurement register. No inventory, no Level 2 — regardless of policy quality.

LEVEL 3

Operationalised

DimensionCriterionEvidence
PEOPLEFunded dedicated capacity (budget line, not volunteer time). Named owner per applicable domain. At least one person trained per applicable framework. Second line able to challenge, not just receive. Human overseers of high-risk systems trained to a defined competence — including authority and ability to override (EU AI Act Art. 14 anchor).Budget allocation; org chart; framework training certificates; a recorded second-line challenge; overseer competence records
PROCESSDeployment gate enforced — with an exception log (a gate with zero exceptions ever recorded is a gate nobody uses). Pre-deployment risk assessment on every system. Escalation path exercised at least once (drill or real). Review cadences producing minutes, not just calendar entries. AI incident response plan with severity taxonomy and regulatory reporting obligations mapped (EU AI Act Art. 73 serious-incident anchor). Vendor AI due diligence in the intake procedure. Fallback or kill-switch defined for every high-risk system.Gate records including exceptions; risk assessments; escalation record; minutes; incident response plan; vendor due-diligence records; documented fallback procedures
TECHModel registry with versions. Inventory in a system (not spreadsheet) once portfolio exceeds roughly 20 systems. Monitoring on high-risk systems. Maintained framework crosswalk. Central evidence repository. Shadow-AI discovery mechanism — network, SaaS, or procurement scanning that tests inventory completeness rather than assuming it.Registry export; monitoring dashboards; crosswalk document; shadow-AI scan results
L3 Gate Test

An examiner walks in cold and audits without the bank preparing anything. If preparation is needed, the organisation is L2 with good marketing.

LEVEL 4

Measured

DimensionCriterionEvidence
PEOPLESpecialists: control testing, fairness metrics, LLM-era validation. Independent challenge function (second line / internal audit) with AI competence. Board committee that interrogates quantified reporting — recorded challenge in minutes, not receipt of slides. The AGMM grid itself independently validated by second line — self-assessed scores don't count at this level.Role profiles; IA reports; committee minutes showing questions asked; second-line sign-off on the maturity grid
PROCESSControl testing calendar executed for two consecutive cycles with stable metric definitions. Residual formula applied consistently. Untested controls scored zero (the integrity rule). Appetite thresholds defined; breaches escalated on record. Regulatory change process with SLA. Oversight effectiveness measured — human override rates tracked to detect rubber-stamp oversight and automation bias. Fallback procedures tested, not just documented (DORA operational-resilience anchor). Vendor control assurance obtained for material vendor AI.Two cycles of test results; scoring methodology document; breach escalation records; override-rate reports; failover test results; vendor assurance reports
TECHAutomated testing for top-residual controls. Scoring engine with audit trail. Dashboards generated from data, not assembled in slides. Drift detection live. Lineage on high-risk training data. Human-override telemetry captured per high-risk system.System-generated reports; pipeline configurations; override telemetry samples
L4 Gate Test

Show two consecutive control-testing cycles with unchanged metric definitions. One snapshot is a point; measurement is a trend.

Divergence 2 — of 3

One measurement cycle is not "Measured." A single snapshot is a point; measurement is a trend. Two consecutive cycles with unchanged metric definitions is the honest gate — it also catches organisations that redefine metrics each quarter to look better.

LEVEL 5

Optimised

DimensionCriterionEvidence
PEOPLEGovernance CoE. Competency framework with assessed (not self-declared) skills. Succession for key roles — governance survives any single departure. Agentic-AI specialists embedded. Level claim externally assured — an independent third party has examined the grid and its evidence.Competency assessments; succession plan; the organisation still functioning after a key exit; external assurance report
PROCESSClosed loop: incident / regulatory change / benchmark result → traceable framework update → measured effectiveness shift, without external prompting. Simulated regulatory exam annually. Board reviews and revises appetite annually (a never-revised appetite statement is decoration). Sustain rule enforced: any level claim lapses after 12 months without reassessment — maturity decays; the model must too.One complete traced loop; sim-exam report; appetite revision history; reassessment calendar with completed cycles
TECHContinuous monitoring across domains. Multi-framework reconciliation automated (one risk entry satisfies RBI + EU AI Act + NIST simultaneously). Tamper-evident audit trails. Per-step agentic observability.Live dashboards; reconciliation mapping; agentic trace samples
L5 Gate Test

Show one complete loop — incident → catalogue update → control change → measured effectiveness improvement — that nobody outside the system initiated.

04 · Imbalance Diagnostics

Dimension-imbalance diagnostics — the most useful output

Healthy ascent keeps People ≥ Process ≥ Tech, because people build process and process specifies tech. Two inversion patterns are reliable failure signatures:

PatternNameWhat it predicts
Tech ≥ Process + 2 Shelfware Bought a governance platform; nobody runs it. Tooling spend wasted; false confidence at board level.
Process ≥ People + 2 Paper governance Beautiful documentation nobody can execute. Fails the first real incident or examination.

These should surface automatically in the assessment output. They are more actionable than the level itself: "you're Level 3" prompts a shrug; "your Risk Management domain is a shelfware pattern — Tech 4, Process 2" prompts a budget reallocation.

05 · Applicability

The applicability rule

Divergence 3 — of 3

Earlier logic made organisation level the minimum across all domains. That breaks on Agentic AI: a bank with zero agents deployed would be dragged to Level 1 overall by a domain it doesn't operate. Wrong incentive — it punishes prudence.

Rule

A domain may be marked N/A only if there are no deployments AND none on a 12-month roadmap. N/A domains are excluded from the floor but flagged in the profile. The moment an agent pilot is approved, Agentic AI Governance becomes applicable at whatever level the organisation actually has — usually Level 1, which is exactly the honest signal you want before the pilot ships, not after.

06 · Critical Review

Seven gaps found on adversarial review — and where they were closed

The level criteria were reviewed from the examiner's chair: what would an RBI, ECB, or MAS examiner ask that the model could not answer? Seven gaps surfaced. Each is now embedded in the criteria above, marked in bold.

GapWhy it mattersClosed at
Vendor & third-party AI Most bank AI is vendor-supplied — fraud engines, scoring services, copilots. A governance model that only sees in-house builds misses the majority of the estate. EU AI Act deployer obligations and RBI outsourcing guidelines both apply. This is the first question an examiner asks. L2 inventory & intake; L3 due diligence; L4 vendor assurance
Shadow AI An inventory is only as good as its completeness — and completeness must be tested, not assumed. Unsanctioned employee GenAI use is the largest unmanaged surface in most institutions. L2 acceptable-use policy; L3 discovery mechanism
AI incident response The original L3 had escalation but no incident plan. EU AI Act Art. 73 mandates serious-incident reporting with deadlines; without a severity taxonomy mapped to reporting obligations, a bank discovers its duty mid-incident. L3 incident response plan with regulatory reporting map
Human oversight effectiveness EU AI Act Art. 14 requires oversight with real competence and authority to override. Oversight that never overrides is rubber-stamping — automation bias in production. Override rates are the measurable signal. L3 overseer competence; L4 override-rate measurement and telemetry
Assessment integrity A maturity model scored by the people it measures converges on flattery. The grid itself needs a validation chain that hardens with level. L4 second-line validation of the grid; L5 external assurance
Resilience & fallback What happens when the model fails at 2 a.m.? DORA makes operational resilience a legal requirement for EU financial entities. Kill-switches must exist by L3 and be tested by L4 — a documented failover that has never run is a hope, not a control. L3 fallback defined; L4 fallback tested
Level decay The original model treated maturity as a ratchet — once claimed, held forever. Teams change, portfolios grow, regulations move. A stale claim is a false claim. L5 sustain rule: claims lapse after 12 months without reassessment

A residual observation: the first five gaps share one root cause — the original criteria implicitly assumed the bank builds its own AI and governs willing participants. Real estates are mostly bought, partly hidden, and overseen by humans who defer to the machine. The revised criteria govern the estate a bank actually has.

07 · Research-Driven Revision

Twelve gaps from the industry landscape — classified, and dispositioned

An extensive review of ~35 published frameworks (CMMI AIM, MITRE AI MM, Microsoft RAI MM, Credo AI, OWASP AIMA, TCS 5A, Capgemini, Deloitte's banking index, FINOS AIGF, MAS Veritas, and others) surfaced twelve capabilities present somewhere in the landscape but absent from the AGMM. Each was classified by which integrity property of a maturity model it attacks — and that classification determines where the fix lives.

ClassIntegrity property attackedFix location
A · CONTENTConstruct validity — the grid misses real capabilityIn the model: domain and criteria changes. Validity failures are the ones an examiner exploits.
B · METHODReliability — same organisation, different scoresThis methodology paper: administration rules. Putting reliability machinery into the scored grid bloats it.
C · ECOSYSTEMAdoptability — can't scale beyond its authorRoadmap, deliberately outside the model: benchmarks and certifications are products built on the model, versioned separately.

Class A — included in the model (this revision)

Gap (source)Disposition
AI Security (OWASP AIMA, Google SAIF, CMMI AIM, FINOS)New 9th domain — Security & Robustness (SR). The grid is now 27 cells. Security could not remain criteria inside Risk Management: it has its own people (red team), process (threat modelling), and technology (runtime detection), so it needs its own P/P/T cells or the floor rule cannot catch a security-specific weakness.
Value realization (MITRE, ServiceNow)Criteria in Strategy & Operating Model, L4: governance ROI reported — incidents avoided, approval time, examination cost. Boards fund what shows return; a model that measures only downside invites defunding.
Governance velocity (Credo AI "Governing at Speed")Criteria in SO, L4: gate cycle time and bypass rate tracked. Wired to the shadow-AI criteria deliberately — slow gates cause the bypass the model penalises elsewhere; the criteria now connect.
Culture & incentives (Microsoft RAI MM)Criteria in SO People axis: L3 concern-raising channel exists; L4 raising concerns demonstrably safe, incentive alignment reviewed; L5 culture measured with survey evidence, not asserted.
Provider concentration risk (FSB, final report Oct 2026)Criteria in SR Process, L4: concentration formally assessed. Placed in Security & Robustness because concentration is a resilience property, ahead of the FSB making it mandatory vocabulary.
Sustainability (RBI FREE-AI Sutra, EY, Wipro)One criterion in SO Process, L5: model efficiency and energy footprint reported. Proportionate to current examiner attention; expands when regulation does.

Class B — administration rules (Section 08)

Assessor protocol, system-level profiles, and proportionality — documented below rather than scored in the grid.

Class C — roadmap, deliberately out of model

ItemRationale for exclusion
Peer benchmark datasetBuilt from anonymised assessment runs over time; seeded from published data (Deloitte: 13% of banks "leading"; McKinsey trust average 2.3/4; EDM: 31% advanced). A dataset is a product on the model, not part of it — and it becomes the moat nobody can copy.
ISO 42001 crosswalkMapping AGMM levels to certification readiness (L3 ≈ certifiable) ties the model to an external standard's revision cycle. Published as a separate crosswalk document.
Machine-readable schemaJSON export of the 27-cell grid for GRC ingestion (FINOS governance-as-code pattern). A tooling feature, not a model property.
08 · Administration

Reliability rules: who scores, at what altitude, and for which organisations

Assessor protocol

Reliability hardens with level, matching the validation chain in Section 03: L2–L3 may be self-assessed against the evidence criteria; L4 requires second-line validation of the grid; L5 requires external assurance. A facilitated assessment (trained facilitator, evidence sampled live) is the recommended default at every level — self-assessment without evidence sampling reliably inflates scores by one level. A full assessor certification protocol is specified separately as AGMM-M-02.

System-level profiles

The 27-cell grid is an organisation-level instrument. High-risk systems additionally receive a system-level profile — the same P/P/T logic applied per system (as MAS Veritas does per use case). The rule connecting the two altitudes: an organisation's level claim is capped by its worst high-risk system profile. A bank at org-L3 running an L1 credit-scoring system is an L1 credit-scoring risk with good paperwork.

Proportionality

Criteria assuming portfolio scale relax for small institutions (small finance banks, NBFCs, fintechs with fewer than ~10 AI systems): the inventory may remain a spreadsheet at L3; dedicated capacity may be fractional; the registry requirement follows portfolio size, not level. What never relaxes: named ownership, evidence for every claim, and the gate tests. Proportionality scales the machinery, never the integrity rules.


Summary

The model in five statements

  • 27-cell grid: 9 domains × People/Process/Tech, each cell scored 1–5 against evidence-named criteria.
  • Domain level = floor of its three cells. Org level = floor across applicable domains, reported alongside frontier level and imbalance spread.
  • Every claim ≥L3 requires a named evidence artifact; assessors sample.
  • Gate tests per level: L2 = produce the inventory; L3 = cold audit; L4 = two stable measurement cycles; L5 = one self-initiated improvement loop.
  • Imbalance diagnostics (shelfware / paper governance) as first-class outputs.
  • The estate as it actually is: vendor and shadow AI in scope from L2; oversight, incident response, and fallback from L3; validation integrity hardening to external assurance at L5; and no level claim survives 12 months without reassessment.

The claim, verified against an extensive review of ~35 published frameworks: the AGMM is the first BFSI-specific AI governance maturity model that scores People, Process, and Technology as separate evidence-gated capability axes in every governance domain. Every word is load-bearing. CMMI AIM (July 2026) has evidence appraisals but is horizontal and does not score P/P/T per domain; TCS 5A names the dimensions but publishes no levels or gates; FINOS and the regulators are BFSI-native but publish catalogues and sound practices, not leveled models; Deloitte's banking index is a survey, not an assessable model. The imbalance diagnostics are unique outright. The 27-cell grid is the natural training deliverable: every workshop participant leaves with their grid.