Last updated:
A governance maturity model is a staged scale for assessing how developed an organization's AI governance is, typically running from ad hoc practice through to measured and continuously improved. It is used to benchmark a starting position, set a realistic target and sequence investment, not to certify compliance.
What do the stages usually describe?
Most models follow the five levels that CMMI established. At the first level activity is ad hoc and depends on individuals. At the second it is repeatable within teams but inconsistent across them. At the third it is defined, with documented policy and roles that apply organization-wide. At the fourth it is measured, with metrics on coverage, cycle time and control effectiveness. At the fifth it is optimizing, with the program adjusting itself based on what the measurements show. The labels vary between frameworks; the progression from personal practice to institutional practice to measured practice does not.
What is being measured?
Maturity is assessed across several dimensions at once, and organizations are rarely at the same level in all of them. Policy asks whether standards exist and are current. Inventory asks what proportion of the AI estate is registered, which is where most programs are weakest. Process asks whether intake, risk assessment and approval happen consistently or by exception. Roles ask whether accountability is assigned and understood. Evidence asks whether controls produce records that survive audit. Monitoring asks whether anything is checked after deployment. A program with strong policy and weak inventory is common and scores worse than its documentation suggests.
How does it relate to ISO/IEC 42001 and the NIST AI RMF?
They answer different questions. ISO/IEC 42001 is a certifiable management system standard: a firm either meets its requirements or does not. The NIST AI Risk Management Framework is voluntary guidance organized around four functions, Govern, Map, Measure and Manage, and does not define maturity levels. A maturity model sits alongside both as a planning instrument, describing how far along a firm is and what to build next. Using a maturity score as evidence of conformity to either standard is a category error, though maturity assessment is a reasonable way to plan the route to certification.
What are the common failure modes?
Self-assessment inflation is the main one, since teams score their own function and read intent as achievement. Assessing at the wrong altitude is another, where a firm scores the whole organization and hides the fact that one business unit is at level four and another at level one. The third is treating the model as the goal: level five is not the right target for every organization, and a firm with a small, low-risk AI estate reaching level three with good coverage is in better shape than one chasing level five with half its systems unregistered.
How should it be used?
Assess against evidence, not opinion, by sampling real systems and checking whether the control record exists. Score by dimension so the weak axis is visible. Set the target level deliberately, based on regulatory exposure and the size of the AI estate. Reassess on a fixed cycle so movement is measurable, and use the gap between dimensions to sequence work, since lifting inventory coverage usually unblocks more than refining policy does.
Real world example:
A retail group assesses its AI governance across five dimensions. Policy scores level three, since a board-approved AI standard exists and is current. Inventory scores level one, because a sampling exercise finds 40 registered systems against roughly 110 in actual use, most of the gap being AI features inside existing SaaS tools. Process scores level two, as intake works in the data science team and nowhere else. The group sets a twelve-month target of level three across all five dimensions instead of level four anywhere, and puts its first quarter into discovery and registration, since the inventory gap makes every other score unreliable. Reassessment is scheduled at six and twelve months against the same evidence-based method.




