This is an open working note. Values and definitions here are working values, they are published so the reasoning is inspectable, not because they are settled.
Status: open. Blocks normative release of the maturity model.
The question
The maturity gates, 90% ownership at Level 1, 80% cadence adherence at Level 2, 70% verification coverage at Level 4, and the rest, are working values. They were set by judgement about what ought to constitute a repeatable practice, not derived from observed distributions.
They are published anyway, because a model whose thresholds are hidden cannot be argued with. But they are labelled pre-normative and they will move.
The constraint that makes this hard
Calibration requires representative data across at least two quarters of recorded activity from multiple organisations. That data does not exist yet.
The available dataset is our own demonstration tenant, and it must not be used for this. A threshold tuned so that the demo lands on a favourable level is the single thing that would make the entire model indefensible, and it would be almost impossible for an outside reader to detect, because the fitted thresholds would look exactly like judged ones.
So the honest position is: the gates are unvalidated, and the temptation to validate them against the one dataset we control is being explicitly refused.
What calibration would actually involve
Distribution, not intuition. For each gate, the observed distribution across practices and organisations. A threshold sitting in a dense part of the distribution will produce unstable level assignments where small changes flip the result.
Separation. Does the gate distinguish organisations that a knowledgeable review panel would place differently under a published rubric? Reviewer opinion is a comparator, not unquestioned ground truth. A threshold that everyone passes or everyone fails carries no information regardless of where it sits.
Stability across cadence. A gate calibrated on monthly practices may behave differently on event-driven ones, where the population is smaller and more variable.
Asymmetry of error. Setting a gate too high delays recognition of genuine progress. Setting it too low recognises practices that are not repeatable. These are not equally bad, and the model should say which way it errs. Current inclination: err high, because a maturity claim that overstates is worse than one that lags.
Reliability and invariance. Do equivalent records produce the same result across reviewers, sectors, sizes and cadences? A gate that separates one cohort but changes meaning in another is not a common maturity threshold.
Sensitivity near the boundary. Publish how conclusions change when the gate, denominator, missing-data rule or observation window moves within a reasonable range. A stable model does not flip a large population on a trivial assumption.
What would change the answer
An initial pilot cohort of roughly eight to fifteen organisations, each with at least two full quarters of connected, reasonably complete execution history, would be enough to begin testing the gates. It is not a claim of statistical sufficiency. The usable sample is the number of independent practices and observation windows after stratifying for cadence, sector and organisation size, not the logo count.
Before looking at outcomes, the calibration plan must publish: inclusion and exclusion rules; the minimum completeness and source coverage; how repeated windows from the same organisation are handled; the stability test around each candidate threshold; the error direction the model prefers; and a holdout set that is not used to select the gates.
Repeated practices from one organisation are not independent observations. The analysis should account for clustering by organisation and practice, retain sector and cadence strata, and report uncertainty rather than presenting a single fitted percentage as truth. Threshold selection, exclusions and the primary tests should be preregistered before the holdout set is opened.
A normative release requires evidence that the selected gates discriminate, remain stable near their boundaries and do not simply reproduce the judgement of the people who chose them. If the available cohort cannot support that claim, the model remains pre-normative regardless of how many organisations are named.
Why this is published rather than kept internal
An assessment model that will not show its thresholds is asking to be trusted on authority. This one is asking to be argued with, and the open questions are the part most worth arguing about.
If you think a gate is in the wrong place, the argument is welcome and it will be recorded against the version.
Industry relationship
C2M2 supplies defined maturity indicator levels and practices; NIST CSF Tiers are explicitly not maturity levels. IO's gate percentages are proprietary working parameters, not values endorsed by DOE, NIST or ISO. Calibration can improve their empirical defensibility; it cannot turn them into an industry standard.