title: "Maturity You Can't Self-Assess" subtitle: "A security maturity model computed from execution evidence, not from an interview" category: "Methodology" reading_time: "8 min" author: "Jacob Revord" document_version: "v3.2" method_version: "v0.2, working draft" updated: "1 August 2026"

Original IO™ Documentation Developed by IO-HQ for the IO™ platform using the InfluenceOS™ methodology. This resource reflects the original IO™ operating perspective. Document version: v3.2 · Revision family: Third generation · Updated: 1 August 2026 © 2026 IO-HQ. IO™, InfluenceOS™, and named IO product features carrying the ™ symbol are trademarks of IO-HQ.
Maturity You Can't Self-Assess
Many widely used security maturity assessments still depend, at their core, on an interview.
Someone arrives with a questionnaire. A team answers it. The answers are scored, banded, and rendered as a level. The organization is told it is a 3.
The problem is not that the people answering are dishonest. It is that they are answering from memory, under time pressure, about a program they are personally invested in, and being asked whether they have a process rather than whether the process actually ran. Nobody in that room is lying. Everybody in that room is guessing.
Then the number ages. It described a moment, and the moment passed. By the time it reaches a board pack, it is a claim about a version of the organization that no longer exists.
There is a different way to do this, and it starts by measuring something else entirely.
Four things wrong with how this is done today
It is self-reported. The people describing the program are the people accountable for it. Even with complete integrity, the answers reflect intention more than practice, what the process is supposed to be, not what happened last Tuesday.
It asks the wrong question. Do you have a process for access reviews? Yes. Did one run this quarter? Did it produce revocations, or sign-offs? Was anything checked afterward? A questionnaire records the existence of a process. It cannot see whether the process executes.
It is a photograph. An annual assessment describes one week and is then quoted for twelve months. A program can improve substantially or quietly collapse without the number moving at all.
And the results are difficult to compare, often even to themselves. When bands are set locally, the assessor changes, or last year's answers are copied forward, the level stops meaning anything stable enough to trend.
The result is a number everybody reports and nobody trusts. Security leaders know their maturity rating is soft. They present it anyway, because the alternative is presenting nothing.
Maturity is not a score
The most common attempt at a fix is to compute maturity from risk scores. Take the numbers, band them, call the top band "optimized."
That produces a redundant model. If the level is a restatement of the score, it tells you nothing the score did not, and the word "maturity" is doing decorative work.
Performance and maturity are different questions:
Domain performance asks how we are doing right now.
Maturity asks whether it is repeatable.
Security practices sit inside IO domains. Domain performance and maturity are reported separately: a domain performance score is not a maturity result, and maturity levels are never calculated by weighting or averaging domain scores. Enterprise maturity is reported through assured level, distribution, evidence coverage and binding constraints.
An organization can perform well and be immature. It happens constantly. A capable team, a quiet quarter, a few strong individuals holding things together, good numbers, no repeatable process, and one resignation away from a very different picture.
That organization is not mature. Its position is fragile. A maturity model that cannot distinguish those two states is not worth running.
Measure the process, not the outcome
If maturity is about repeatability, the evidence is not in the results. It is in how the work was done.
Much of that evidence already exists inside every security program, and almost nobody reads it:
Did the work have an accountable owner? Not only a team or function, but a person or role with clear responsibility for its disposition.
Did it get finished? Not started, not scheduled. Finished, with a date attached.
Did it keep happening? Or did six months of activity compress into the period after an incident, audit, or customer escalation? Continuity is measured against the expected operating cadence of the practice, not an arbitrary clock.
Was it driven by something written down? Does a stated rule exist, does it carry a threshold precise enough to act on, and do findings reference it, or does everyone simply know how things are done here?
Was anything measured afterward? Not the result. Whether a measurement was taken at all.
Did it hold? Or did the same finding return within the defined recurrence window under a different name?
Six questions. In most cases, every one is answerable from records the organization already generates. None requires the organization to rate itself.
The dimension that decides everything
Of those six, one carries more weight than the rest: was anything measured afterward.
Not the pass rate. The coverage, the share of completed work that was measured at all.
Most security programs have no idea. Work is raised, assigned, executed, and closed; the closing is the end of it. Whether the intervention changed anything is not recorded, because nobody requires it and the answer is sometimes uncomfortable.
An organization that completes work without measuring whether it worked cannot demonstrate that its outcomes are empirically managed. It is executing on faith. However good its numbers look, it has no mechanism for discovering that a control has stopped working, a program has stopped landing, or an intervention it has run forty times has never moved the thing it was meant to move.
This is the line most organizations sit on. Owners are named, policy exists, work gets done, and the loop never closes.
Why this hasn't been done before
The method is not complicated. The reason it is not already standard is structural: maturity can be computed reliably from execution evidence only when that evidence can be connected across the whole execution chain.
Few platforms are designed to do that.
Governance and compliance platforms hold the requirement. They know what the control is meant to be and when it was last attested. They do not see the work, and they cannot tell you whether anything changed as a result.
Ticketing and workflow systems hold the task. They know it was assigned and closed. They have no idea what it was supposed to change, or whether it did, a ticket closes when someone clicks close.
Detection and monitoring platforms hold the signal. They see the problem arrive. They do not see the response, and they cannot connect the two.
Assessment tools hold the opinion. They have no execution record at all.
Most tools own one link. The evidence needed for maturity spans the whole chain: the signal that raised the work, the accountable owner, the rule that justified it, the date it was dispositioned, the baseline, and the measurement showing what changed afterward.
What has to be true to compute it
Four things, designed in rather than bolted on:
The baseline is captured before the work starts, wherever possible. Measure only afterward and there is nothing reliable to compare against, you can report activity, but not change. For urgent incident response or newly connected sources, the first reliable observation may serve as the baseline. That exception reduces confidence and may leave the outcome inconclusive.
Rules are structured, not simply written. A policy stating access is revoked "promptly" cannot be evaluated. One stating five working days can, and every finding that breaches it can cite the clause it breached. Policy must be represented as data with thresholds the platform reads, not only as a document filed alongside.
Failures are recorded. An intervention that ran and did not work has to be storable as a fact. Any system that can only record success will always report success, and its maturity finding will be worthless.
And the record is continuous. Not a snapshot taken by a visitor. Evidence accumulates as the work happens, so the level moves when the organization moves, not twelve months later.
Where IO sits
IO is built to connect that chain, which is what makes maturity from execution evidence practical.
Signals arrive from systems already in place. Work is raised with an accountable owner, a stated rule behind it, and a measurement plan agreed before execution. When it completes, the original signal is measured again, and the outcome is recorded as verified, unverified, regressed, or inconclusive.
That last distinction matters. A platform willing to store failure, regression, and uncertainty is a platform whose maturity finding can be trusted. A post-intervention measurement shows what changed after the intervention; it does not prove causation unless the measurement design supports that conclusion.
None of this requires the organization to answer a questionnaire, and none of it can be improved by answering one differently.
The levels
Each level represents a capability cluster, and each depends on the one before it. You cannot verify work that has no accountable owner. You cannot govern work that does not continue.
0, Unmanaged
There is evidence, and it shows no repeatable process.
Work happens. It is largely unowned, discontinuous, unmeasured, and driven by nothing written down.
This is a finding, not an insult, and it must be sayable. A model whose floor is "functioning but reactive" cannot describe organizations that genuinely have nothing, and those organizations exist.
1, Accountable
Work has an accountable owner and reaches a recorded disposition when something forces action.
Ownership is established. Work is completed, formally accepted, or otherwise dispositioned with evidence. The trigger is still external: an incident, an audit, a customer question.
2, Sustained
Work continues across its expected operating cadence, not only when something is on fire.
The program operates across consecutive periods rather than in bursts. This is the first point at which execution can be planned, compared, and expected to recur.
3, Governed
What should be true is written down, precise, and referenced by findings.
Rules exist, carry thresholds specific enough to act on, and are cited when something breaches them. Exceptions have an owner, rationale, approval, and expiry or review date.
4, Verified
Completed work is verified against the thing it was supposed to change.
A valid baseline or first reliable observation is recorded, an expected outcome is stated, and a post-intervention measurement is taken inside a defined window. Negative, neutral, regressed, and inconclusive outcomes remain in the record.
Few programs consistently get here. Measuring afterward is optional, invisible, and occasionally uncomfortable, so it rarely becomes part of normal execution unless the system requires it.
5, Adaptive
Measurement changes behavior.
Interventions that demonstrably do not work are modified or retired rather than repeated. Recurring findings trigger root-cause action. The organization still fails, but evidence changes what it does next.
Why a level is all-or-nothing
Levels are not averaged. Every published gate at a level must be met, and every gate below it must still hold. This does not require every work item to be perfect; it requires each defined condition to meet its published threshold.
The dimensions are not comparable. Ownership at 80% and verification coverage at 70% are not the same quantity. Averaging them produces a number that cannot be explained when someone asks what it is made of.
A weighted score is a grade; all-or-nothing is a diagnosis. "Capped at Accountable because verification coverage is 22%" tells you what to do on Monday. "You scored 61" does not.
And weighted models get gamed, not maliciously, just rationally. You lift the cheap dimension to shift the average and leave the expensive one alone. All-or-nothing points you at the condition actually blocking you.
Every result names its binding constraint: the practice, the unmet condition, the published threshold, the current measurement, and the remaining gap.
Thresholds are fixed, published, and versioned. They are not set by the organization being measured. Internal targets may sit alongside the standard, but they never replace it. Threshold changes apply prospectively under a named model version, so movement is never created by silently moving the bar.
"We can't assess this" is not a low score
There is an important difference between an organization with no repeatable process and one whose evidence has not yet reached us.
The first is a finding. The second is a gap in instrumentation. Reporting the second as a low level is not merely inaccurate; it attributes a process weakness the available evidence cannot support.
So not yet assessable is held separately from every level, never rendered on the same scale, and always carries its reason: a critical source is not connected, evidence coverage is below the published minimum, the observation window is incomplete, or the sample is insufficient.
The count of areas that cannot be assessed is itself a finding, and an honest one. "Four of eight domains cannot be assessed because evidence is not flowing from them" is more useful than an amber tile implying partial maturity.
When performance and maturity disagree
This is where the model earns its place. Performance indicators and the maturity finding are supposed to be able to diverge.
| Performance and maturity | Executive reading |
|---|---|
| Strong performance, low maturity | Strong but fragile. The result is real, but not yet reproducible, and may depend on a few individuals or favorable conditions. |
| Weak performance, high maturity | Underperforming but controlled. The organization knows where it stands, measures change, and has a credible improvement trajectory. |
| Strong performance, high maturity | Demonstrably strong. Results are good and the process producing them is repeatable. |
| Weak performance, low maturity | Uncontrolled exposure. Establish ownership, execution discipline, and evidence before chasing isolated performance gains. |
The second row is the one boards consistently misread. A security program with weak current performance and a working measurement loop can be in better shape than one with excellent current numbers and no reliable explanation for them.
On established frameworks
This model is informed by the maturity progression familiar from CMMI and by the structure of C2M2 Version 2.1, the US Department of Energy's Cybersecurity Capability Maturity Model. Where a crosswalk is published, it is anchored to C2M2, which is publicly available and built for industrial and operational technology environments.
It is not an appraisal, and it does not claim alignment or compliance with any of them. Compliance is an assessor's determination, made by an assessor. This is a measurement, made from evidence.
The relationship is complementary rather than competitive. A framework describes what good looks like. This describes what you are actually doing and where that lands you relative to it. If you have already invested in a framework, that investment is what makes the mapping worth having.
We reference practice identifiers. We do not reproduce framework text.
What this does not cover
A model that claims to cover everything is not credible.
This one says nothing about business continuity and disaster recovery planning, or personnel vetting. Physical Security is an IO domain with observable performance and verification measures, but a Physical Security maturity finding is not produced: the model requires defined material practices, attributable execution lineage, representative evidence coverage, a completed assessment window and calibrated gates, and until those are satisfied IO reports domain performance and verification outcomes instead. It says nothing about any area producing no execution record, since the entire method depends on one.
Those are real parts of a security program and they need assessing by other means. Naming that is what makes the covered parts believable.
The short version
Nobody should be asked to rate their own security maturity, and nobody should have to wait for an annual assessment to discover that it changed.
The evidence showing whether a security program is repeatable is already being generated through normal execution. Ownership, disposition, continuity, governance, verification, and adaptation are all there. They are rarely read as one chain, because most systems hold only one part of the path from the signal that raised the work to the measurement showing whether it landed.
We did not ask. We watched.
This is the security maturity model IO is built around. The full specification, including threshold gates, evidence confidence bands, scope rules, and denominators, is published separately.
The demo runs on a fictional industrial company with two quarters of recorded security activity. No installation. Pick a seat and walk in.