This method is implemented through the IO Verify™ operating loop.
Most security programmes can tell you what they did last quarter. Very few can tell you whether it worked. This document describes the mechanism that answers the second question, and, more importantly, the constraints that stop it answering it dishonestly.
Industry grounding
Verification is not an IO invention. NIST and ISO methods already expect monitoring, measurement, assessment and continual improvement. NIST CSF 2.0 expresses cybersecurity outcomes; NIST SP 800-53 includes control assessment and continuous-monitoring concepts; ISO/IEC 27001 requires performance evaluation and improvement; NIST AI RMF separates Measure from Manage.
The industry norm, however, leaves organisations substantial freedom in what counts as evidence and how effectiveness is represented. IO adds an opinionated operating rule: define the measure before execution, capture the baseline at the accountability event, re-measure from the same source, preserve unsuccessful results, and keep coverage separate from pass rate. This is a product method, not a claim that the cited frameworks require IO's exact sequence.
It also draws a hard line between three related activities:
- Monitoring observes conditions over time.
- Assessment evaluates whether a control or requirement is satisfied.
- IO verification tests whether a named intervention was followed by the declared, measurable change.
The third can support the first two. It does not replace an assessor's judgement and does not prove causation without a stronger evaluation design.
Primary references: NIST CSF 2.0 (opens in a new tab) · NIST SP 800-53 Rev. 5 (opens in a new tab) · NIST AI RMF 1.0 (opens in a new tab) · ISO/IEC 27001 (opens in a new tab)
1. The measure is defined on the playbook, not on the outcome
A verification measure is part of a playbook's definition, written when the playbook is authored rather than when the work completes. It names four things:
- The signal the intervention is meant to change
- The measure and its unit
- The population it is computed over
- The direction that would count as improvement
A playbook without all four can prescribe work. It cannot verify it, and it is marked that way rather than quietly producing an unverifiable task.
The product sometimes calls this package the verification criterion. The terms are not interchangeable: the criterion is the complete rule, measure, unit, population, direction, noise threshold, source and re-measurement window. The measure is the number computed under that rule.
2. The baseline is captured at acceptance, not at completion
This is the single most important constraint in the method, and it is the one most commonly got wrong.
If the baseline is read when the work finishes, the person doing the work has already influenced the number, and the comparison is between a post-intervention observation and a reconstruction of what came before. Reconstructions drift toward whatever makes the result readable.
So the measure is taken at the moment a named person accepts the action, before any of the work starts, from the same source that will be read again afterwards. Acceptance and baseline capture are one event. Work that was never formally accepted has no baseline and is not verifiable, which is a finding about the operating model, not a gap in the instrument.
In this specification, acceptance is the approval event. Product copy should not imply that approval and acceptance are two separate opportunities to choose a baseline. There is one accountable event, one timestamp and one baseline.
There is one narrow exception for urgent response or a newly connected source: the first reliable observation may establish a baseline for future work. It cannot retrospectively verify the intervention that already occurred. That record is normally inconclusive, carries reduced confidence and names why no pre-intervention baseline existed.
3. Four results, and one state that is not a result
| Result | Meaning |
|---|---|
| Verified | Re-measured, moved in the declared direction, beyond the noise threshold for that measure |
| Unverified | Re-measured, did not move materially. The work ran and could not be shown to have worked. |
| Regressed | Re-measured, moved against the declared direction |
| Inconclusive | Re-measured, but the comparison cannot be relied on, the population changed, the source gapped, or the window was disturbed |
Pending is not a result. Work awaiting its re-measurement date has no outcome yet, and it is held apart from the four above with its due date attached. Reporting pending inside the result set would convert an absence into a measurement, which is the failure this whole method exists to prevent.
A partially executed playbook is not eligible for a clean verified, unverified or regressed result. If re-measurement still occurs, the result is inconclusive and the incomplete execution is the stated reason.
The distinction between unverified and inconclusive is worth holding precisely. Unverified is a finding about the intervention: it ran, it was measured, nothing moved. Inconclusive is a finding about the measurement: it cannot be trusted enough to say. Collapsing the two flatters the programme, because it converts measurement failures into intervention neutrality.
4. Coverage is not the pass rate
Two numbers get confused constantly, and only one of them describes the programme's honesty.
- Verification coverage, the share of completed work that was measured at all
- Verification pass rate, of the work that was measured, the share that came back verified
An organisation with 20% coverage and a 95% pass rate knows almost nothing about itself and looks excellent. An organisation with 80% coverage and a 60% pass rate knows a great deal and looks worse. Coverage is reported first, always, because it governs how much the pass rate can be relied on.
5. Effectiveness attaches to the playbook, never to the person
When a verification regresses, the result is recorded against the playbook and lowers that playbook's observed effectiveness inside the tenant. It is never recorded against whoever executed the work.
Cross-tenant effectiveness is not implied. IO does not pool tenant records. Any future shared benchmark would require a separately defined aggregation and privacy policy, contractual authority and a minimum sample that prevents an organisation or cohort from being inferred.
This is a design decision with a specific purpose. A method that scores individuals produces individuals who avoid measurable work. A method that scores playbooks produces playbooks that get retired when they stop working, which is the behaviour worth having.
Observed effectiveness affects ranking. Automatic suppression requires a published, versioned rule for the minimum sample, maturity context and failure threshold. Until that rule exists, product copy should say a weak playbook is ranked lower, not that it automatically stops being prescribed.
6. Failure is stored as fact
A regressed verification is not corrected, softened, re-baselined, or removed when the next quarter looks better. It stays in the record, attributed to the playbook that produced it, and it feeds forward: playbooks with poor observed effectiveness are ranked lower when the next exposure is matched, and forecasts inherit real history rather than optimism.
The uncomfortable version of this is the useful one. A board pack that reports what regressed is more credible than one that does not, and a vendor willing to show its own failures is making a claim that a vendor with something to hide cannot copy.
7. What verification does not establish
A post-intervention measurement shows what changed after the intervention. Attributing that change to the intervention requires a measurement design that supports the claim, a control population, a stable window, an isolated variable, and most operational programmes do not have one.
So the language is precise and stays precise: the measure moved in the intended direction following this intervention. Not this intervention caused the improvement. The difference is not pedantry. It is the difference between a claim that survives a sceptical reader and one that does not.
8. Reading a verification record
Every record carries: the playbook, the accepted action, the accountable owner, the measure and its unit, the baseline value and its capture date, the re-measured value and its date, the population, the declared direction, the result, and, where the result is inconclusive, the reason it could not be relied on.
If a reader cannot reconstruct the result from that record without trusting the number printed beside it, the record is incomplete.