IO Verify™, the surface that holds the answer.
Verification is the argument. Verify is where it is kept. It reads live from the completed work and returns four things: what each intervention returned, how much of the work was measured at all, which playbooks actually move a measure, and where residual risk moved and because of what.
Before any result means anything: how much was measured at all.
Verify leads on coverage per domain, the share of completed work that carried a verification measure. Read it as the confidence interval on everything below it. Where coverage is low, the results on that domain are a sample, and Verify presents them as one.
Completed work
Everything the Execute step finished in the period, per domain, the denominator, whether or not it was measured.
Measured share
How much of that work had a baseline captured at approval and a re-measure taken afterwards. This is the number that qualifies every other number on the page.
Unmeasurable, and why
Work that could not be verified, with the reason attached, no criterion agreed, source unavailable, or population no longer comparable.
Four states, and none of them is 'complete'.
The measure moved in the intended direction beyond the noise threshold, taken from the same source and the same population as the baseline.
The measure did not move enough to claim anything. The intervention is recorded as unproven at this maturity for this cohort rather than quietly counted as done.
The measure moved the wrong way after the work completed. The case is named, the execution record is attached, and it appears in reporting at the same prominence as an improvement.
The comparison could not be made safely, the source changed shape, the population is no longer equivalent, or the playbook was only partly executed. Verify says so instead of producing a number.
What that looks like on the surface itself.
A real capture from the live application, taken against a demonstration tenant. Select it to enlarge.
A domain that got worse, and the instrument saying so.
This is what the instrument does when the news is bad. Figures are from a live demonstration tenant that re-seeds, so read them as an illustration of the shape rather than as fixed statistics.
Resiliency in the demonstration tenant reads lower in the current window than in the prior one: around 61 against roughly 76 at the last read, which carries the domain down a band, from managed to defined. It is the only domain in that tenant to change band, and it changed in the wrong direction.
Both windows stay on the record. The prior figure is not overwritten and the decline is not smoothed into a trend line, because the two windows are the comparison. The band change is derived from the same counted signals as the score, so a reader can walk the drop back to the rows underneath it.
IO does not assert a cause. Nothing in the evidence establishes which piece of work drove the decline, and manufacturing an explanation would be the exact pattern this platform exists to prevent. What it reports is the shape: a measured decline, a band change, both windows retained, and a named owner asked to account for it.
Most security reporting only knows how to show numbers going up. An instrument that cannot report its own decline is not an instrument.
Ranked, including the ones that do not work.
Effectiveness is attached to the playbook, never to the person who executed it. A playbook that has failed to move its measure across enough executions is ranked as such and stops being prescribed at that maturity until something changes.
Attached to the method
The unit of judgement is the playbook and the maturity level it ran at. Verify does not produce a leaderboard of owners, and it is not an input to anyone's performance review.
The failures stay visible
Playbooks that produced no change, or produced regressions, remain in the ranking with their record intact. Removing them would make the effective ones look better than the evidence supports.
Prescription reads this
The ranking feeds back into what gets prescribed next, which is the only mechanism on the platform by which a method can be retired on evidence.
Named, with the before and after shown.
A regressed case is listed individually: the intervention, the measure, the baseline captured at approval, the value returned at re-measure, and the interval between them. No aggregation that lets a bad result disappear inside a good average.
Each carries the execution record, who accepted it, what steps ran, and what was blocked, because the useful question after a regression is which part of the method failed, not who was on the ticket.
A callback starts the clock. It does not settle the result.
Where an action is executed by your own automation rather than by hand, that automation reports the result back to IO. The callback closes the baseline window and starts the measurement clock: it records the execution state, executed, failed, or not attempted, and the time the work happened, and it sets the re-measurement date running.
It does nothing else. It cannot set a verification outcome, move residual risk or alter a capability score. The verification outcome is still established only by re-reading the measure from the source system at the re-measurement date, from the same source and the same population as the baseline.
A callback records that work happened, and when. Whether it worked is established only by re-measuring the source system.
Every movement is attributed to the verification that caused it.
Residual risk on this platform is not a workshop opinion that gets refreshed quarterly. It moves when a verification result lands, and each movement links back to the specific result behind it, which means any figure on a board pack can be traced to the measure that produced it.
Watch a regression get reported.
The demo includes verification results that went the wrong way, because real programmes produce them.