Articles

Why We Publish the Ones That Didn't Work

In one demonstration tenant, 183 interventions are recorded as regressed. The measure they were meant to improve moved the wrong way, and the record says so, permanently, attributed to the playbook that produced it.

Publishable articleDocument v3.14 min read

In one demonstration tenant, 183 interventions are recorded as regressed. The measure they were meant to improve moved the wrong way, and the record says so, permanently, attributed to the playbook that produced it.

That is a deliberate choice and an uncomfortable one to explain to a prospect. Here is the reasoning.

A programme where nothing fails is not measuring

Security interventions fail regularly. Awareness campaigns that move nothing. Access reviews that produce revocations which get reinstated within a month. Controls that push people into workarounds nobody sees. Anyone who has run a programme knows this.

So a verification ledger showing a 98% success rate is not evidence of an excellent programme by itself. It may reflect excellent performance, or one of three problems: the measures are chosen to be easy, the measurement happens only where success is likely, or failures are being quietly removed. The most common is the first, and it is usually not deliberate, it is what happens when the person choosing the measure is also accountable for the result.

The pass rate is not the interesting number. Coverage is. An organisation measuring 20% of its completed work and passing 95% of what it measures knows almost nothing about itself and looks excellent. One measuring 80% and passing 60% knows a great deal and looks worse.

Failure attaches to the playbook, never to the person

This is the design decision that makes the rest of it survivable.

When a verification regresses, the outcome is recorded against the playbook, the encoded method, and it lowers that playbook's observed effectiveness inside the tenant. It is never recorded against whoever executed the work. No cross-tenant effectiveness claim is made from customer records.

The reason is behavioural and unavoidable. A system that scores individuals on verification outcomes produces individuals who avoid measurable work, choose soft measures, and quietly stop accepting anything that could regress. Within two quarters the ledger is meaningless and the people are not to blame, the instrument taught them to do that.

A system that scores playbooks produces something else: methods that get retired when they stop working. That is the behaviour worth having, and it only exists if the failure lands on the method rather than the practitioner.

What a regression is worth in a board pack

More than a success, and this is not a rhetorical flourish.

A residual risk score that rose because an intervention was measured and found to have made things worse tells a board four things at once: the intervention happened, it was measured, the measurement was honest, and the register reflects reality rather than intention. No successful verification carries that much information, because a success is also what you would see if the measure were rigged.

The most valuable line in a quarterly pack is usually the one nobody wanted to write.

The commercial argument, stated plainly

There is a self-interested version of this and it would be dishonest to pretend otherwise.

A populated failure ledger is a stronger trust signal than a product that shows only success, but it is not proof of honesty. Failures can also be seeded, selectively framed or measured against trivial criteria. The commercially credible claim is narrower: IO has a defined place to retain regressions, shows their derivation, and does not require operators to erase them before reporting.

But the argument only works if the derivation is real. The demonstration tenant is synthetic, so the underlying before-and-after observations are seeded. The result labels are then computed by the same verification rule used elsewhere; they are not decorative outcome fields chosen after the fact. The manifest must publish both facts together so nobody mistakes synthetic evidence for customer history.

What this does not claim

A regression shows the measure moved the wrong way after the intervention. It does not establish that the intervention caused it. Attributing causation requires a measurement design most operational programmes do not have, a control population, an isolated variable, a stable window.

So the language stays precise: the measure moved against the declared direction following this intervention. Not this intervention made things worse.

The distinction survives a sceptical reader. The stronger claim does not.

The measurement principle underneath it

This is an application of a familiar incentive problem: when a measure becomes a target, people adapt to the target. Publishing coverage, unsuccessful results and measure definitions makes gaming more visible; attaching outcomes to the playbook rather than the person reduces the incentive to avoid difficult work. It does not eliminate selection bias, metric gaming or management pressure, so those remain governance concerns.

Next step

See it before you talk to anyone.

Two quarters of recorded activity across sixteen seats. One click, no install.