Experiment analysis & readouts
Readouts that settle a question instead of reopening it
We turn a finished experiment into a decision somebody can sign: the effect with its range rather than a bare verdict, guardrails read against limits agreed before launch, and a document short enough that whoever pays for the change reads all of it.
Most of that work happens before the test is live, because a rule written while the answer is unknown is the only kind anybody can be held to afterwards.
- Nothing priced as a share of what a test wins
- Month to month, with no contract to work out of
The problem
Why a finished test still does not produce a decision
The statistics are rarely why a result sits unactioned for three weeks. What stalls a decision is that nobody wrote down, in advance, what would count as one.
A verdict with no size attached
The dashboard says the variation won. Nobody can say by how much, within what range, or whether that range repays the engineering it would take to keep.
The rule arrived after the numbers
A metric nobody named up front turns out to be the one that moved, so it becomes the metric that mattered all along. That is a story assembled backwards.
Nothing was allowed to veto the win
Revenue per session is up, and returns, support contacts and page weight were never on the scorecard, so a change costing more than it earns still reads as a success.
The decision was never recorded
Six months on, the same idea returns from somebody who joined after it was tested. With no written disposition, every test stays available to argue about.
So the work is not the arithmetic. It is agreeing the rule before launch, applying it honestly afterwards, and leaving a document that answers what the next person will ask.
Reading a result: the size first, the verdict second
A significance verdict answers one narrow question about whether an observed gap is surprising. It puts no number on what the gap is worth, and worth is the only thing a business can act on. That number lives in the measured effect plus the range around it, which the NIST handbook defines as a span of values likely to contain the true one. Two results can clear the same threshold and mean entirely different things once you see how wide their ranges are.
That distinction is worth learning rather than delegating, and our article on what statistical significance actually claims works it through on real numbers. If your platform reports probabilities instead of p-values, the Bayesian and frequentist comparison explains what changes and what does not.
Calibration matters as much as method. The Microsoft and LinkedIn authors of Seven Rules of Thumb report that at a site running thousands of experiments a year most fail, and those that succeed move key metrics by 0.1% to 1.0% once diluted to overall impact. Against that, a claimed double-digit jump earns the scepticism the same paper attaches to Twyman’s law: any figure that looks interesting or different is usually wrong.
The rule gets agreed before the test launches
One artefact removes most of the arguments a readout can start. Before a variation goes live, four things are written down and circulated: the primary metric, the smallest effect worth acting on, the date the test stops, and the guardrails able to overrule a positive result. Agreeing all four while the answer is still unknown is what makes them binding afterwards.
The stop date is skipped most often, and it is a calculation rather than a preference. Our sample size and duration calculator turns your baseline and smallest useful effect into visitors and weeks; the runtime arithmetic behind it covers what to do when that answer is longer than anyone hoped.
Fixing the date is also what makes the significance threshold mean what the tool claims. Dmitriev and colleagues at Microsoft put it plainly in A Dirty Dozen: monitoring continuously and halting the instant a threshold is crossed raises the Type-I error rate, and they name both early stops and quiet extensions as common mistakes. They also warn that a significant reading on an underpowered metric tends to exaggerate the real effect, so a large win on a thin slice is a lead, not a result.
The metrics allowed to overrule a winner
A guardrail is not a success metric. The same Microsoft team define guardrails as metrics that do not indicate whether the feature worked, but that you refuse to significantly harm in exchange for a ship decision, small movements expected and large ones not permitted. For a store that means returns and refunds, support contact rate, delivery promise accuracy, and paint timings on the pages the test touches.
Their thresholds belong in the pre-launch rule, because a limit set afterwards is a negotiation. Our guide to guardrail metrics names the five an ecommerce test should carry, and the four dispositions once one is tripped covers the awkward case where the winner is also the culprit. A guardrail that has never stopped anything is decoration.
What we build
What an analysis engagement covers
Before launch
- A single deciding metric, plus the smallest movement worth having
- Runtime and stop date calculated from your own baseline
- Guardrails chosen and limited while nobody knows the answer
While it runs
- Assignment integrity watched throughout, so a broken split surfaces early
- Movement reported without a verdict, so a glance cannot become a decision
- The only conditions that may end a test early, fixed before it starts
The readout
- Effect and interval on the named metric, at the size the data supports
- Guardrails read against agreed limits, not against significance alone
- A recommendation in plain words, with reasoning a sceptic can take apart
After the decision
- A disposition recorded for every test, winners and losers alike
- Archive entries written to be searched, not only read
- What this result makes worth testing next, and what it retires
What the readout document has to survive
The test of a readout is whether it holds up in front of somebody who sat in none of the meetings and has no reason to be generous. That reader asks three things in order: what did you say you would measure, what did you find, and what are you doing about it. A document that cannot answer all three is a slide deck.
So it opens with the disposition and the reasoning, reproduces the pre-launch rule unedited, then gives the effect with its interval, the guardrail panel against its agreed limits, the checks proving the arms were comparable, and any segment declared in advance. Our significance article publishes the reporting standard behind that shape, and it stays short enough that a finance director reads all of it.
One check sits ahead of the rest, because it can invalidate them. Where the split a tool actually delivered has drifted off the ratio you asked for, you do not have a result yet. Run that yourself with our free SRM checker before anyone reads a conversion number.
Losers and flat results are most of a programme
Winners are the minority of any honest programme. That is not a programme working badly, it is a programme finding out, and the value survives only if the finding is written where the next person looks. So every result leaves with a disposition and a row in the archive, dull ones included.
| Outcome | What it licenses | Where it goes |
|---|---|---|
| The primary metric moved and no guardrail objected | Shipping it, at the size the interval supports rather than the headline | Rebuilt natively, then watched for the first weeks after removal of the test code |
| The primary metric moved the wrong way | Ruling the idea out in this form, which is a genuine saving | Archived with the reasoning, so it stops being re-proposed each quarter |
| The interval spans zero after the planned runtime | Saying the true difference sits under this test’s detection floor | Archived as a bound on the idea, with the effect the test could have seen |
| A guardrail broke its agreed limit | Refusing the win until the cost is understood, whatever the primary metric did | Held, investigated, then re-run in a form that does not pay for itself twice |
The archive is what stops a settled question being reopened by somebody who joined last month. Ours is free and ungated as the experiment archive template: twelve fields, a fixed tag vocabulary, worked examples. Take it and run this yourself. Our catalogue of the ways tests fail is mostly ideas the whole room felt certain about.
How it runs
Four steps, and the first one is before launch
The order matters because most of what makes a readout trustworthy is decided before there is anything to read.
- 1
Agree the rule
We write the decision rule with you: the metric that decides, the effect worth having, the stop date the arithmetic gives, and the guardrails with their limits, circulated before launch.
- One metric named as the decider
- Stop date calculated, not guessed
- Guardrail limits set in advance
- 2
Watch without judging
Assignment and data quality are checked from the first day, while movement is reported without a verdict. Only a condition written down before launch stops a test early.
- Assignment verified on the first day
- No interim verdict issued
- Stopping conditions fixed up front
- 3
Read it against the rule
At the stop date the result is read against what was agreed, in that order: data quality, the primary metric with its interval, the guardrails, then anything pre-declared.
- Data quality cleared first
- Effect reported with its interval
- Guardrails read before the verdict
- 4
Write the decision down
A readout short enough to be read in full, a disposition attached whichever way it went, and an archive entry a colleague can find in a year without asking anyone.
- Disposition recorded either way
- Archive entry written and tagged
- Next questions named, not implied
Where this fits in a testing programme
Teams buy this on its own, having built a programme they cannot get decisions out of. No tool changes and nothing is handed across. We will also say when the analysis is fine and the problem is upstream.
If the events feeding the numbers cannot be trusted in the first place, tracking and QA validation is the pass that comes before this one. If the shortage is builds rather than conclusions, A/B test development is the service that ships them, with the readout included in each.
The answers to your questions.
You do, against a rule the two of us wrote before launch. Our job is to make that rule explicit while nobody yet knows the answer, apply it without flinching once the numbers arrive, and say plainly when the test cannot settle the question. No readout we write asks you to take a verdict on trust: the rule, the arithmetic and the reasoning are in the document, so you can disagree on evidence rather than authority.
Often, though what can be recovered depends on what was recorded at the time. Where the raw assignment and event data survive we can check the split, recompute the effect with an interval and see whether the guardrails were ever examined. What cannot be recovered is a decision rule nobody wrote, and inventing one after seeing the results is the exact failure to avoid.
It is written up like any other outcome, because an inconclusive test still narrows the field. A range spanning zero once the planned runtime expires says any real effect is below what this test could see, which is a genuine bound on the idea. The two wrong answers are running on until it turns positive, and dropping it quietly so the same suggestion returns next quarter.
Read access to the experiment results and the analytics behind them is usually enough, and an export works where access is awkward. We configure nothing, run nothing, and ask nobody to move tools. Where the platform and your analytics report noticeably different numbers, that gap gets examined first, since it decides how far either can be trusted.
Usually not as a standing arrangement, and saying so costs us nothing. On a few tests a year the higher-value fix is the rule and the archive, both of which you can run yourself with our free templates and calculators. It earns its keep on a programme producing results faster than anyone can honestly read them, or one where the same decisions keep being reopened.
Sitting on a result nobody can act on?
Send us the test, the metric it was meant to move and what the tool reported. We will tell you what it can honestly support, and whether that needs us at all.
Show us the result ↗