Over fourteen months we ran forty-two experiments on one client’s checkout: Northbeam Outfitters, an outdoor gear retailer whose numbers are already public on our work page. Fifteen won. Nineteen came back flat. Eight lost. Checkout completion went from 57.4% to 67.7%, worth about £1.07m a year at their £142 average order.
This is the whole story, including the losers, because the losers carried most of the lessons. It is also a story about restraint: the single most valuable decision in the programme was not a test at all, but a shipping subsidy the board had already budgeted and never had to spend.
Why publish it at this level of detail? Because the industry runs on case studies with the losses removed, and a scoreboard missing its right-hand columns teaches the wrong lesson: that testing is a machine you feed ideas into and winners fall out of. It is not. It is a machine for finding out, and what it found out here is the story.
The setup: a leak everyone had normalised
42.6% of everyone who started this checkout never finished it. Nobody inside the business had looked hard at that number precisely because it is normal; retail abandonment tables say so. The working theory was price sensitivity on shipping, since a 65-litre pack costs real money to deliver, and the planned fix was a subsidy.
Six weeks of research said otherwise. Joining fourteen months of order data to product dimensions showed baskets with an oversize item abandoning at 51.3% against 36.8% for everything else. Thirty phone calls to abandoners found almost nobody refusing to pay: nineteen of thirty said some version of I wanted to check something first. Delivery cost, arrival date, returns terms. They were not objecting. They were leaving to find out, and mostly not coming back.
The constraints shaped everything that followed. Roughly 180,000 sessions a month with a 3.4% checkout-start rate meant about 6,100 checkout entries, so any effect below 4% relative was undetectable in a sensible run time, and checkout tests had to run one at a time to avoid splitting an already thin funnel three ways. October to January carried 54% of annual revenue and sat behind a hard ten-week code freeze, so the peak we most wanted to influence was the one we could not touch. Fifty-eight hypotheses were written; forty-two survived scoping to become experiments.
The scoreboard, and the four tests that mattered
Every hypothesis had to name the specific thing a shopper was leaving to find out and the specific place we would tell them instead. That rule threw out most of the existing backlog, including an express-pay reshuffle and three separate field-reduction proposals that session data showed were solving nothing.
- Oversize delivery rate shown on the cart page: +6.2% completion at 98% confidence. The surcharge scared people less than the mystery did.
- A dated arrival window instead of a 3 to 5 working days range: +4.9% at 97%. Trips have dates; ranges make people leave to calculate.
- Returns terms inside the payment step: +3.1% at 96%, concentrated in footwear baskets, where can I send boots back after trying them indoors was the second most coded reason for leaving.
- Guest checkout as the default path: +2.2% at 95%, on a checkout where 71% of entries were first-time buyers.
- The instructive loser: a security and payment badge row, standard advice everywhere, cost 2.8% at 93%. On this audience, badges raised the question they claimed to answer.
- The instructive flat: a free-shipping progress bar, the most requested feature internally, moved completion 0.3%, which is to say nothing.

The secondary metrics moved with the primary one, which is how you know a completion gain is real rather than borrowed. Shipping-step exits fell from 21.4% to 12.9%. Median time in checkout dropped from 214 seconds to 148, not because fields were removed, but because nobody was leaving mid-flow to hunt for a returns policy. Sitewide conversion carried the gain through, from 1.95% to 2.30% on unchanged traffic, which at a £142 average order is roughly 630 extra orders a month.
What the losers taught
The eight losses were written up in the same detail as the winners, and one write-up has since done more work than several winners: the badge-row result has stopped two other teams in the business from adding trust badges to their own flows. A documented loss is a fence; an undocumented one is a lesson you will pay for again. Most of the ways tests fail before launch appeared in this programme’s first rejected backlog, which is exactly where you want to catch them.
The flat results earned their keep too. Nineteen tests moving nothing sounds like waste until you price the alternative: nineteen changes shipped on instinct, each one a permanent piece of interface justified by a feeling. The flat tests are why this checkout stayed clean.
The freeze deserves its own mention, because respecting it was a decision made against internal pressure both years. Ten peak weeks with nothing shipping feels unbearable when a test is one week from significance. What made it bearable was using the freeze deliberately: read-only research on the year’s richest traffic, and the write-ups that turned eight losses into fences other teams still walk along. Peak taught more than it earned, which was the plan.
What transfers to your checkout
Not the specific winners. Showing an oversize delivery rate early is the answer to a question this audience was asking; yours is asking something else. What transfers is the shape: find the question people leave to answer, answer it inside the flow, and read the result by device, because mobile started at 48.2% completion here and did most of the work on the way to 61.1%. The general patterns live in our checkout UX guide; the statistical discipline that kept these verdicts honest is covered in our statistics guide.
And one number to keep: four winners compounding to about +17.4% is the argument for programmes over projects. No single test here was a headline. The accumulation was worth a million pounds a year. If your checkout has a number everyone has normalised, that is usually where test development pays for itself first.



