Skip to content
5 min read · A/B testing

Why We Kill Our Own Redesigns (Evidence Over Ego)

MWritten byM R Q MotinFounder & CEO
Updated on 6 August 2026
An A/B test results panel where a redesigned product page variant loses to the control on revenue per visitor, £2.40 against £2.50, a 4% drop with a do-not-ship verdict

Last quarter our design team produced the best product page we have ever made. It is now deleted. It ran for five weeks against the page it was meant to replace, lost by 4% on revenue per visitor, and at Optyv the rule is absolute: if it cannot beat the control, it does not ship. Including, especially, our own work.

Evidence over ego is the closest thing we have to a company constitution, and this essay is what it means in practice: what the rule costs us, what it earns the brands we work with, and how to install a version of it in your own team. It is not a slogan; it is a specific operating rule with a body count of beautiful deliverables behind it.

The page we deleted

The redesign was better by every standard designers are trained on: cleaner hierarchy, stronger photography, a tighter buy box, whitespace where the old page had clutter. We tested it the way we test everything, a 50/50 split with revenue per visitor as the primary metric, and the control won: £2.50 against the variant’s £2.40, a 4% drop the test was confident about. The recordings suggested why. The clutter we had removed included the delivery estimate and the returns line, and shoppers who could not find them hesitated exactly where they used to click. The variant was archived with its learning, and the learning shipped instead: the old page kept its answers and gained the tighter buy box in a follow-up test that won.

Killing it was not a close call, and that is what the process is for. The decision had been made before the test started: revenue per visitor was named as the primary metric, the sample was planned, and the number came in where it came in. Nobody had to win an argument against the design team; the design team read the result out themselves, logged the learning, and opened the next test. The discomfort is real and it is also brief, which is the part outsiders never quite believe.

An experiment archive entry logging a killed product page redesign, recording the £2.50 to £2.40 revenue-per-visitor result, a kill decision, and the learning that the delivery estimate and returns line must stay near the buy button

Why agencies hate this rule

The economics of most agencies make this rule impossible, which is worth understanding before you hire anyone. An agency that bills for the redesign has sold the deliverable before knowing whether it works; the deliverable is the product, so it must ship, and it must be praised on launch day. Testing it honestly afterwards is all downside: at best the invoice is validated, at worst the flagship project of the quarter is exposed as a loss. So full redesigns launch on faith, the case study is written from launch-week screenshots, and nobody is paid to check the revenue three months later. None of the people involved are dishonest; the incentive does all the work.

It is the same incentive that scopes redesigns as single dramatic cutovers rather than staged, tested rollouts: the reveal is billable, the discipline is not. When we say evidence over ego, the ego in question is usually not the designer’s. It is the invoice’s.

What the rule costs us

Honesty about the trade-offs, since a manifesto without costs is just marketing. Our portfolio grows more slowly than our output, because a meaningful share of our best-looking work is dead and we will not present killed designs as if they shipped. Sales calls get awkward when a prospect wants their favourite idea built and we say the base rate is against it; most tests lose, including ours, and we say so before contracts are signed. And the design team carries a discipline most studios never ask of theirs: watching strong work die on a number, repeatedly, and going back to the file with the learning instead of the grudge.

There is a quieter cost too: we cannot promise outcomes, only process, and process is a harder sale than a wall of before-and-afters. A studio that ships everything can show you forty launches a year. We can show you the subset that survived, and then explain, to a prospect who wanted reassurance, why the graveyard is the credential.

What it earns clients

The return is that every change we ship has beaten the incumbent on your revenue, not on our taste. Nothing in the account is decoration; the losers were caught at 50% of traffic and billed as information rather than rolled out as regret. Our 42-test checkout case study shipped 15 winners; the other 27 tests were the fee for knowing, and the checkout that resulted carries an 18% completion gain nobody has to take on faith. Clients stay for that: our client retention is 96%, and we credit this rule for most of it, because a partner who will kill their own work is a partner whose recommendations mean something.

It also changes what a long engagement looks like. Because losers die at half traffic and winners compound, the account gets safer with time rather than staler: the archive of what your audience has already rejected grows, the priors sharpen, and the third year of testing starts from a hundred answered questions the first year did not have.

Three practices to steal

You do not need our whole operating system to get the core of this. Three practices carry most of the value, and none of them needs a tool:

  • Write the kill criteria before the work starts: the primary metric, the sample, and the number below which the design dies, agreed while everyone is still calm.
  • Separate the maker from the verdict: whoever built the thing presents the result, but the pre-agreed number makes the decision, so nobody has to defeat a colleague to kill a design.
  • Give losers the same ceremony as winners: log every result with its learning, and cite dead tests in planning the way you cite live ones.

The habit compounds quietly. A team that has killed three of its own designs on evidence stops arguing from seniority, because everyone has watched the number outrank the loudest voice in the room. If you would rather borrow the discipline than build it from scratch, get in touch: we are easy to work with, provided you do not need your ideas to win.

Sharein𝕏f

The questions people ask first

It is the operating rule that a pre-agreed metric, not seniority or taste, makes every shipping decision, and that the rule applies most strictly to your own work. In practice it means any change, including a beautiful redesign, dies if it cannot beat the current version in a controlled test.

Because looking better and selling better are different properties, and controlled tests keep finding them pointing in opposite directions. A cleaner page that hides the delivery estimate can lose revenue while winning every design review; when the two conflict, the revenue evidence decides.

Rarely, because the incentives point the other way: a deliverable that has been billed must ship, and testing it honestly can only embarrass the invoice. Ask any agency how many of their own designs they killed on results in the last year; the number tells you whether testing is a process or a slide.

With the number, quickly, alongside what the test taught and what the follow-up variation will try. A loss caught at 50% of traffic is cheap information rather than shipped damage, and clients who hear about losses promptly trust the wins, which is the entire value of reporting both.

MWritten byM R Q MotinFounder & CEO

Motin founded Optyv and still sits in on the readouts. He is happiest when a test proves him wrong in public.

Do you like what you see?

Optimize your store