Skip to content
6 min read · A/B testing

Ecommerce A/B Testing: A Practical Guide for DTC Brands

AWritten byAnika RahmanExperimentation Lead
Updated on 21 July 2026
Two variants of the same ceramic vase product page side by side, labelled Variant A and Variant B, with a results strip showing 2.4% against 2.8% and a +17% uplift badge

Most of the test programmes we get asked to rescue aren't test programmes at all. They're redesigns with a stopwatch attached: someone shipped a new page, watched the numbers for a fortnight, and called it a win because revenue rose in a month that also contained a payday, a heatwave and the start of a sale.

A/B testing is the discipline of making that judgement honestly. Two versions of a page live at the same time, traffic split at random, one number agreed in advance. Everything else — the tooling, the dashboards, the little green significance badges — is decoration on top of those four conditions. Get one of them wrong and the result is not slightly wrong. It is meaningless.

This is roughly the conversation we have with every DTC team before we start A/B testing with them. It will not make you a statistician. It will stop you spending six weeks proving nothing, which is worth more.

A test is a question, not a change

"Let's try a green button" is not a hypothesis. It is a guess with no failure condition. A hypothesis names three things: the friction you believe exists, the change you think fixes it, and the number that should move as a result.

Written out, it sounds like this: because exit surveys show delivery anxiety on the product page, showing the delivery date above the fold will raise add-to-cart by at least 5%. Notice what that sentence does. It commits you to a number before you see the data, which means it can be wrong — and a test that cannot be wrong cannot teach you anything.

If you cannot fill in that sentence, you will ship the variant when it is up and shrug when it is down. That is exactly what you would have done without running a test, only slower and with more meetings.

We reject about half the test ideas clients bring us at this stage. That rejection is free. Discovering the same thing after six weeks of live traffic costs a quarter of your testing year.

Evidence first, ideas second

The best predictor of a useful test programme is boring: a written backlog where every idea has a piece of evidence attached to it. Not a Slack thread. Not the founder's Saturday-morning thoughts. A list, with sources.

Most of that evidence is already sitting in your business, unread. Session recordings show people hunting for a size chart that is three scrolls down. Support tickets repeat the same delivery question forty times a month. On-site search logs are a list of things customers expected to find and did not. Checkout analytics tell you which field people abandon on. None of this requires new tooling — it requires somebody spending a Tuesday reading it.

Rank what comes out of that by expected impact against effort, and be ruthless about the top of the list. A backlog of forty ideas where the top five are genuinely evidence-backed will beat a backlog of two hundred brainstormed ones, because you only get to run so many tests a year and every slot you spend on a hunch is a slot you did not spend on a signal.

An experiment backlog interface with hypothesis, evidence, impact and effort columns, the top three rows highlighted and tagged session recordings, support tickets and search logs

The maths that decides whether you can test at all

Here is the calculation that kills more test programmes than any other, usually because nobody ran it. If your baseline conversion rate is 2.4% and you want to detect a 5% relative lift — that is 2.4% moving to 2.52%, not to 7.4% — you need somewhere around 210,000 visitors per variant to see it clearly.

Run that on a template getting 30,000 visitors a month and your "two-week test" is a five-month test. In practice you will stop it early, see a muddle, and record it as a loss. You will then have learned something false about your customers and carried it into the next three decisions.

When the maths says no, you have two honest options and one dishonest one. You can test a bigger change, because large effects need far less traffic to detect than small ones. You can test somewhere with more traffic — the collection page and the cart usually see far more people than a single product page. Or you can run it anyway and call whatever happens a result, which is the option most teams take.

Reading the result without fooling yourself

Decide the runtime before you launch and do not look at the numbers to decide when to stop. Checking daily and stopping the moment a variant goes green is called peeking, and it manufactures winners out of noise — with enough looks, almost any test will cross the line at some point on its way to nowhere.

Run for whole weeks, never part ones. Tuesday traffic does not behave like Sunday traffic, and a test that starts on a Monday and ends on a Thursday has quietly weighted your result towards weekday buyers.

And treat losses as information rather than embarrassment. A variant that loses by 8% has told you something true about your customers that you can act on everywhere else on the site. The genuinely wasted tests are the inconclusive ones — the flat results from underpowered tests, which tell you nothing at all and still cost you six weeks of traffic.

What to test first on an ecommerce store

If you are starting from nothing, these five tend to earn their traffic across almost every store we work on, because each attacks a question the customer is already asking:

  • Delivery and returns information on the product page, moved above the fold rather than buried in an accordion.
  • The product gallery — how many images, what order, and whether a scale or in-use shot beats another studio angle.
  • Guest checkout prominence, and how early you ask for an account.
  • Collection page filtering and sort order, especially on mobile where the default sort does most of the merchandising.
  • The cart drawer versus a full cart page, which is a genuinely different experience and rarely tested by the teams that argue about it.

When not to A/B test

Testing is not always the right tool, and pretending otherwise wastes months. If something is plainly broken — a form that fails on mobile Safari, a size chart nobody can find — fix it. You do not need an experiment to tell you that broken is worse than working.

If your traffic genuinely cannot support the test, do the qualitative work instead: watch recordings, run five user sessions, read the tickets. And if the change is a brand or legal decision, own it as one rather than hiding behind a number that was never going to decide it for you.

The programmes that compound are unglamorous. A written backlog, honest sample maths, fixed runtimes, and a willingness to lose in public. If you want a second pair of eyes on yours, get in touch — we reply within one working day.

Sharein𝕏f
AWritten byAnika RahmanExperimentation Lead

Anika has designed and read out 300+ experiments at Optyv. She joined from a quant background, still checks everyone's sample-size maths, and once killed the founder's favourite redesign with a two-week test. Nobody has let him forget it.

Do you like what you see?

Optimize your store