Personalisation is what testing vendors sell you when the demo needs a wow moment. Sometimes it is genuinely the right tool: a returning customer should not be greeted like a stranger, and a shopper in Oslo should not read delivery promises written for London. But every personalised experience is a segment-sized A/B test, and most stores do not have segment-sized traffic.
Ecommerce personalisation means showing different visitors different experiences based on what you know about them: their purchase history, their location, their traffic source, the state of their cart. A/B testing shows every visitor the same candidate experiences and measures which one wins. The first multiplies the experiences you maintain; the second reduces them to the single best one. That difference, not the technology, is what should decide between the two.
Every personalised experience is a test you now own
The uncomfortable arithmetic is that personalisation does not remove the need for evidence, it multiplies it. Show returning visitors a different homepage and you have created two experiences, each of which can be better or worse than the one it replaced, for its own audience. The honest way to know is to test within each segment, and each of those tests draws on a smaller pool of traffic than a whole-site test would.
Skip the evidence and you are guessing twice instead of once, while paying to maintain every variant forever. We have audited stores running seven personalised homepage variants that nobody had measured since launch, at least one of which was quietly underperforming the default it replaced. Personalisation platforms rarely surface this, because a report that says your rules are losing is not a report the platform is keen to build.
The traffic maths that decides
Work one example through. A store doing 80,000 sessions a month at a 2.4% conversion rate wants to know whether a new homepage layout wins. Run it as a whole-site test and detecting a 10% relative lift needs roughly 65,000 visitors per variant, which at 40,000 sessions per variant per month resolves in about seven weeks. Slow, but workable.
Now split the audience four ways: new, returning, paid, organic. The largest segment might carry 30,000 sessions a month, so the same test inside it gets 15,000 sessions per variant per month and needs over four months to reach the same answer. The smaller segments never get there at all. Our guide to A/B testing sample size walks through the calculation properly; the short version is that segmentation divides your traffic while multiplying the evidence you owe.

Where personalisation reliably pays
New versus returning is the split that earns its keep most often. A first-time visitor needs orientation: what you sell, why you are credible, what it costs to try you. A returning customer needs none of that. One fashion store we worked with swapped its brand-story hero for a restock and new-in module for returning visitors, and that segment’s conversion moved from 3.1% to 3.4%, a 10% relative lift at 95% confidence. The experience was not cleverer; it just stopped repeating an introduction.
Geography is the second dependable win, because it is barely personalisation at all: it is accuracy. Showing a German visitor German delivery times and duty-inclusive prices is telling the truth earlier. The same logic covers currency, returns terms and which warehouse the order actually ships from, and none of it needs a segmentation strategy meeting.
Source-matching and cart-state complete the reliable list. A visitor arriving from an ad about one product should land on that product and its promise, not on a generic homepage. A visitor returning after three days with a full cart needs a route back to it, not a fresh pitch. All of these share a property worth noticing: the difference between segments is a fact you know, not a taste you are guessing at.
Where it becomes a money pit
AI product recommendations on a small catalogue are the classic sinkhole. Recommendation engines need breadth to say anything a bestsellers row could not; on a catalogue of eighty products, shoppers-also-bought converges on the same six items a merchandiser would have picked in an afternoon, at a monthly platform fee. The demo looks like magic because the demo runs on a marketplace-sized catalogue you do not have.
Micro-segmentation is the other. Every rule you add halves somebody’s sample. By the time a store is customising for lapsed high-value weekend mobile shoppers, the segment is a few hundred sessions a month, the experience can never be validated, and the rule set has become something nobody dares touch. If a segment is too small to test in, it is too small to personalise for.
Personalise or test? A working rule
The decision is five questions, not a platform demo:
- Is the difference between segments a fact you know, like location, login state or cart contents, rather than a hunch about taste? Facts personalise well; hunches test poorly.
- Would every segment benefit from the same improvement? If yes, run one whole-site A/B test and bank the shared winner first.
- Does each segment carry enough traffic to validate its own experience within a quarter? If not, it cannot be personalised honestly.
- Can you name the metric each personalised rule must beat, and the number it currently sits at? A rule nobody measures is a liability.
- Who maintains each variant after launch? Every experience you add is a page somebody updates forever, through every rebrand and replatform.
Shared winners come first for a reason: a whole-site test resolves faster, its result compounds across every visitor, and it leaves your traffic intact for the next question. On Shopify specifically, the mechanics of running those tests without corrupting their numbers have their own traps, which our Shopify A/B testing guide covers in detail.
Validate with holdbacks or you are guessing
Once a personalised experience ships, the discipline that keeps it honest is a holdback: a slice of the segment, usually 5 to 10%, that keeps seeing the default experience. A rule that is genuinely better should beat its holdback continuously, not just in the launch-month deck. When the gap closes, and it often does as novelty fades or the audience shifts, the rule has earned retirement. Personalisation without holdbacks is a collection of anecdotes with a dashboard.
That instrumentation, building the segments, the variants and the measurement without flicker or sample corruption, is engineering rather than strategy, and it is the same engineering an experiment programme runs on; it is what A/B test development exists to do. If you are weighing a personalisation platform against a testing programme and want a second opinion on the traffic maths first, get in touch.



