You shipped the winner and the lift evaporated inside a month. Nobody lied to you and nothing was miscalculated. The test was real, the number was real, and what it measured was the reaction to a change rather than the value of the change. That is the novelty effect, and it is the most polite way an experiment can deceive you.
The term gets its one-line definition in our A/B testing glossary. This post does the two jobs a definition cannot: catching it while the test is still running, and proving after rollout whether the win you shipped is still on the site.
Novelty, and its twin
Novelty is a returning visitor reacting to the fact that something changed. The new banner gets clicked because it is new. The moved button gets found because the eye lands on whatever moved. Give it three weeks, the visitor stops noticing, behaviour drifts back to baseline, and the lift goes with it.
The mirror image is the primacy effect, and it is the reason you cannot simply discount early results and call the job done. Returning visitors sometimes resist a change they have muscle memory against, so the variant gets punished at the start and recovers as familiarity builds. A change that reads as a small early loser can be a real winner you killed on day four. Both effects share the one property that makes either of them detectable: they exist only in people who remember the old version.
Detection one: split new against returning
This is the sharpest diagnostic available and almost nobody runs it, mostly because it is a report nobody thinks to open on a test that is winning. A first-time visitor has no memory of the previous design, so they cannot experience novelty or primacy at all. Whatever they show you is the effect of the change itself, uncontaminated by recognition.
So split the result. A promotional badge test we ran reported +11% overall. Split by visitor type, new visitors were up 3% and returning visitors up 22%, with returning traffic making up 40% of the sample. That is not an interesting segment finding to note in the appendix. It is a warning printed in bold: the effect lives almost entirely in the people who could tell that something had changed.
It reads just as usefully in the other direction. When new and returning move together, novelty is not the explanation and you can ship with more confidence than the headline number alone would justify. Be careful with the arithmetic, though: the new-visitor slice is a fraction of your sample, so it is noisier and will often be inconclusive on its own. You are comparing directions rather than declaring a second significant result, and our note on A/B testing statistics explains why that distinction is not pedantry.
Detection two: slice the lift by week
Plot relative lift week by week rather than cumulatively, which is the opposite of what every dashboard shows you by default. A genuine effect is noisy and roughly flat. A novelty effect decays. The badge test above ran +19% in week one, +12% in week two, +7% in week three and +3% in week four, and that is not a winner cooling off. It is a shape, and once you have seen it twice you will recognise it in a glance.

One honest caveat. Week one is the smallest sample and therefore the wildest, so a high first week followed by a lower second week is the most ordinary thing in experimentation and proves nothing at all on its own. Decay becomes a finding when it is consistent across four or more weeks and the new-visitor segment does not show the same slope. Two weeks of decline is a reason to keep the test running, never a reason to call it.
Which changes are most vulnerable
- Motion and animation. Anything that moves is a magnet on first sight and wallpaper by the fourth visit.
- New badges, banners and ribbons on a familiar template. The strongest novelty producers we see, and by some distance the most common.
- Colour changes to elements a regular customer already knows exactly where to find.
- Layout reshuffles that relocate a known control to somewhere better in theory.
- Popups, interstitials and gamified elements, which begin as a surprise and end as an obstacle.
- Anything shipped to a returning-heavy audience: subscription stores, replenishable products, and traffic bought back through retargeting.
- Anything launched in the same week as a campaign, a promotion or a seasonal peak, where the change and the calendar are impossible to tell apart afterwards.
The inverse is a genuine reassurance. Repairing something broken, adding information that was missing, showing delivery cost earlier, making a page faster: none of these depend on surprise, and in our archive they almost never decay. If the mechanism of your win is that the customer now knows something they needed to know, novelty is not your risk. If the mechanism is that the customer noticed, it is.
Segment-level experiences are a special case worth naming, because a treatment aimed specifically at returning customers is the most novelty-exposed thing you can ship. The holdback discipline described in our piece on ecommerce personalization exists for precisely this reason.
Holdback validation is the only real proof
Detection during a test is inference. Proof arrives after rollout and costs 10% of your traffic. When the winner ships, keep a randomly assigned 10% holdback on the control for four to six weeks, then read the gap at the end of that window rather than at the start. If the effect was novelty, the two lines converge and you get to watch it happen. If it was real, the gap holds and you have a number you can put in a forecast without flinching.
The badge test settled at +4% against its holdback. Not the +11% that was reported, and not zero either, which is the usual outcome: a novelty-inflated winner rarely evaporates completely, it settles at a fraction, and the fraction is the figure that belongs in your archive and in whatever forecast somebody built on the original result. Across the 1640+ tests in our archive, roughly one shipped winner in nine has come back materially smaller under holdback, and a handful have come back at nothing at all.
Two habits make this routine rather than heroic: run the new-against-returning split on every test as a matter of course, and put a holdback on any winner whose mechanism was attention rather than information. Neither costs much. Together they prevent the quarter where every test won and the revenue did not move, which does more damage to a testing programme than a run of honest losses ever does. That habit, and the build discipline underneath it, is what experiment engineering means in practice, and we are happy to look at any result that fits this pattern.



