BFCM traffic is enormous, cheap to test on, and made of people who are not your customers. A normal November week at a mid-size store might bring 60,000 sessions at a 2.2% conversion rate, 42% of them new visitors, 11% using a discount code. The peak weekend brings 180,000 sessions at 3.6%, 68% new, and 74% using a code. Whether that is an opportunity or a trap depends entirely on what you plan to do with the answer.
This post owns one question and refuses the rest of the season: does a result measured on promotional traffic generalise to the eleven months when nothing is on sale? The preparation work, the fix list and the code freeze belong to our BFCM checklist and are not repeated here. What follows is the reading problem, which is the part that survives long after the season ends.
Peak is a different population, not more of the same
An experiment measures the effect of a change on the people who saw it. That sentence is uncontroversial until the people change underneath you. Peak traffic differs from normal traffic on almost every dimension a test usually targets: the new-visitor share roughly doubles, price sensitivity is the reason most of them are there at all, a large slice arrives through deal aggregators and voucher sites rather than your usual channels, mobile share rises, and a meaningful proportion are buying for somebody else, which changes how they treat sizing, delivery dates and returns policy.
That matters because the mechanisms most tests rely on are exactly the ones peak has already saturated. Hesitation, urgency, price framing and first-purchase trust are all being supplied in bulk by the sale itself. A change that works by supplying one of them has nothing left to add, and a change that works against them looks stronger than it is. Neither reading describes February.
What generalises and what does not
The split is not random. It follows the mechanism the change relies on, so the question to ask of any peak result is whether its mechanism depends on the season being on.
- Travels: build quality and defect fixes. A faster template, a repaired mobile breakpoint or a size filter that finally works helps every population there is, and a fix that wins at peak will not stop winning in March.
- Travels: factual information that answers a universal question. Delivery estimates, returns terms and stock clarity are read the same way by a bargain hunter and a full-price buyer, though the size of the effect will differ between them.
- Does not travel: urgency and scarcity. Peak traffic arrives pre-urgent, and a countdown winning against a Black Friday deadline is measuring the deadline rather than the countdown.
- Does not travel: price and offer framing. Every element of how a discount is presented gets read against every other discount that visitor saw the same morning, and that comparison set does not exist in February.
- Does not travel: first-time trust elements. Peak converts sceptical visitors that normal trading never converts at all, so a trust badge or a reviews block can test weak simply because the discount was already doing that job.
- Does not travel: navigation and discovery aimed at gift traffic. Somebody buying for another person browses differently, so a category restructure that wins in December can lose in a month when visitors already know what they came for.
The practical consequence is that a peak result needs a label rather than a verdict. Anything in the first group can be shipped on the strength of a peak reading, because the population difference is working against you rather than for you: an effect that survives a crowd of bargain hunters will not weaken when calmer traffic returns. Anything in the second group gets logged, held and rerun. Making that call takes ten seconds at readout and saves the argument in March, when the number stops reproducing and somebody starts blaming the tool.
The one thing peak is uniquely good for
Test the promotion itself. Offer framing, banner wording, whether the discount is expressed as a percentage or an amount, sitewide against category-level, the free-delivery threshold during a sale: these are all questions whose answers you want to apply to peak traffic, and peak traffic is the only sample that can answer them. Log the entry as promo-only in your archive, read it again next November, and you accumulate a body of seasonal evidence rather than starting each year with the same argument in a meeting room.
The second genuine advantage is sample. A checkout test on a thin funnel that would take five weeks in normal trading can reach its planned sample in four days at peak, which makes the season the one window where small effects on deep funnel steps are detectable at all. That is a real opportunity, provided the change being tested is one whose mechanism travels, and provided nobody is tempted to reallocate traffic mid-run because the numbers arrived early.

Reading a promo test without lying to yourself
Four habits do most of the work. Split every peak result by new against returning visitors, because the two arms of a promo test can hide completely different effects behind a blended number. Split again by whether a discount code was applied. Tag the archive entry promo-only so that a future reader knows what population produced it. And rerun anything you intend to keep, from zero, once traffic composition has returned to normal, which for most stores means mid-January at the earliest.
One example makes the point better than the rule does. An urgency banner tested over a peak weekend returned +9% on conversion at 97% confidence, on traffic that was 68% new. The same variation rerun in January returned +1% and did not reach significance. Nothing was broken in either test. The first one measured a population that was already in a hurry, and shipping it on the strength of that number would have meant carrying a permanent banner for an effect that only exists four days a year. This is a close cousin of the novelty effect, except the thing that fades is the audience rather than the surprise.
What we actually run during peak
In practice, very little, and deliberately. Anything already live before the freeze keeps running if it has a clean sample plan and a guardrail set. New launches during peak are limited to promo tests and defect fixes. Everything else becomes read-only research on the richest traffic of the year: recordings, funnel reads and support-ticket themes that fill the January backlog. On Shopify the usual platform constraints still apply and are worth knowing before you commit to any peak test, which our Shopify testing guide covers method by method.
The honest summary is that peak is a good place to learn about peak and a poor place to learn about your store. Treat it that way and the season stops being a gamble with your roadmap. If you would rather have the promo tests built, read and archived properly by a team who does this every November, our A/B test development work is built for exactly that window, and you can get in touch before the calendar makes the decision for you.



