Two weeks is not an answer, it is a habit. It became the default because it is a round number that covers two weekends, and because it is long enough to feel rigorous without anybody having to do arithmetic. How long a test needs to run is not a matter of taste or diligence. It falls out of four numbers, and if you have those four numbers you can work it out before you brief anything.
The uncomfortable part is that the answer is often much longer than two weeks, and sometimes long enough that the honest conclusion is not to run the test at all. That is worth knowing on day zero rather than on day forty.
Duration is an output, not a decision
A test finishes when it has collected enough visitors per variant to distinguish the effect you are looking for from noise. That number depends on how often your page converts today and how big a change you want to be able to detect. Traffic then converts it into calendar time. Nothing about the two-week convention enters into it.
The four inputs, and what each one actually controls:
- Baseline conversion rate: how the page under test converts now. Not sitewide, and not last year. A product template converting at 2.5% and a checkout step converting at 60% need wildly different samples.
- Minimum detectable effect: the smallest improvement you want the test to be able to find. This is the input teams get wrong most often, and it dominates everything else.
- Significance and power: conventionally 95% and 80%. Significance is how often you are willing to call a result real when it is not. Power is how often you would catch a real effect of that size, and 80% still means missing one real winner in five.
- Weekly traffic to the tested page: sessions that actually reach the template you are changing, which is usually a fraction of the number in the dashboard headline.
Work a real case. A store converting at 2.5% that wants to detect a 10% relative lift, at 95% significance and 80% power, two-sided, needs about 64,200 visitors per variant. Two variants means roughly 128,400 visitors in total. At 25,000 sessions a week reaching that page, the test runs six weeks. Not two.

The effect size is the lever, not the patience
Sample size scales with the square of the effect you are chasing, which has a consequence most teams underrate. Take the same 2.5% store. Looking for a 5% relative lift needs about 250,800 visitors per variant, or roughly twenty-one weeks at that traffic. Looking for 20% needs about 16,800, or two weeks. Same page, same statistics, same traffic, and a ten-fold difference in how long you wait.
So when a test comes back as a four-month commitment, the fix is almost never to run it longer or to relax the significance threshold. It is to test something bolder. A rebuilt product page beats a reworded button by a margin that shows up in the calendar, not just in the result. If you cannot make the change bigger, move the test to a template with more traffic. Both of those are real answers. Waiting is not.
You can run your own numbers through the sample size and duration calculator, which shows the same trade-off as a table rather than making you compute it four times.
Round up to whole weeks, always
Once you have a duration, round it up to a whole number of weeks. Buying behaviour is weekly. Tuesday traffic does not look like Saturday traffic, payday clusters, and a B2B-leaning store can see half its orders on two weekdays. Stopping a test after nine days means one group saw an extra Monday, and you have quietly compared two different populations.
The same logic applies at the short end. If the maths says four days, run seven. A test that finishes inside a week has not seen a full cycle, and a result built on one weekend is a result about that weekend.
Checking daily is not diligence, it is a second experiment
The sample size you calculated assumes you look once, at the end. Every time you check the dashboard and consider stopping, you take another shot at crossing the significance line by chance. Peek daily for a fortnight and the real false-positive rate is several times the 5% the tool reported. This is the single most common way a technically valid test produces a wrong answer, and no amount of extra traffic repairs it, because the problem is the stopping rule rather than the sample.
Two defences. Write the sample size down before launch and treat it as the finish line, or use a sequential testing method that is built to be monitored and adjusts its thresholds accordingly. What does not work is planning with a fixed-horizon calculator and then stopping whenever the chart looks good, which is what most teams do by default. It is also worth knowing what the number at the end does and does not claim: statistical significance is not the probability that your variant wins, and one minus the p-value is not a confidence level.
When the honest answer is that you cannot run it
Past about eight weeks, a test stops failing for statistical reasons and starts failing for ordinary ones. Your prices change. A supplier runs late. You start a sale, a competitor starts a bigger one, your traffic mix shifts as a campaign spends out. The two groups are no longer drawn from the same world, so whatever the result says at the end is not answering the question you asked at the start. Running it anyway costs the traffic and produces a number you should not act on, which is worse than no number at all.
Below roughly 30,000 sessions a month to a single template, most split tests will not finish in a useful window. That is not a reason to stop improving conversion. It is a reason to get the evidence somewhere else: session recordings, checkout drop-off, unmoderated testing with five people, and fixing what is plainly broken without demanding proof first. We have written up what to do instead at low traffic in full.
What to do before you brief the next one
Pull the baseline for the specific template, pull its weekly sessions, decide the smallest lift that would genuinely change what you ship, and calculate. If the answer is under four weeks, stop optimising the maths and go and argue about the hypothesis instead, because that is now the thing most likely to waste the month. If it is five to eight, look hard at whether the change is bold enough to be worth the calendar. If it is longer, say so out loud before anyone builds anything.
None of this requires a specialist. It requires doing the arithmetic before the work rather than after, which is most of what separates a testing programme that compounds from one that produces a lot of inconclusive tests and a quiet loss of faith. If you want the calculation done against your actual numbers, that is where our A/B testing work starts too, and we will tell you when the answer is that the test is not worth running. Otherwise, get in touch and bring the four numbers.



