If your site converts at 2% and you run 20,000 visitors through a test, you can only reliably detect enormous wins. Most teams are not running bad tests — they are running tests their traffic can never resolve. Here is how to work out what yours can.
There is a specific kind of disappointment that arrives about six weeks into a conversion optimisation programme. The tests have run, the results are inconclusive, and somebody asks whether testing works.
It works. What usually went wrong is arithmetic: the tests were asking a question the available traffic could never answer.
The uncomfortable arithmetic
The number of visitors a test needs depends on three things: your current conversion rate, the size of the improvement you want to detect, and how confident you want to be.
The critical relationship is that required sample size scales with roughly the inverse square of the effect size. Halve the improvement you want to detect and you need about four times the traffic. That single fact explains almost every failed testing programme.
Work it through with realistic numbers. Global e-commerce conversion rates average around 2.66%, with Shopify stores averaging nearer 1.40% and the top decile above 4.7%. Take a store at 2%:
- To detect a 50% relative lift — 2% to 3% — you need roughly 3,800 visitors per variant at 95% confidence and 80% power. Very achievable.
- To detect a 10% relative lift — 2% to 2.2% — you need about 80,000 per variant.
- To detect a 5% relative lift — 2% to 2.1% — you need in the region of 315,000 per variant.
Now notice what a 10% lift is worth. On a store doing $200,000 a month, it is $20,000 a month, forever. It is easily the most valuable result on that list, and it is the one most sites cannot detect.
The four ways teams fool themselves
Peeking. Watching a test daily and stopping it the moment it crosses 95% confidence is the most common error in the discipline, and it is worse than it sounds. Significance calculated repeatedly on accumulating data will cross the threshold by chance far more often than 5% of the time. Fix the sample size and the duration in advance, then look once. If you need to monitor continuously, use a method built for it — sequential testing or a Bayesian approach with a defined stopping rule — not a fixed-horizon calculator checked every morning.
Stopping on a winner too early. Related and even more expensive, because the early leader in a test is disproportionately likely to be a fluke that regresses. Effects measured at the moment of stopping are systematically overstated.
Running for less than a full business cycle. Weekday and weekend buyers differ, payday weeks differ, and a campaign landing mid-test changes the traffic mix. Run in whole weeks, minimum two, and note anything unusual that happened during the window.
Testing a dozen things at once and reporting whichever won. Every additional comparison raises the chance that something crosses 95% by accident. Twenty simultaneous variants will produce a winner from pure noise about as reliably as a coin does.
What to do when you do not have the traffic
Most sites do not, and this is where the discipline gets useful rather than academic.
Test radical changes, not refinements. If you can only resolve a 20% effect, do not test button colours. Test a fundamentally different page: a different offer, a different structure, a different first screen. Big swings produce big effects or clear failures, and both are information.
Move the test upstream. A form completion happens more often than a purchase; an add-to-cart happens more often still. Optimising a step with ten times the volume gets you an answer in a tenth of the time, and steps early in a funnel are where the largest leaks usually are anyway.
Use evidence that is not a test. Session recordings, funnel drop-off analysis, form field analytics, support tickets and five moderated user sessions will tell you more in a week than an underpowered test will in two months. The Baymard data is a free head start here: extra costs shown late (40%), delivery too slow (20%), card security concerns (19%), forced account creation (18%) and a checkout that is too long (17%) are the documented reasons people abandon. Check your own checkout against that list before you test anything.
Fix the things that do not need testing. A checkout with 11 form fields against a documented average of 11.3 and a practical minimum of 8 does not need an experiment; it needs three fields removed. Neither does a page failing Core Web Vitals, a broken mobile layout, or a form that fails silently on error. Ship those and spend your testing capacity on genuine uncertainty.
What a defensible test looks like
- A hypothesis with a reason. Not we think a shorter form will convert better, but form analytics show 40% of abandonments happen on the phone field, so removing it should raise completion.
- A primary metric decided in advance. One. Secondary metrics are for understanding what happened, never for declaring victory.
- A sample size and end date calculated before launch, from your real conversion rate and the smallest effect worth acting on.
- A minimum of two full weeks, in whole-week increments.
- One decision at the end: ship, discard, or iterate. Inconclusive is a legitimate result and should be recorded as one — a test that fails to find a large effect has told you the large effect is not there.
- A log entry either way. The compounding value of a testing programme is the archive of what did not work, which stops the same idea being re-proposed every eighteen months.
Questions we get asked
Is 95% confidence a rule? It is a convention, not a law. For a low-risk change on a page with limited traffic, 90% with a clear-eyed view of the risk is a defensible business decision. For a change that is expensive to reverse, be stricter.
Bayesian or frequentist? Bayesian methods report probability that a variant is better, which most stakeholders find easier to act on, and they handle continuous monitoring more gracefully. Neither approach creates statistical power out of traffic you do not have.
Can I test on paid traffic to get volume faster? Yes, and be aware you are then measuring the behaviour of paid visitors, which is not always the behaviour of your organic audience. It is a good way to accelerate a test whose result you intend to apply to paid landing pages in particular.
What about personalisation instead of testing? Same maths, harder. Every segment you personalise for divides your sample again. Personalisation earns its keep at volume and rarely below it.
The honest summary
What to record about every test
A testing programme compounds through its archive, not its wins. Six fields, kept in one place, are enough.
- The hypothesis and the evidence behind it, so a future team can see what you believed and why.
- The audience and the page, including device split — a result on mobile is not automatically a result on desktop.
- The planned sample size, duration and primary metric, recorded before launch. This is what stops a result being reinterpreted after the fact.
- The outcome, with the confidence interval, not just the point estimate. A 12% lift with an interval spanning zero to 25% is a very different fact from a 12% lift with an interval of 8% to 16%.
- Whether it shipped, and what happened afterwards. Post-launch performance often disagrees with the test, and that disagreement is worth knowing about.
- What you would test next. The most valuable output of most tests is a better question.
Teams that keep this record stop re-running the same experiment every eighteen months, which on a low-traffic site is a significant share of the available testing capacity.
Testing is not the point. Deciding well is the point, and testing is the most reliable tool for it when you have the traffic to use it. When you do not, the answer is not to run underpowered tests and read tea leaves — it is to use qualitative evidence, fix the known problems, and save experimentation for the questions where the answer genuinely is not knowable in advance.
That is how our conversion rate optimisation programmes are structured: research and known-issue remediation first, experimentation where it can actually resolve, and a documented record of both.
Looking for more? Browse all resources.











