Conversion

Why your A/B test lies: statistical significance and sample size

Five ways an A/B test crowns a winner that doesn't exist. What statistical significance means and how much traffic you need — with a sample size calculator.

Kaengrowth buddy team · 7 min read

Variant B is up 30% after three days. The tool glows green, the team celebrates, the change ships. Two months later the conversion rate is right where it was. Nobody knows why.

The reason is simple: that difference never existed. An A/B test doesn’t measure truth, it measures a sample — and a sample can lie in several dependable ways. This article describes them and shows how to avoid each.

What statistical significance actually says

Flip a coin ten times and seven heads is perfectly normal. The coin isn’t rigged — small samples just wobble. Website visitors behave the same way: two completely identical pages will show different conversion rates after a few days purely by chance.

Statistical significance answers one question: if there were no real difference between the variants, how often would we see a gap this large by chance alone? The usual 95% threshold means you accept a false alarm roughly one time in twenty.

What significance does not say:

  • that variant B is better “with 95% certainty”,
  • how big the difference really is,
  • that the difference will hold after you ship.

And above all: it only holds if the test was run by the rules. Most lies come from breaking them.

Five ways your A/B test lies to you

1. You peek, and stop when it looks good

The most common mistake. The gap between variants wobbles during a test — A leads, then B. If you check every day and stop the moment the tool first shows significance, you aren’t picking a winner. You’re picking the moment chance happened to push the numbers the right way.

The effect isn’t cosmetic: with daily checks, the share of false winners can climb from the planned 5% into the tens of percent.

The fix: set sample size and duration up front. You may look; you may only decide at the end.

2. The sample is too small — so wins look huge

An underpowered test has two properties. It usually misses real improvements. And when it does “catch” something, the result is almost always exaggerated — with a small sample, only results inflated by chance make it past the significance bar. That’s why so many tests report +40% and a fraction of it, or nothing, survives the rollout.

The fix: calculate the sample you need (calculator below) and distrust big wins from small numbers.

3. You track twenty metrics and one comes up

At a 95% threshold, one metric in twenty comes up “significant” by chance. If you comb through conversions, clicks, time on page and scroll depth after the test, each split by mobile, desktop, new and returning visitors, you will always find a winner. It means nothing.

The fix: one primary metric, chosen before launch. Treat segments and secondary metrics as a source of new hypotheses, not as proof.

4. The test didn’t run whole weeks

People behave differently on a Tuesday morning than on a Sunday evening. A test from Wednesday to Monday over-weights one part of the week. The first days distort things too: returning visitors react to novelty differently than they will a month from now.

The fix: always whole weeks, ideally at least two. And expect part of the early effect to be curiosity.

5. Broken measurement or an uneven split

If traffic should split fifty-fifty and one variant has noticeably more visitors, something is wrong — a redirect, caching, a blocked script, bots. The same goes for a conversion event that doesn’t fire reliably in one variant. A test like that can’t be repaired, only thrown away.

The fix: after day one, check the visitor ratio between variants and that conversions are recorded in both. How to verify step tracking is covered in the guide to funnel analysis in GA4.

How many visitors you need

Two numbers set the required sample: your current conversion rate and the smallest lift you want to detect reliably. The lower the conversion rate and the smaller the difference you’re after, the more people you need.

Sample size calculator

visitors per variant
visitors in total (2 variants)
weeks to run

Target conversion rate:

Two-sided test, 5% significance level, 80% power. Run whole weeks so the test covers weekdays and weekends.

Two examples worth remembering:

  • 2% conversion rate, looking for a 20% relative lift (to 2.4%): roughly 21,000 visitors per variant.
  • Same site, looking for a 10% lift: roughly 81,000 per variant. Half the effect costs about four times the sample.

An uncomfortable but useful consequence: small tweaks — a button color, one word in a headline — produce small differences, and most sites have no chance of measuring those.

What to do with low traffic

Test bigger changes. A new page structure, a different offer, a shorter form. A bigger intervention can produce a bigger difference, which a smaller sample can detect.

Test where the traffic is. A step most visitors pass through collects its sample faster than the last page of checkout.

Don’t stretch a test past eight weeks. Campaigns, seasons and the site itself change in that time. A test that needs a quarter is a signal to redesign it.

Lean on qualitative data. When five out of five people can’t find the submit button during user testing, you don’t need an A/B test for that. Obvious barriers get fixed, not tested — the recurring ones are listed in how to increase website conversion rate.

Tools and more on testing with modest traffic are covered in A/B testing tools after Google Optimize.

Pre-launch checklist

  1. We have a hypothesis shaped like “because we see X, we believe change Y will lift Z”.
  2. We have one primary metric.
  3. We know the required sample and, from it, the duration in whole weeks.
  4. The test fits within eight weeks; otherwise we redesign it.
  5. After day one we check the traffic split and conversion recording in both variants.
  6. We decide at the end — whatever the numbers look like along the way.

None of this is complicated. What’s hard is sticking to it every time, especially when a variant leads after three days and everyone wants to celebrate. Which is why this discipline is increasingly held by software: an AI CRO agent calculates the sample up front, doesn’t stop the test early, and tells you when a change simply can’t be measured on your traffic.

Frequently asked questions

What does statistical significance mean in an A/B test?

It tells you how unlikely it would be to measure a gap this large between variants if there were no real difference. A 95% threshold means you accept a false alarm roughly one time in twenty. It does not tell you how big the difference is, or that the variant is better with 95% certainty.

How many visitors do I need for an A/B test?

It depends on your current conversion rate and the smallest lift you want to detect. At a 2% conversion rate and a 20% relative lift, it’s roughly 21,000 visitors per variant. Halving the lift you look for roughly quadruples the sample.

How long should an A/B test run?

Until it collects the sample you calculated up front, and always in whole weeks so it covers weekdays and weekends. In practice two to eight weeks. A test that would need longer is better redesigned.

What if I don’t have enough traffic for A/B tests?

Test bigger changes that can produce bigger differences, test steps with more traffic, and for smaller tweaks rely on qualitative research and fixing obvious barriers instead of a test that has no chance of proving anything.


The takeaway

An A/B test doesn’t lie out of malice. It lies when you ask before it has anything to say, or when you ask twenty times and pick the answer you like. A sample set in advance, one metric, whole weeks and a decision at the end — that’s enough for the numbers you change your site by to mean something.

If you’d rather have someone else hold that discipline, Kaen is an AI agent for conversion rate optimization: he designs the test with the sample it needs, measures it, and calls the result only when there’s a result to call. Nothing ships without your approval.

Keep reading

All articles →