A/B Testing
Show two versions to separate groups at the same time and measure the difference. The only method that establishes cause, not correlation — if you have the traffic and the discipline to run it right.
▸ Try the interactive toolThe framework
An A/B test is a controlled experiment in which two or more versions of a product element are shown to separate groups of users simultaneously, to establish a causal relationship between a change and an outcome. Version A (control) is the current experience; version B (variant) has the change. Random assignment and simultaneous exposure are what make the result causal rather than correlational.
A/B testing is the gold standard for proving a change caused an improvement — but only when run with discipline. It demands a clear hypothesis, enough traffic to reach statistical significance, a single well-defined change, and the patience not to peek and stop early. Run loosely, it produces false confidence: significance declared too soon, multiple metrics cherry-picked, or a result read into pure noise.
Random assignment (groups differ only by the change) + simultaneous exposure (same conditions) + a pre-defined success metric + statistical significance. Remove any one and you're back to correlation.
Try it yourself
What it prevents
Almost every other method shows correlation; A/B testing is the one that shows cause. But its rigour is fragile — small process slips turn a causal test into confident noise.
| The shortcut | What it costs | What it gives you instead |
|---|---|---|
| Correlation mistaken for cause | “We shipped X and the metric rose” ignores everything else that changed. | Random, simultaneous groups isolate the change as the only difference. |
| Peeking and stopping early | Calling a winner the moment it looks good inflates false positives. | A pre-set sample size and duration prevent premature, false conclusions. |
| Testing too much at once | Changing several things means you can't attribute the result. | One well-defined change keeps the result interpretable. |
| Cherry-picking metrics | Hunting through many metrics for any significant one. | A single pre-defined success metric prevents false discoveries. |
Step by step
Start from a clear hypothesis
State what you're changing, the effect you predict, and why — before you build the test. A test without a hypothesis is a fishing expedition (see Tool 13).
Change one thing, define one metric
Isolate a single change and a single primary success metric, set in advance. Multiple changes or roaming metrics make the result uninterpretable.
Calculate the sample size first
Determine how many users and how long you need for a trustworthy result before starting. This is what stops you from peeking and stopping early on noise.
Run simultaneously with random assignment
Expose both versions at the same time to randomly assigned groups, so the only systematic difference between them is the change itself.
Wait for significance, then decide
Run to the pre-set sample size; don't stop the moment it looks good. Read the result against the one pre-defined metric, honestly — including “no difference” as a valid outcome.
A short illustration
A team launched an A/B test and, three days in, the variant was winning by a wide margin. Excited, they declared victory and shipped it. Over the following weeks the metric drifted back to where it started — the early “win” had been noise that hadn't yet averaged out.
Had they calculated the required sample size in advance and waited for it, they'd have seen the effect was never real. Stopping early on an encouraging-but-premature result — “peeking” — is one of the most common ways A/B tests produce false positives, because random fluctuation looks like signal before enough data accumulates.
The experiment design + result
The deliverable is a pre-registered design (hypothesis, single change, primary metric, sample size) and an honestly-read result.
| Discipline | Why it matters |
|---|---|
| One change | So you can attribute the result |
| Pre-defined metric | Prevents cherry-picking a false win |
| Pre-set sample size | Prevents peeking and stopping on noise |
| Random, simultaneous groups | Makes the result causal, not correlational |
Why it matters
A/B testing is the only method that turns “the metric moved after we shipped” into “the change caused the metric to move.” That causal certainty is its unique gift — and it's entirely dependent on discipline.
Every shortcut attacks the causality. Peek and stop early, and random noise masquerades as a win. Change two things, and you can't say which mattered. Hunt through ten metrics, and one will look significant by chance alone. Each of these is tempting precisely because it gets you to a confident answer faster — which is the problem, because the confidence is fake. The teams that get real value from A/B testing treat the pre-commitments (one change, one metric, one sample size, decided up front) as non-negotiable, because those commitments are the only thing standing between a causal experiment and an expensively-disguised guess. And they accept “no significant difference” as a real, useful result rather than a failure to be massaged away.
Common mistakes
Calling a winner before the sample size is reached turns noise into a false win. Pre-set it and wait.
You can't attribute the result. Isolate one change per test.
Searching many metrics for any significant one manufactures false findings. Pre-define one.
A null result is real information, not a failure. Don't massage it into a win.
When not to use it
- Not enough traffic. A/B tests need volume to reach significance; with few users, prefer qualitative methods or fake doors.
- Big, strategic, one-way decisions. Some changes are too large or too rare to A/B test — use judgement and qualitative evidence.
- You haven't formed a hypothesis. Testing without a hypothesis is fishing. Frame it first (Tool 13).
Where this sits in the toolkit
Every A/B test starts from a precise, falsifiable hypothesis (Tool 13).
Analytics finds where a change might help; A/B testing proves whether it does (Tool 09).
A/B testing is the high-rigour end of the lean-experiment spectrum (Tool 15).
Significance, sample size, and power draw on the statistical foundations in Module 5.
Find the flaw in a tempting test
Imagine a test where the variant is winning big after two days. List the reasons you should not ship it yet.
Then name the one pre-commitment that would have protected you from the temptation to stop early.
If you reached “I should have calculated the sample size before starting,” you've understood why A/B testing's discipline matters more than its math — the process is what makes the result real.
New tools and AI deep-dives, occasionally.
No spam. Unsubscribe anytime.