Show two versions to separate groups at the same time and measure the difference. The only method that establishes cause, not correlation — if you have the traffic and the discipline to run it right.
▸ Try the interactive toolAn A/B test is a controlled experiment in which two or more versions of a product element are shown to separate groups of users simultaneously, to establish a causal relationship between a change and an outcome. Version A (control) is the current experience; version B (variant) has the change. Random assignment and simultaneous exposure are what make the result causal rather than correlational.
A/B testing is the gold standard for proving a change caused an improvement — but only when run with discipline. It demands a clear hypothesis, enough traffic to reach statistical significance, a single well-defined change, and the patience not to peek and stop early. Run loosely, it produces false confidence: significance declared too soon, multiple metrics cherry-picked, or a result read into pure noise.
Random assignment (groups differ only by the change) + simultaneous exposure (same conditions) + a pre-defined success metric + statistical significance. Remove any one and you're back to correlation.
Almost every other method shows correlation; A/B testing is the one that shows cause. But its rigour is fragile — small process slips turn a causal test into confident noise.
| The shortcut | What it costs | What it gives you instead |
|---|---|---|
| Correlation mistaken for cause | “We shipped X and the metric rose” ignores everything else that changed. | Random, simultaneous groups isolate the change as the only difference. |
| Peeking and stopping early | Calling a winner the moment it looks good inflates false positives. | A pre-set sample size and duration prevent premature, false conclusions. |
| Testing too much at once | Changing several things means you can't attribute the result. | One well-defined change keeps the result interpretable. |
| Cherry-picking metrics | Hunting through many metrics for any significant one. | A single pre-defined success metric prevents false discoveries. |
State what you're changing, the effect you predict, and why — before you build the test. A test without a hypothesis is a fishing expedition (see Tool 13).
Isolate a single change and a single primary success metric, set in advance. Multiple changes or roaming metrics make the result uninterpretable.
Determine how many users and how long you need for a trustworthy result before starting. This is what stops you from peeking and stopping early on noise.
Expose both versions at the same time to randomly assigned groups, so the only systematic difference between them is the change itself.
Run to the pre-set sample size; don't stop the moment it looks good. Read the result against the one pre-defined metric, honestly — including “no difference” as a valid outcome.
A team launched an A/B test and, three days in, the variant was winning by a wide margin. Excited, they declared victory and shipped it. Over the following weeks the metric drifted back to where it started — the early “win” had been noise that hadn't yet averaged out.
Had they calculated the required sample size in advance and waited for it, they'd have seen the effect was never real. Stopping early on an encouraging-but-premature result — “peeking” — is one of the most common ways A/B tests produce false positives, because random fluctuation looks like signal before enough data accumulates.
The deliverable is a pre-registered design (hypothesis, single change, primary metric, sample size) and an honestly-read result.
| Discipline | Why it matters |
|---|---|
| One change | So you can attribute the result |
| Pre-defined metric | Prevents cherry-picking a false win |
| Pre-set sample size | Prevents peeking and stopping on noise |
| Random, simultaneous groups | Makes the result causal, not correlational |
A/B testing is the only method that turns “the metric moved after we shipped” into “the change caused the metric to move.” That causal certainty is its unique gift — and it's entirely dependent on discipline.
Every shortcut attacks the causality. Peek and stop early, and random noise masquerades as a win. Change two things, and you can't say which mattered. Hunt through ten metrics, and one will look significant by chance alone. Each of these is tempting precisely because it gets you to a confident answer faster — which is the problem, because the confidence is fake. The teams that get real value from A/B testing treat the pre-commitments (one change, one metric, one sample size, decided up front) as non-negotiable, because those commitments are the only thing standing between a causal experiment and an expensively-disguised guess. And they accept “no significant difference” as a real, useful result rather than a failure to be massaged away.
Calling a winner before the sample size is reached turns noise into a false win. Pre-set it and wait.
You can't attribute the result. Isolate one change per test.
Searching many metrics for any significant one manufactures false findings. Pre-define one.
A null result is real information, not a failure. Don't massage it into a win.
Every A/B test starts from a precise, falsifiable hypothesis (Tool 13).
Analytics finds where a change might help; A/B testing proves whether it does (Tool 09).
A/B testing is the high-rigour end of the lean-experiment spectrum (Tool 15).
Significance, sample size, and power draw on the statistical foundations in Module 5.
Imagine a test where the variant is winning big after two days. List the reasons you should not ship it yet.
Then name the one pre-commitment that would have protected you from the temptation to stop early.
If you reached “I should have calculated the sample size before starting,” you've understood why A/B testing's discipline matters more than its math — the process is what makes the result real.