PM Mapped
Home / Module 3 / A/B Testing
12
MODULE 3 · QUANTITATIVE · TOOL 12

A/B Testing

Show two versions to separate groups at the same time and measure the difference. The only method that establishes cause, not correlation — if you have the traffic and the discipline to run it right.

▸ Try the interactive tool
SolvesShipping on opinion and mistaking random noise for a “win.”
Category · Quantitative Research Complexity · Mid–Advanced Time to apply · 1–4 weeks per test Pairs with · Testable Hypotheses
A WHAT IT IS

The framework

An A/B test is a controlled experiment in which two or more versions of a product element are shown to separate groups of users simultaneously, to establish a causal relationship between a change and an outcome. Version A (control) is the current experience; version B (variant) has the change. Random assignment and simultaneous exposure are what make the result causal rather than correlational.

A/B testing is the gold standard for proving a change caused an improvement — but only when run with discipline. It demands a clear hypothesis, enough traffic to reach statistical significance, a single well-defined change, and the patience not to peek and stop early. Run loosely, it produces false confidence: significance declared too soon, multiple metrics cherry-picked, or a result read into pure noise.

WHAT MAKES IT CAUSAL

Random assignment (groups differ only by the change) + simultaneous exposure (same conditions) + a pre-defined success metric + statistical significance. Remove any one and you're back to correlation.

TRY IT

Try it yourself

B WHY IT MATTERS

What it prevents

Almost every other method shows correlation; A/B testing is the one that shows cause. But its rigour is fragile — small process slips turn a causal test into confident noise.

The shortcutWhat it costsWhat it gives you instead
Correlation mistaken for cause“We shipped X and the metric rose” ignores everything else that changed.Random, simultaneous groups isolate the change as the only difference.
Peeking and stopping earlyCalling a winner the moment it looks good inflates false positives.A pre-set sample size and duration prevent premature, false conclusions.
Testing too much at onceChanging several things means you can't attribute the result.One well-defined change keeps the result interpretable.
Cherry-picking metricsHunting through many metrics for any significant one.A single pre-defined success metric prevents false discoveries.
C HOW TO RUN IT

Step by step

1

Start from a clear hypothesis

State what you're changing, the effect you predict, and why — before you build the test. A test without a hypothesis is a fishing expedition (see Tool 13).

2

Change one thing, define one metric

Isolate a single change and a single primary success metric, set in advance. Multiple changes or roaming metrics make the result uninterpretable.

3

Calculate the sample size first

Determine how many users and how long you need for a trustworthy result before starting. This is what stops you from peeking and stopping early on noise.

4

Run simultaneously with random assignment

Expose both versions at the same time to randomly assigned groups, so the only systematic difference between them is the change itself.

5

Wait for significance, then decide

Run to the pre-set sample size; don't stop the moment it looks good. Read the result against the one pre-defined metric, honestly — including “no difference” as a valid outcome.

D IN PRACTICE

A short illustration

IN PRACTICEthe peeking trap

A team launched an A/B test and, three days in, the variant was winning by a wide margin. Excited, they declared victory and shipped it. Over the following weeks the metric drifted back to where it started — the early “win” had been noise that hadn't yet averaged out.

Had they calculated the required sample size in advance and waited for it, they'd have seen the effect was never real. Stopping early on an encouraging-but-premature result — “peeking” — is one of the most common ways A/B tests produce false positives, because random fluctuation looks like signal before enough data accumulates.

The lesson: the rigour is the method. Random assignment makes an A/B test causal, but peeking, multiple metrics, and stopping early quietly destroy that rigour and turn a causal instrument into a confident-noise generator.
E THE ARTIFACT

The experiment design + result

The deliverable is a pre-registered design (hypothesis, single change, primary metric, sample size) and an honestly-read result.

DisciplineWhy it matters
One changeSo you can attribute the result
Pre-defined metricPrevents cherry-picking a false win
Pre-set sample sizePrevents peeking and stopping on noise
Random, simultaneous groupsMakes the result causal, not correlational
F THE SO-WHAT

Why it matters

THE KEY INSIGHT

A/B testing is the only method that turns “the metric moved after we shipped” into “the change caused the metric to move.” That causal certainty is its unique gift — and it's entirely dependent on discipline.

Every shortcut attacks the causality. Peek and stop early, and random noise masquerades as a win. Change two things, and you can't say which mattered. Hunt through ten metrics, and one will look significant by chance alone. Each of these is tempting precisely because it gets you to a confident answer faster — which is the problem, because the confidence is fake. The teams that get real value from A/B testing treat the pre-commitments (one change, one metric, one sample size, decided up front) as non-negotiable, because those commitments are the only thing standing between a causal experiment and an expensively-disguised guess. And they accept “no significant difference” as a real, useful result rather than a failure to be massaged away.

G MISTAKES & LIMITS

Common mistakes

Peeking and stopping early

Calling a winner before the sample size is reached turns noise into a false win. Pre-set it and wait.

Testing multiple changes at once

You can't attribute the result. Isolate one change per test.

Cherry-picking metrics

Searching many metrics for any significant one manufactures false findings. Pre-define one.

Ignoring “no difference” results

A null result is real information, not a failure. Don't massage it into a win.

When not to use it

H CONNECTS TO

Where this sits in the toolkit

Requires → Testable Hypotheses

Every A/B test starts from a precise, falsifiable hypothesis (Tool 13).

Aimed by → Behavioural Analytics

Analytics finds where a change might help; A/B testing proves whether it does (Tool 09).

Part of → the experiment toolkit

A/B testing is the high-rigour end of the lean-experiment spectrum (Tool 15).

Underpinned by → statistics

Significance, sample size, and power draw on the statistical foundations in Module 5.

TRY IT YOURSELF

Find the flaw in a tempting test

Imagine a test where the variant is winning big after two days. List the reasons you should not ship it yet.

Then name the one pre-commitment that would have protected you from the temptation to stop early.

If you reached “I should have calculated the sample size before starting,” you've understood why A/B testing's discipline matters more than its math — the process is what makes the result real.