PM Mapped
Home / Module 5 / Common Experimentation Mistakes
19
MODULE 5 · EXPERIMENTATION · TOOL 19

Common Experimentation Mistakes

Six experimentation failures that are common, expensive, and entirely preventable. Where the design checklist is the proactive discipline, this is the diagnostic awareness of how experiments go wrong.

▸ Try the interactive tool
SolvesPeeking, p-hacking, and shipping false wins.
Category · Experimentation Complexity · Mid Time to apply · Reference Pairs with · The Experiment Design Checklist
A WHAT IT IS

The framework

This tool catalogues six experimentation failures that are common, expensive, and entirely preventable. Where the design checklist (Tool 18) is the proactive discipline of setting experiments up well, this catalogue is the diagnostic awareness of how experiments go wrong — the failure modes to recognise in your own and others' work.

The six are deadly because each produces a confident, wrong conclusion that looks legitimate from the inside. Peeking and novelty effects manufacture false wins; underpowered tests and sample ratio mismatch produce meaningless results that masquerade as findings; multiple comparisons guarantee a spurious winner; ignoring guardrails lets you 'win' while doing damage. Recognising them is most of the defence — and the running of an experiment, when the temptation to peek and stop early is strongest, is exactly when this awareness matters most.

WHY THESE ARE DANGEROUS

Each produces a confident, wrong conclusion that looks legitimate from the inside. The temptation peaks while the experiment runs — especially the urge to peek and stop early. Awareness is most of the defence.

TRY IT

Try it yourself

B WHY IT MATTERS

What it prevents

These mistakes are dangerous precisely because the flawed experiment looks successful to the people running it — the false win feels like a real one.

The shortcutWhat it costsWhat it gives you instead
Peeking & early stoppingStopping on a good peek inflates false positives invisibly.A pre-set sample and no-peeking rule prevent it.
Novelty effectA week-one win reverses once the newness fades.A minimum duration lets novelty wash out.
Multiple comparisonsTest enough things and one 'wins' by pure chance.One primary metric and correction for multiple comparisons.
Ignoring guardrailsWinning the primary while harming retention or revenue.Guardrail metrics catch the hidden damage.
C THE BREAKDOWN

The six mistakes

Each is common, expensive, and preventable — and each produces a confident conclusion that looks right from the inside, which is what makes them dangerous.

MistakeWhat goes wrongHow to prevent it
Peeking & early stoppingStopping when results look good inflates false positives — each peek is a new testPre-set sample size and duration; never stop early on a favourable peek
Novelty effectUsers engage more with anything new; a week-one win can reverse by week threeRun a minimum duration (e.g. two weeks) so the novelty fades
Underpowered testsToo small a sample to detect the effect; 'no result' is meaninglessCalculate required sample size before running (Tool 17)
Multiple comparisonsTesting many metrics/variants means something 'wins' by chanceOne primary metric; correct for multiple comparisons
Ignoring guardrailsWinning the primary metric while quietly harming anotherTrack guardrail metrics alongside the primary
Sample ratio mismatchTraffic isn't splitting as intended — the experiment is brokenCheck the actual split; investigate any imbalance before trusting results
D IN PRACTICE

A short illustration

IN PRACTICEthe temptation while it runs

An experiment had been live only a few days when the early activation numbers looked good for the variant, and the pull to call it and ship was immediate. That pull — peeking and early stopping — is the single most common cause of false positives, because every time you look and consider stopping, you're effectively running a new test and inflating the odds of a fluke.

Two of the six mistakes were in play at once. Even if the early numbers held, they might be a novelty effect — users engage more with anything new for a week or two, and a variant winning early can normalise or reverse by week three. The discipline that defeats both is the same: a pre-set minimum duration and sample, run without peeking, so the result reflects sustained value rather than early shine or random noise.

The lesson: the experimentation mistakes are most tempting precisely while an experiment runs and early numbers look good. Recognising peeking and novelty effects for what they are — manufacturers of false wins — is what lets a team hold the discipline when the pull to stop early is strongest.
E THE ARTIFACT

The mistakes checklist

The deliverable is using this catalogue diagnostically — before trusting any experiment, checking it against the six failure modes.

If you're tempted to…Recognise it as…
Stop early on a good resultPeeking — run to the planned sample
Trust a week-one winNovelty — wait out the minimum duration
Celebrate a small-sample resultUnderpowered — it may be noise
Pick the metric that wonMultiple comparisons — one primary only
F THE SO-WHAT

Why it matters

THE KEY INSIGHT

These six mistakes share one trait that makes them lethal: each produces a result that looks like success from the inside. The flawed experiment doesn't feel flawed — the false win feels exactly like a real one, which is why awareness is the primary defence.

The most important practical insight is when the danger peaks: while the experiment is running and early numbers look favourable. That's the moment the pull to peek, stop early, and ship is strongest — and it's the moment the discipline matters most. Peeking inflates false positives because each look is effectively another test; novelty effects make early wins evaporate as the newness fades; both are defeated by a pre-set minimum duration and sample run without interruption. The other mistakes — underpowering, multiple comparisons, ignored guardrails, sample ratio mismatch — round out the failure modes that turn experiments into confident, wrong conclusions. This catalogue is the diagnostic counterpart to the design checklist: the checklist prevents these in setup, this recognises them when the temptation arrives.

G MISTAKES & LIMITS

Common mistakes

Treating awareness as enough

Knowing the mistakes doesn't stop them mid-experiment. Pre-commit to the disciplines that prevent each.

Peeking when numbers look good

The strongest temptation and the most common false-positive cause. Run to the planned sample.

Trusting early wins

Novelty fades. Wait out the minimum duration before believing a result.

Winning the primary, ignoring guardrails

A 'win' that harms another metric isn't a win. Watch the guardrails.

When not to use it

H CONNECTS TO

Where this sits in the toolkit

Counterpart to → the Design Checklist

The checklist (Tool 18) prevents these proactively; this recognises them diagnostically.

Rooted in → Statistical Foundations

Most of the six are statistical failures the fundamentals (Tool 17) explain.

Pairs with → Metrics Anti-Patterns

Like Tool 27, this is a catalogue of failures to recognise and avoid.

Protects → A/B Testing

Avoiding these keeps A/B tests (Module 3, Tool 12) trustworthy.

TRY IT YOURSELF

Catch yourself peeking

Imagine an experiment three days in, with the variant looking good. Write down the argument for stopping and shipping now — then identify which two of the six mistakes that argument represents.

Describe the single discipline that defeats both.

If you spotted peeking and the novelty effect — and prescribed a pre-set minimum duration run without peeking — you've learned to recognise the false wins exactly when they're most tempting.