PM Mapped
Home / Module 5 / The Experiment Design Checklist
18
MODULE 5 · EXPERIMENTATION · TOOL 18

The Experiment Design Checklist

A pre-experiment routine that eliminates most experimentation errors. Its central insight: most bad experiments fail not from bad analysis after the fact, but from bad setup before they start.

▸ Try the interactive tool
SolvesExperiments that produce noise you mistake for signal.
Category · Experimentation Complexity · Mid Time to apply · Before each experiment Pairs with · Statistical Foundations
A WHAT IT IS

The framework

The experiment design checklist is a pre-experiment routine that, if followed, eliminates the vast majority of experimentation errors. Its central insight is that most bad experiments are not bad because of analysis errors after the fact — they're bad because of setup errors before the experiment ever ran. The cheapest place to fix an experiment is in its design.

The checklist forces you to settle the things that, left undecided, doom an experiment: a clear, falsifiable hypothesis; a single primary metric (and guardrail metrics); a pre-calculated sample size and duration; a pre-set significance threshold; a plan to avoid peeking; and a check for confounds (novelty effects, seasonality, contamination between groups). Run through it before every experiment and you prevent the errors that no amount of clever post-hoc analysis can rescue — because an experiment broken in design can't be fixed in analysis.

THE CHECKLIST (before running)

Falsifiable hypothesis · single primary metric + guardrails · pre-calculated sample & duration · pre-set significance threshold · no-peeking plan · confound check (novelty, seasonality, contamination).

TRY IT

Try it yourself

B WHY IT MATTERS

What it prevents

Most experimentation failures are baked in before the experiment runs — so the leverage is entirely in the setup, which a checklist makes systematic.

The shortcutWhat it costsWhat it gives you instead
Setup errorsAn experiment broken in design can't be saved in analysis.The checklist catches the fatal setup errors before running.
No clear hypothesisWithout a falsifiable prediction, any result is rationalisable.The checklist forces a clear, falsifiable hypothesis.
Multiple primary metricsMany metrics means something always 'wins' by chance.A single primary metric prevents cherry-picking the winner.
Unplanned sample/durationStopping ad hoc invites peeking and false positives.Pre-calculated sample and duration remove the temptation.
C HOW TO RUN IT

Step by step

1

State a falsifiable hypothesis

Write the prediction clearly enough that the experiment could prove it wrong. A vague hypothesis produces a result you can interpret any way you like.

2

Choose one primary metric (plus guardrails)

Pick the single metric that decides the experiment, and guardrail metrics to ensure you don't win on the primary by harming something else. Multiple primary metrics guarantee a spurious 'winner.'

3

Pre-calculate sample size and duration

Using the statistical fundamentals (Tool 17), determine how many users and how long, before running. This is what removes the temptation to stop early.

4

Set the significance threshold and no-peeking plan

Fix the p-value threshold in advance and commit to not stopping early on favourable peeks. Decide these before any data exists.

5

Check for confounds

Before launching, check for things that could contaminate the result: novelty effects, seasonality, overlap or contamination between control and variant. Catching these in design saves a worthless experiment.

D IN PRACTICE

A short illustration

IN PRACTICEfixing it in design, not analysis

A team kept running experiments that produced ambiguous or untrustworthy results, then spent enormous effort trying to rescue them with clever analysis after the fact. The analysis could never fully fix what the setup had broken — vague hypotheses, multiple competing metrics, no pre-set sample size.

Adopting a design checklist moved the effort to where it actually paid: before running. Each experiment now started with a falsifiable hypothesis, one primary metric, a pre-calculated sample and duration, a fixed significance threshold, and a confound check. The results became clean and trustworthy — not because the analysis got smarter, but because the experiments were no longer broken before they began.

The lesson: the cheapest and only reliable place to fix an experiment is its design. Most experimentation failures are setup errors that no post-hoc analysis can rescue — which is exactly why a pre-experiment checklist eliminates the majority of them.
E THE ARTIFACT

The pre-experiment checklist

The deliverable is a completed checklist before every experiment — hypothesis, metric, sample, threshold, no-peeking plan, confound check.

Decide before runningPrevents
Falsifiable hypothesisRationalising any result
Single primary metricCherry-picking a winner
Sample size & durationPeeking, under-powering
Significance thresholdMoving the goalposts
Confound checkNovelty/seasonality fooling you
F THE SO-WHAT

Why it matters

THE KEY INSIGHT

The checklist embodies a simple, powerful truth: experiments succeed or fail in their design, not their analysis. The leverage is entirely up front, so a few minutes of disciplined setup prevents the errors that hours of clever analysis can't.

This reframes where a PM should spend their experimentation effort. The instinct is to focus on analysing results — but a broken experiment yields untrustworthy results no matter how sophisticated the analysis, while a well-designed one is often trivial to read. Every item on the checklist closes off a specific, common failure: a vague hypothesis lets you rationalise any outcome; multiple primary metrics guarantee a spurious winner; no pre-set sample invites peeking; unchecked confounds (novelty, seasonality) masquerade as effects. By forcing these decisions before any data exists — when there's no result to bias them — the checklist removes the temptations and errors that corrupt experiments from the inside. It's the operational front end of the statistical foundations, turning principles into a routine anyone can follow.

G MISTAKES & LIMITS

Common mistakes

Fixing experiments in analysis

A design-broken experiment can't be rescued after. Get the setup right first.

Vague hypotheses

An un-falsifiable prediction lets you rationalise anything. State it clearly.

Multiple primary metrics

Something always wins by chance. Choose one primary metric plus guardrails.

Skipping the confound check

Novelty and seasonality masquerade as effects. Check before launching.

When not to use it

H CONNECTS TO

Where this sits in the toolkit

Operationalises → Statistical Foundations

The checklist turns the fundamentals (Tool 17) into a repeatable routine.

Prevents → Common Experimentation Mistakes

Following the checklist heads off most of the six mistakes (Tool 19).

Built on → Testable Hypotheses

The falsifiable-hypothesis item is the discipline from discovery (Module 3, Tool 13).

Applies to → A/B Testing

It's the setup discipline for every A/B test (Module 3, Tool 12).

TRY IT YOURSELF

Run the checklist on an experiment

Take an experiment idea. Walk it through the checklist: is the hypothesis falsifiable? Is there one primary metric? Have you a pre-calculated sample and a fixed threshold? Any confounds?

Note which items you'd have skipped if working from instinct — those are where experiments usually break.

If the checklist surfaces a setup flaw you'd otherwise have discovered only in a confusing result, you've seen why experiments are won or lost in design, not analysis.