A pre-experiment routine that eliminates most experimentation errors. Its central insight: most bad experiments fail not from bad analysis after the fact, but from bad setup before they start.
▸ Try the interactive toolThe experiment design checklist is a pre-experiment routine that, if followed, eliminates the vast majority of experimentation errors. Its central insight is that most bad experiments are not bad because of analysis errors after the fact — they're bad because of setup errors before the experiment ever ran. The cheapest place to fix an experiment is in its design.
The checklist forces you to settle the things that, left undecided, doom an experiment: a clear, falsifiable hypothesis; a single primary metric (and guardrail metrics); a pre-calculated sample size and duration; a pre-set significance threshold; a plan to avoid peeking; and a check for confounds (novelty effects, seasonality, contamination between groups). Run through it before every experiment and you prevent the errors that no amount of clever post-hoc analysis can rescue — because an experiment broken in design can't be fixed in analysis.
Falsifiable hypothesis · single primary metric + guardrails · pre-calculated sample & duration · pre-set significance threshold · no-peeking plan · confound check (novelty, seasonality, contamination).
Most experimentation failures are baked in before the experiment runs — so the leverage is entirely in the setup, which a checklist makes systematic.
| The shortcut | What it costs | What it gives you instead |
|---|---|---|
| Setup errors | An experiment broken in design can't be saved in analysis. | The checklist catches the fatal setup errors before running. |
| No clear hypothesis | Without a falsifiable prediction, any result is rationalisable. | The checklist forces a clear, falsifiable hypothesis. |
| Multiple primary metrics | Many metrics means something always 'wins' by chance. | A single primary metric prevents cherry-picking the winner. |
| Unplanned sample/duration | Stopping ad hoc invites peeking and false positives. | Pre-calculated sample and duration remove the temptation. |
Write the prediction clearly enough that the experiment could prove it wrong. A vague hypothesis produces a result you can interpret any way you like.
Pick the single metric that decides the experiment, and guardrail metrics to ensure you don't win on the primary by harming something else. Multiple primary metrics guarantee a spurious 'winner.'
Using the statistical fundamentals (Tool 17), determine how many users and how long, before running. This is what removes the temptation to stop early.
Fix the p-value threshold in advance and commit to not stopping early on favourable peeks. Decide these before any data exists.
Before launching, check for things that could contaminate the result: novelty effects, seasonality, overlap or contamination between control and variant. Catching these in design saves a worthless experiment.
A team kept running experiments that produced ambiguous or untrustworthy results, then spent enormous effort trying to rescue them with clever analysis after the fact. The analysis could never fully fix what the setup had broken — vague hypotheses, multiple competing metrics, no pre-set sample size.
Adopting a design checklist moved the effort to where it actually paid: before running. Each experiment now started with a falsifiable hypothesis, one primary metric, a pre-calculated sample and duration, a fixed significance threshold, and a confound check. The results became clean and trustworthy — not because the analysis got smarter, but because the experiments were no longer broken before they began.
The deliverable is a completed checklist before every experiment — hypothesis, metric, sample, threshold, no-peeking plan, confound check.
| Decide before running | Prevents |
|---|---|
| Falsifiable hypothesis | Rationalising any result |
| Single primary metric | Cherry-picking a winner |
| Sample size & duration | Peeking, under-powering |
| Significance threshold | Moving the goalposts |
| Confound check | Novelty/seasonality fooling you |
The checklist embodies a simple, powerful truth: experiments succeed or fail in their design, not their analysis. The leverage is entirely up front, so a few minutes of disciplined setup prevents the errors that hours of clever analysis can't.
This reframes where a PM should spend their experimentation effort. The instinct is to focus on analysing results — but a broken experiment yields untrustworthy results no matter how sophisticated the analysis, while a well-designed one is often trivial to read. Every item on the checklist closes off a specific, common failure: a vague hypothesis lets you rationalise any outcome; multiple primary metrics guarantee a spurious winner; no pre-set sample invites peeking; unchecked confounds (novelty, seasonality) masquerade as effects. By forcing these decisions before any data exists — when there's no result to bias them — the checklist removes the temptations and errors that corrupt experiments from the inside. It's the operational front end of the statistical foundations, turning principles into a routine anyone can follow.
A design-broken experiment can't be rescued after. Get the setup right first.
An un-falsifiable prediction lets you rationalise anything. State it clearly.
Something always wins by chance. Choose one primary metric plus guardrails.
Novelty and seasonality masquerade as effects. Check before launching.
The checklist turns the fundamentals (Tool 17) into a repeatable routine.
Following the checklist heads off most of the six mistakes (Tool 19).
The falsifiable-hypothesis item is the discipline from discovery (Module 3, Tool 13).
It's the setup discipline for every A/B test (Module 3, Tool 12).
Take an experiment idea. Walk it through the checklist: is the hypothesis falsifiable? Is there one primary metric? Have you a pre-calculated sample and a fixed threshold? Any confounds?
Note which items you'd have skipped if working from instinct — those are where experiments usually break.
If the checklist surfaces a setup flaw you'd otherwise have discovered only in a confusing result, you've seen why experiments are won or lost in design, not analysis.