Module
Back to courses
Product ManagementAdvanced

Design and analyze an A/B test

Decide whether to test, write the protocol, size it, analyze without fooling yourself, and decide.

48 steps~3 hrLevel: Advanced · Regular practice: you already use the tools or work on this topic.

Know when an A/B test is the right method, design it so that it really answers your question, then read its result without fooling yourself. You learn to recognize when not to test (too little traffic, B2B with small numbers) and the alternatives, to write a hypothesis, a primary metric and guardrails, to choose the randomization unit, to compute a sample size and a duration, to interpret a p-value and a confidence interval correctly, to avoid the pitfalls (peeking, multiple comparisons, novelty effect, sample ratio mismatch, Simpson's paradox), to place Bayesian and frequentist approaches, to choose a tool, then to decide and document. You finish with the protocol of an experiment on your product, its analysis and a decision memo.

What you will be able to do

  • Decide whether an A/B test is the right method (volume, B2B, reversibility, time, ethics) and choose a suitable alternative otherwise.
  • Write a testable hypothesis, choose a sensitive, attributable primary metric, guardrails and the exposure event to log.
  • Choose the randomization unit and prevent interference between groups, including in B2B.
  • Compute a sample size and a duration from the baseline rate, the minimum detectable effect, the significance level and the power.
  • Analyze a result (p-value, confidence interval, practical significance, guardrails) and read a tool's Bayesian or frequentist output.
  • Detect and avoid analysis pitfalls (peeking, multiple comparisons, novelty effect, sample ratio mismatch, Simpson's paradox).
  • Decide based on rules written before the test and document the experiment in a reusable memo.

Prerequisites

  • Have already worked on a digital product in production, with usage data
  • Be able to compute a percentage and read a table; no statistics training is required, calculations are detailed step by step
  • Recommended course before this one: Instrument a product and use data to make decisions, which covers the tracking plan, data quality and the pitfalls of observational analysis.
  • To arbitrate between several bets before testing them, see the course Prioritize and build an outcome-based roadmap.

Syllabus

What will I be able to do by the end of this course, and in what order?

  1. Objective · By the end of this overview, you will know what you are going to produce (the protocol of an experiment on your product, its analysis and a decision memo), which two questions Atelio will settle throughout the course and which four modules take you there.

What does an A/B test prove that an analysis does not, and when is it better to do without one?

  1. Objective · By the end of this lesson, you will be able to explain why a pre/post comparison or a comparison between users and non-users does not measure the effect of a change, what randomization brings, and to use the vocabulary of an A/B test correctly (control, variant, unit, exposure).

  2. Objective · By the end of this lesson, you will be able to estimate whether the available volume allows an A/B test (minimum detectable effect), to recognize the other cases where a test is not the right method, and to choose a suitable alternative, especially in B2B.

What must be settled before launching the test so that its result is readable and credible?

  1. Objective · By the end of this lesson, you will be able to write a testable, quantified hypothesis, choose a sensitive and attributable primary metric, guardrails with their threshold, and specify the exposure event to log.

  2. Objective · By the end of this lesson, you will be able to choose the unit to randomize (user, account, session, business object, cluster) by weighing consistency of experience, interference and number of units, and to spot the B2B and marketplace cases where groups influence each other.

  3. Objective · By the end of this lesson, you will be able to compute by hand and with a free calculator the sample size of a test on a rate, from the baseline rate, the minimum detectable effect, the significance level and the power, to derive a duration in full weeks, and to know what to do if it is out of reach.

How do I read a test result without over-interpreting it or missing a problem?

  1. Objective · By the end of this lesson, you will be able to compute the gap between two rates, its 95% confidence interval and its p-value, to interpret them correctly (and spot wrong interpretations), and to judge whether the effect is large enough to matter, guardrails included.

  2. Objective · By the end of this lesson, you will be able to spot peeking, multiple comparisons, a novelty effect, a sample ratio mismatch or a Simpson's paradox in a test or its report, to check a mismatch by calculation, and to choose the right countermeasure.

  3. Objective · By the end of this lesson, you will be able to explain to a team the difference between a frequentist and a Bayesian reading of a test, to read a tool's output in both frameworks, and to choose an experimentation tool suited to your constraints, with a free path.

How do I turn a result into a decision, and the test into lasting learning?

  1. Objective · By the end of this lesson, you will be able to apply a decision rule written before the test to a conclusive, negative or inconclusive result, to plan what comes next (rollout, long-term holdout, new iteration), and to write a reusable experiment memo.

  2. Objective · By the end of this lesson, you will have written the protocol of an experiment on your product (or a reasoned alternative plan), analyzed a real result or the practice dataset provided, written the decision memo, self-assessed the whole with a grid, and you will leave with an exit kit and a 7-day and 30-day plan.