Module
Module 4 of 5Lesson 1 of 3~18 min

Read a p-value and a confidence interval

The p-value is one of the most misread numbers in product reviews, and those misreadings lead to unjustified rollouts. This lesson teaches you to analyze an A/B test by reading the difference, its confidence interval and its p-value without misinterpretation. You then know whether the effect deserves a rollout, guardrails included.

Lesson objective

By the end of this lesson, you will be able to compute the gap between two rates, its 95% confidence interval and its p-value, to interpret them correctly (and spot wrong interpretations), and to judge whether the effect is large enough to matter, guardrails included.

Topics covered

  • p-value
  • confidence interval
  • statistical significance
  • practical significance
  • A/B test analysis

Where it fits

Analyze the result

How do I read a test result without over-interpreting it or missing a problem?

Lessons in this module

  1. Read a p-value and a confidence interval (this lesson)
  2. Analysis pitfalls: peeking, multiple comparisons, novelty, SRM, Simpson
  3. Bayesian or frequentist, and which tool to choose

What you will learn in the course

This lesson is part of the course Design and analyze an A/B test

  • Decide whether an A/B test is the right method (volume, B2B, reversibility, time, ethics) and choose a suitable alternative otherwise.
  • Write a testable hypothesis, choose a sensitive, attributable primary metric, guardrails and the exposure event to log.
  • Choose the randomization unit and prevent interference between groups, including in B2B.
  • Compute a sample size and a duration from the baseline rate, the minimum detectable effect, the significance level and the power.
  • Analyze a result (p-value, confidence interval, practical significance, guardrails) and read a tool's Bayesian or frequentist output.
  • Detect and avoid analysis pitfalls (peeking, multiple comparisons, novelty effect, sample ratio mismatch, Simpson's paradox).
  • Decide based on rules written before the test and document the experiment in a reusable memo.