Skip to content
Module
Product and data

Statistical significance

A result is statistically significant when the observed difference would be unlikely if the change had no real effect, given a threshold chosen in advance. In practice teams compute a p-value: the probability of observing a gap as big as the one measured, or bigger, assuming the change does nothing. When it falls below the threshold, often 5%, the result is called significant. A confidence interval complements it with the range of effect sizes compatible with the data.

Why it matters for a PM

Significance protects a PM from shipping changes that only looked good by chance, yet it is often misunderstood. A significant result can be too small to matter, and a non-significant one can hide a real effect that the test lacked the power to detect. Reading the interval, not just the verdict, leads to better decisions.

Example

A test of a new pricing page shows a 2% lift in sign-ups with a p-value of 0.03. Its confidence interval spans a lift of 0.2% up to 3.8%. The result is significant, but the plausible effect may be tiny, so the team weighs it against the cost of maintaining two page designs.

Key points

  • A p-value does not tell you how likely the variant is to win.
  • Statistical power is the chance of detecting a real effect of a given size; it depends on the sample size.
  • Practical significance asks whether the gain is worth the cost of the change.
  • Checking results repeatedly, or testing many metrics, inflates false positives unless you correct for it.

Common mistakes

  • Reading “not significant” as “no effect”.
  • Lowering the bar after the fact to obtain a significant result.
  • Presenting a win without the width of its confidence interval.

Go further with Module

The courses and lessons that cover this concept: