Statistical significance
A result is statistically significant when the observed difference would be unlikely if the change had no real effect, given a threshold chosen in advance. In practice teams compute a p-value: the probability of observing a gap as big as the one measured, or bigger, assuming the change does nothing. When it falls below the threshold, often 5%, the result is called significant. A confidence interval complements it with the range of effect sizes compatible with the data.
Why it matters for a PM
Significance protects a PM from shipping changes that only looked good by chance, yet it is often misunderstood. A significant result can be too small to matter, and a non-significant one can hide a real effect that the test lacked the power to detect. Reading the interval, not just the verdict, leads to better decisions.
Example
A test of a new pricing page shows a 2% lift in sign-ups with a p-value of 0.03. Its confidence interval spans a lift of 0.2% up to 3.8%. The result is significant, but the plausible effect may be tiny, so the team weighs it against the cost of maintaining two page designs.
Key points
- A p-value does not tell you how likely the variant is to win.
- Statistical power is the chance of detecting a real effect of a given size; it depends on the sample size.
- Practical significance asks whether the gain is worth the cost of the change.
- Checking results repeatedly, or testing many metrics, inflates false positives unless you correct for it.
Common mistakes
- Reading “not significant” as “no effect”.
- Lowering the bar after the fact to obtain a significant result.
- Presenting a win without the width of its confidence interval.
Go further with Module
The courses and lessons that cover this concept: