What p < 0.05 Actually Says
A p-value is one of the most quoted numbers in science and one of the most consistently misread. The usual paraphrase — "the probability the result is just chance" — is not what it means.
This page states what a p-value actually measures, explains why a large p does not prove there is no effect, and why a significant result can still be too small to matter.
Then fifteen questions check whether it landed.
The sentence almost everyone gets backwards
Ask around and you will hear that p = 0.03 means there is a 3 % chance the result is down to luck, or a 97 % chance the effect is real. Both are wrong, and they are wrong in the same way: they have the conditional the wrong way round.
A p-value tells you about the data, given an assumption. It does not tell you about the assumption, given the data.
What it actually measures
A p-value is the probability of getting data at least as extreme as yours, assuming the null hypothesis is true.
That assumption is not a footnote; it is the whole setting. You temporarily grant that there is no effect, then ask how surprising your data would be in that world. A small p means the data sit awkwardly with that assumption. It does not measure how likely the assumption is, because the calculation took the assumption as given.
So 1 − p is not the probability the alternative is true. There is no route from a p-value alone to the probability that a hypothesis is correct — that would need to take account of how plausible the hypothesis was to begin with, and a p-value never sees that.
Where 0.05 came from
Nowhere in particular. It is a convention that stuck, not a property of the universe, and nothing physical changes between p = 0.049 and p = 0.051.
The threshold you pick is the rate of one particular mistake you are willing to accept: rejecting a null hypothesis that was actually true. That is a type I error, and its probability is the significance level α you chose. The opposite mistake — failing to reject a false null — is a type II error with probability β, and 1 − β is the power of the study.
A large p does not mean there is no effect
p > 0.05 means your data were not surprising enough, under the no-effect assumption, to reject it. That is absence of evidence, and it is not the same thing as evidence of absence.
A small, badly powered study will routinely fail to detect effects that are really there. Reporting that as "no effect" turns a limitation of the study into a claim about the world.
Significant and important are different words
"Statistically significant" says only that the effect is distinguishable from zero. It says nothing about whether the effect is big enough to care about.
Sample size is what drives them apart. With a million observations, an effect far too small to matter in practice will clear p < 0.05 comfortably. With twelve observations, an effect worth acting on may not.
This is why an effect size and a confidence interval are worth more than a p-value on its own: they say how large the effect might be, not merely that it is probably not exactly zero. A 95 % confidence interval that excludes zero corresponds to p < 0.05 in the same test — but it also tells you the range, which the p-value never does.
How p-values get abused
Multiple comparisons. Run twenty independent tests on data with no real effects and, at α = 0.05, one is expected to come out "significant" anyway. Report only that one and you have manufactured a finding. Corrections exist — the Bonferroni correction simply divides α by the number of tests — and their crudeness is part of the point.
p-hacking. Trying analyses, subgroups and exclusions until something crosses the threshold, then presenting it as though it had been the plan. The individual steps each look defensible, which is exactly what makes it hard to spot from the outside.
Neither of these is a flaw in the p-value itself. They are what happens when one number is asked to carry a decision it was never built to make.
What This Quiz Covers
- What the p-value is conditional on
- Why 1 − p is not the probability the effect is real
- Why 0.05 is a convention
- Absence of evidence against evidence of absence
- Statistical significance against practical importance
- Multiple comparisons and p-hacking
Cletica
Want to create your own quiz?
Build surveys and quizzes, share with anyone, collect responses — free to start.
Try Cletica for free →