Sample Size and Statistical Power

This is the thing that keeps me muttering at my laptop during the week, and occasionally shouting at the ocean on a Saturday morning swim. Bad statistics aren't always dishonest. Often they're just underpowered.
What "statistical power" actually means
Power is the probability that a study will detect a real effect if one actually exists. The conventional benchmark, set out in Cohen's foundational 1988 work on statistical power analysis, is 0.80 , meaning an 80% chance of catching a true signal. Fall below that and you're running a study with a meaningful probability of a false negative: you conclude "no effect found" when there genuinely was one.
Power is determined by three interacting factors: the size of the effect you're looking for, the variability in your measurements, and your sample size. Effect size and variability are usually constrained by biology or the research question itself , you can't negotiate with the endocannabinoid system about how consistent it feels like being. (Our glossary on the endocannabinoid system gives a decent primer on why receptor-level responses can vary substantially between individuals.) What a researcher can control is sample size. That's why it matters so much.
The false positive problem on the other end
Here's where it gets uncomfortable. Underpowered studies don't just miss real effects; they also produce inflated effect size estimates when they do find something significant. This is sometimes called the "winner's curse." A small trial that crosses the p<0.05 line by luck tends to report an effect much larger than the true one, because small samples are noisy, and only the extreme noise clears the significance bar.
So the paradox is this: you run a 30-person study, find a dramatic result, publish it, and the scientific community is impressed, until a 300-person replication finds a much smaller or nonexistent effect. This has played out repeatedly in psychology, nutrition, and increasingly in phytomedicine research. The 2015 Open Science Collaboration's large-scale replication project (published in Science) found that roughly 60% of published psychology findings failed to replicate at the same effect size. The pattern is not unique to that field.
Why this matters when reading cannabinoid research
A lot of early research on cannabidiol, cannabigerol, and the broader family of phytocannabinoids was done in very small samples or in cell and animal models that don't straightforwardly translate to humans. That's not a criticism of the researchers, early-stage work is necessarily exploratory, and you have to start somewhere. The problem is when those findings get reported as if they were definitive.
Take bioavailability research as a concrete example. The absorption of an oral cannabinoid compound varies substantially depending on whether it's taken with food, what formulation is used, and individual metabolic differences routed through first-pass metabolism. A study of 20 fasted adults will produce numbers that may look precise but are highly context-specific. Report those numbers without that caveat and you've produced something that looks like knowledge but functions more like folklore.
The same applies to peptide research. Peptide pharmacology involves complex questions about receptor affinity, stability, and; again, absorption and degradation. Small mechanistic studies are genuinely useful for building hypotheses. They should not be the sole basis for any claim about what a compound does in a living person.
What a proper power calculation looks like
Before a well-designed trial recruits a single participant, the researchers should have specified: the minimum effect size they consider clinically meaningful, the expected variability in the primary outcome measure, their desired power (usually 0.80 or 0.90), and their significance threshold (usually 0.05, though there's a live debate about whether 0.005 should become the new standard, as Benjamini and colleagues argued in a 2018 Nature Human Behaviour paper).
From those inputs, a sample size follows mathematically. This calculation should be reported in the methods section of any trial you read. If it isn't there, that's a question worth asking. Not every absence means misconduct, sometimes it's just sloppy reporting; but its presence is a basic mark of methodological transparency.
A few other things worth checking in methods sections:
- Was the primary outcome pre-registered before data collection (e.g., on ANZCTR or ClinicalTrials.gov)? Pre-registration substantially reduces the risk of outcome switching.
- Did the study report confidence intervals alongside p-values? A p-value of 0.04 combined with a confidence interval that nearly includes zero tells a different story than the headline number suggests.
- Were subgroup analyses pre-specified or post-hoc? Post-hoc subgroup findings in small studies are, to be blunt, barely more than anecdote.
The "n of 1" problem in this space specifically
Honestly, one of my bigger frustrations with how plant-based medicine gets discussed publicly is the weight given to individual experience. Someone takes a full-spectrum extract and reports a change. That's real to them. It is not evidence of a mechanism. Human beings are spectacularly bad at isolating variables in their own lives, sleep changed, stress changed, diet changed, the product changed, and we credit the product. This isn't unique to alternative health; it's how brains work.
I'm not saying personal experience is worthless. But it sits at the bottom of the evidence hierarchy for a reason. Controlled trials, adequately powered and pre-registered, exist precisely because we cannot trust our own perception of cause and effect. The history of medicine is littered with things that "worked" in individual cases and failed in trials. And the reverse: things that seemed to do nothing but turned out to have genuine mechanistic effects, once we could measure properly.
A note on interpreting research in regulated contexts
In Australia, the TGA evaluates evidence quality before substances can receive regulatory approval for specific indications. Products listed on the ARTG have met a defined evidentiary threshold, that threshold varies by listing versus registration pathway, and it's worth understanding the distinction. Substances accessed via the Special Access Scheme or via an Authorised Prescriber arrangement operate under different oversight frameworks to fully approved products.
When you read about a compound that's available through one of those pathways, the question of underlying evidence quality is still entirely open. Regulatory approval is necessary but not sufficient for certainty about effect sizes. And; this is the part people don't always hear, the absence of regulatory approval does not mean the evidence is absent. It means the approval process hasn't been completed. These are different claims.
I'd argue, as a general principle, that any discussion of cannabinoid or peptide research that doesn't foreground study size, power, and replication status is doing readers a disservice. Not maliciously, usually. But the effect on public understanding is the same either way.
How to build a quick mental checklist
When I read a study, and I read a lot of them, usually with a very strong coffee and the cattle dog demanding attention from the other end of the couch; I run through roughly this sequence:
Sample size first. Is there a power calculation? Does the reported N match what was calculated? (Dropout is common; a study powered for 80 that completes with 51 is a different study.) Then: was it pre-registered? Then: what were the confidence intervals? Then: has it been replicated, and by whom?
If a study clears all of those checkpoints, I start paying real attention. Most don't clear them all. That doesn't mean the research is worthless, early-phase work is genuinely important for the field. It means the conclusion should be "interesting preliminary finding, needs replication" rather than "this compound does X."
The short version is: p-values are not the story. Sample size and power are the infrastructure the whole thing rests on. Get comfortable reading those numbers and you'll read research very differently, and probably more sceptically than most headlines would prefer.
Sources
- Open Science Collaboration, "Estimating the reproducibility of psychological science"; NCBI / Science
- Benjamin et al., "Redefine statistical significance", NCBI / Nature Human Behaviour
- Evidence-based regulation framework, TGA (Australian Government)
- Australian New Zealand Clinical Trials Registry; ANZCTR
, Hannah Bui, Evidence & Research Literacy Writer
]]>Common questions
- What is statistical power in simple terms?
- Statistical power is the probability that a study will detect a real effect if one genuinely exists. A power of 0.80 means an 80% chance of catching a true signal. Studies with low power risk missing real effects — called false negatives — and can also inflate the apparent size of effects when they do find something.
- How small is too small for a study sample?
- There's no single universal cutoff — the required sample size depends on the size of the effect being studied and the variability in the outcome measured. What matters is whether the researchers calculated a required sample size before the study began (a power calculation) and whether recruitment matched that target. A study of 20 people might be adequate for some questions and wildly insufficient for others.
- What is a p-value and why isn't it enough on its own?
- A p-value is the probability of observing a result at least as extreme as the one found, assuming the null hypothesis (no effect) is true. A p-value below 0.05 is conventionally called 'significant', but it doesn't tell you how large the effect is, how precise the estimate is, or whether the finding will replicate. Confidence intervals and effect sizes carry more practical meaning.
- What is pre-registration and why does it matter?
- Pre-registration means the researcher publicly records their hypotheses, primary outcomes, and analysis plan before data collection starts — usually on a registry like ANZCTR or ClinicalTrials.gov. This prevents 'outcome switching', where a researcher analyses multiple outcomes and only reports the one that turned out significant. Pre-registered findings carry substantially more weight.
- Does this affect how I should read cannabinoid or phytomedicine research?
- Yes. Many early studies in this space used small samples, animal or cell models, and were not pre-registered. That makes them useful as hypothesis-generating work but not as firm evidence of a mechanism or effect in humans. Always check whether a finding has been replicated in an adequately powered, pre-registered human trial before treating it as established.
Related reading
Preclinical vs Clinical ResearchPreclinical research and clinical trials are not the same thing. Here's how to tell them apart — and why that gap matters more than most headlines let on.
Randomised Controlled Trials ExplainedRCTs are the gold standard for testing whether something works — but they have real limits. Here's how to read them without being fooled.
Real-World Evidence vs Trial DataRCT data and real-world evidence each tell a different story. Here's how to read both without getting fooled — and why the gap between them matters.
Pre-Registration and Publication BiasPublication bias shapes what science "knows". Here's how pre-registration works, why it matters, and how to spot studies that may have moved the goalposts.
Observational Studies vs RCTsObservational studies and RCTs answer different questions. Knowing which is which changes how much weight you should give any headline claim.
Understanding Statistical SignificanceP-values, sample sizes, confidence intervals — knowing what these actually mean can save you from a lot of misleading headlines. Here's how to read them honestly.
I am the resident sceptic. I write about how to read studies without getting fooled, and the history of how we got here. Sea swimmer year-round, statistics nerd, op-shop devotee, and owner of one very opinionated cattle dog.
BSc Statistics
More from Hannah Bui
The Origins of Cannabis ProhibitionCannabis wasn't always a prohibited substance. Understanding how prohibition came to pass reveals as much about politics and race as it does about pharmacology.
The Hierarchy of EvidenceNot all studies are created equal. Here's how to read the hierarchy of evidence without getting fooled by headlines, hype, or a single promising trial.
Pre-Registration and Publication BiasPublication bias shapes what science "knows". Here's how pre-registration works, why it matters, and how to spot studies that may have moved the goalposts.
Conflicts of Interest in ResearchWho funded the study? That single question can change how you read almost any cannabinoid or phytomedicine research. Here's how to spot funding bias.
Confounding Variables ExplainedConfounding variables are the reason a convincing study can still be wrong. Here's how to spot them — and why it matters in plant-based medicine research.
Why Anecdotes Are Not EvidenceAnecdotes feel compelling, but they're not evidence. Here's how to spot the difference — and why it matters when you're reading cannabinoid or phytomedicine research.