Real-World Evidence vs Trial Data

By Hannah Bui · 22 May 2026 · 7 min read
a man looking through a microscope at something
Roughly 85% of clinical trial participants are white, male, and younger than 65 , despite the fact that most chronic conditions skew older, more diverse, and often female. That single statistic from a 2021 analysis published in JAMA Network Open does more to explain the limits of trial data than any methodology lecture I could write.

I think about it a lot. Especially when someone forwards me a study summary and says "but it was a randomised controlled trial , doesn't that mean it's settled?"

Short answer: not quite. The long answer is what this article is about.

What "real-world evidence" actually means

Real-world evidence (RWE) is data collected outside the controlled conditions of a clinical trial. It comes from patient registries, electronic health records, insurance claims databases, post-market surveillance programmes, and observational studies. The patients taking the substance are doing so in the genuine mess of daily life , alongside other medications, variable diets, different sleep patterns, stress loads, and the rest of it.

Randomised controlled trial data (RCT data) is the inverse. Researchers recruit a tightly defined population, randomly assign participants to an intervention or a placebo, control as many variables as possible, and measure a pre-specified outcome over a defined period. It is, by design, an artificial environment. That's its strength. It's also its limitation.

Neither is "better." They answer different questions. Confusing the two is where most science communication goes wrong.

Why RCTs are the gold standard; and what that phrase actually means

The randomised controlled trial became the dominant method for evaluating pharmaceutical interventions because randomisation is the only reliable way to control for confounding, the problem of unknown variables that differ between people who choose to take a substance and people who don't.

Say a group of people decide to start taking a particular herbal preparation. They are, as a group, probably more health-conscious, more likely to exercise, more likely to have social support. Any positive outcome you measure might have nothing to do with the preparation. It might reflect the kind of person who seeks out that preparation. This is called healthy-user bias, and it haunts observational research.

Random assignment breaks that link. If assignment to treatment is truly random, the health-conscious people and the less health-conscious people should, on average, end up distributed across both arms. That's why RCT results are more likely to reflect a genuine causal relationship between the intervention and the outcome.

But "gold standard" doesn't mean "complete picture." It means "best method for establishing causation under controlled conditions." Those conditions are a long way from your kitchen bench at 8pm.

Where RCTs fall short in the real world

The homogeneity problem I mentioned up top is real. Trial populations tend to underrepresent women, older adults, people on multiple medications, people from non-English-speaking backgrounds, and people with complex comorbidities. The TGA's post-market surveillance framework and the FDA's own guidance documents both acknowledge that trial results may not generalise to the full range of patients who ultimately use an approved product.

There's also the duration problem. A 12-week trial tells you something about 12 weeks of exposure. It says almost nothing about five years of use. Long-term safety signals, rare adverse events, cumulative effects on organ systems, interaction effects that only emerge over time; typically show up in post-market data, not trial data. That's not a failure of the trial design. It's a structural feature of it.

And then there are the dosing and formulation questions that trials often can't capture cleanly. Bioavailability, the proportion of a substance that actually reaches systemic circulation, varies with the route of administration, food intake, individual metabolic variation, and formulation. First-pass metabolism alone can reduce oral bioavailability dramatically and variably between individuals. Trials standardise this as best they can. Real-world use doesn't.

Where real-world evidence falls short

Confounding, confounding, confounding. I'll admit I was a bit too enthusiastic about large observational datasets when I first started reading this literature; the sample sizes looked impressive and I let that do too much work. I got that wrong for a while.

RWE can be riddled with selection bias, reporting bias, and immortal-time bias (a particularly sneaky one where how you define the start of "exposure" can make a drug look protective when it isn't). Electronic health records often capture diagnoses inconsistently. Patient registries have their own enrolment biases. Post-market surveillance in Australia under the TGA's Special Access Scheme captures some real-world data on unapproved products, but completeness depends heavily on prescriber reporting.

RWE is observational. At its best, it generates hypotheses and flags signals. It should not, on its own, be used to establish causation, even when the numbers look compelling.

The specific challenge in cannabinoid and phytomedicine research

This tension matters a lot in the area of phytomedicine and cannabinoid science. Partly because the RCT base is still comparatively thin, decades of legal restriction on cannabis research left a significant gap in the controlled-trial literature. Partly because the compounds themselves are complex.

Take cannabidiol. The endocannabinoid system it interacts with is distributed across many tissue types and involves multiple receptor subtypes; including the CB1 and CB2 receptors, with effects that differ substantially depending on baseline endocannabinoid tone, which is itself individually variable. Conducting a controlled trial that captures that variability cleanly is genuinely hard. Real-world registries may actually have more heterogeneous and therefore (in some ways) more representative data on individual variation than small, tightly controlled trials do.

But that doesn't make the RWE causal. The gap between "patients in this registry reported X" and "this compound caused X" is large, and researchers are often more careful about that distinction than the popular summaries of their work are.

Pharmacokinetics is another area where the lab-to-life gap is real. A trial measuring plasma concentration curves under controlled fasting conditions is measuring something quite different from what happens when someone takes the same product after a high-fat meal, or alongside other substances. The full-spectrum vs isolate question, and the contested science around the entourage effect; is a good example of an area where RCT data is sparse and the RWE (largely from patient surveys and registry data) generates interesting hypotheses but not conclusions.

How to read a study without getting fooled

A few practical habits I've developed over years of reading this literature.

First: check who was in the trial. Sample size matters, but sample composition matters more. A 400-person trial that enrolled only young males tells you less than you think about a general population.

Second: look at what was actually measured. Surrogate endpoints (a biomarker, a lab value) are not the same as clinical endpoints (quality of life, functional outcomes). Many trials use surrogates because they're faster and cheaper to measure. That's legitimate science, but it's limited science.

Third: for RWE, ask what was adjusted for. A well-conducted observational study will use statistical methods, propensity scoring, multivariate regression, to try to account for confounding. If the paper doesn't discuss confounding at all, be cautious. If it claims causation from observational data without strong instrumental variable analysis or a natural experiment design, be very cautious.

Fourth: check the funding and the registration. Pre-registered trials (registered on ClinicalTrials.gov or ANZCTR before results are known) are more trustworthy than unregistered ones. Industry-funded studies are not automatically wrong, but the effect-size inflation in industry-funded trials is documented and significant.

I had a long conversation about this with a researcher at the Menzies Institute for Medical Research on Argyle Street in Hobart a few months back; she made the point that the most important question you can ask about any study is simply: what was this designed to show, and what would it have missed by design? That question alone will carry you a long way.

My honest assessment: the gap between how researchers use these tools and how the results are communicated publicly is a bigger problem than the tools themselves. "RCT says X" in a headline almost never captures what the RCT actually said, or the assumptions it rested on. Learning to go back to the abstract, and ideally the methods section, is the single most useful research-literacy skill there is. It's not glamorous. But neither is getting fooled by a press release.

Sources

, Hannah Bui, Evidence & Research Literacy Writer

]]>

Common questions

Is real-world evidence less reliable than RCT data?
Not less reliable — differently reliable. RCTs are better at establishing whether a compound caused a particular effect (causation), because randomisation controls for confounding. Real-world evidence is better at capturing how a product performs across diverse populations over longer timeframes. The two are complementary, not competitive. The mistake is treating one as a substitute for the other.
Why do clinical trials often use such narrow populations?
Mainly for practical and statistical reasons. Narrow inclusion criteria reduce variability, which makes it easier to detect a signal with a smaller sample. Recruiting a more diverse population takes longer and costs more. Historically there were also regulatory and liability reasons that discouraged including pregnant people, older adults, and those with complex health histories. Regulators including the TGA and the FDA have increasingly pushed for broader enrolment, but change has been slow.
What is confounding and why does it matter in observational studies?
Confounding happens when a third variable is linked to both the exposure and the outcome, making it look like there's a direct relationship when there might not be. For example, people who seek out particular supplements may also exercise more, eat better, or have higher health literacy — all factors that independently affect health outcomes. Without randomisation, it's genuinely difficult to separate the effect of the supplement from the effect of the kind of person who takes it.
How does Australian regulation use real-world evidence?
The TGA uses real-world data primarily in post-market surveillance — monitoring approved and provisionally approved products after they reach the market. Data from the Special Access Scheme and the Authorised Prescriber pathway can contribute to the evidence base for less-established substances. The TGA has also published guidance on how RWE can complement trial data in regulatory submissions, though it does not replace the requirement for controlled trial evidence in most approval pathways.
What is pre-registration and why should I look for it in a study?
Pre-registration means the researchers publicly recorded their hypotheses, methods, and primary outcome measures before collecting or analysing data. It's done through registries like ANZCTR (Australia and New Zealand) or ClinicalTrials.gov. Pre-registration makes it much harder to engage in 'outcome switching' — changing the primary outcome after seeing the data to find a statistically significant result. Studies that are not pre-registered are not automatically invalid, but the absence of pre-registration is worth noting when evaluating findings.

Related reading

About the author
HB
Hannah Bui
Evidence & research-literacy writer · Hobart, TAS

I am the resident sceptic. I write about how to read studies without getting fooled, and the history of how we got here. Sea swimmer year-round, statistics nerd, op-shop devotee, and owner of one very opinionated cattle dog.

BSc Statistics

More from Hannah Bui