Making Sense of the Numbers: How to Decode Statistics Without a Degree

September 25, 2025
Paige Reagan

You’re skimming through a new study someone posted in a practitioner group. The headline makes a big claim: TheraFlame significantly reduces inflammation.” Curious, you click through, but by the second paragraph your eyes glaze over: p-values, confidence intervals, hazard ratios … and what’s a Kaplan-Meier curve? You quickly scroll to the conclusion, hoping for something … anything … you can actually understand. Sound familiar?

As functional practitioners, we strive to make decisions informed by research. But for many of us, the language of statistics can feel intimidating, overly technical, and downright unfamiliar. We scan the abstract, squint at the figures, and secretly hope someone on Instagram has already broken it down, instead of digging into the data ourselves. 

If this is you, you’re not alone, and you’re not failing as a clinician. Interpreting the literature can feel like decoding a secret language. Yet, understanding just a few key concepts can go a long way in helping you feel more confident and less misled by what the research actually says, so you can make clinical decisions that truly reflect the evidence.

Here’s a no-stress breakdown of some of the most common statistical concepts and why they matter for your clinical practice. You don’t need a PhD in biostatistics, just a curious mind and desire to make better use of the research that guides your care.

Understanding the Assumptions Behind the Numbers

Before diving into research statistics, it’s helpful to understand what researchers are actually testing and how the process is set up from the beginning.

What many people don’t realize is that scientific testing doesn’t start by trying to prove something is true. It starts with assuming the opposite.

Research defines two competing ideas, called hypotheses, which are educated guesses about what might be true. The null hypothesis is the default assumption: that there’s no effect, no difference, or no relationship. The alternative hypothesis is what they hope to support: that there is a meaningful effect or difference.

Here’s the twist. Even though the alternative is the outcome the researchers are interested in, they don’t test it directly. Instead, they assume the null is true and then use data to see if that assumption holds up.

It’s a bit like a court case. The null hypothesis is the defendant, assumed innocent unless there’s strong evidence to say otherwise. The researcher, like a prosecutor, gathers data to challenge that assumption. If the results would be very unlikely under the assumption of “no effect,” we might reject the null and consider the alternative a better explanation.

Let’s apply this to our working example:

A study claims that TheraFlame significantly reduces inflammation.

  • The null hypothesis says TheraFlame has no effect
  • The alternative hypothesis says the supplement does reduce inflammation

The researchers collect data from a sample of participants, run the analysis and ask: If TheraFlame really did nothing, how likely would we be to see a result like this? This is where statistical tools, like the p-value, come into play.

P-value: How likely is this result due to chance?

The p-value is one of the most common, and commonly misunderstood, statistics in research.

At its core, a p-value helps you to evaluate how consistent the data are with the null hypothesis. It tells you the probability of seeing a result as extreme as what we observed (or even more extreme), if the null hypothesis were true.

This fits into the logic we just walked through: we assume the null is true, gather data, and then ask: How surprising is this result given that assumption?

In most studies, a p-value less than 0.05 is considered statistically significant. That means there is less than a 5% chance of seeing a result this extreme just by random chance if the null were true. In other words, researchers are accepting no more than a 1-in-20 risk of being wrong when they reject the null hypothesis.

So, if p<0.05, we usually reject the null and consider the alternative (that there is an effect) more plausible. When p>0.05, we don’t reject the null, though that doesn’t prove it’s true. It just means the evidence wasn’t strong enough to rule it out.

Let’s go back to TheraFlame. Suppose the study finds that TheraFlame reduces C-reactive protein, a marker of inflammation, by 2 mg/L, with p=0.03. This means that there’s a 3% probability of getting a result this extreme (or more extreme) if the supplement truly had no effect. That’s relatively unlikely under the “no effect” assumption, so researchers may reject the null and conclude that TheraFlame likely does reduce inflammation.

Think of it this way: if the null hypothesis were true, a result like this would only show up 3 times out of 100. That doesn’t prove the supplement works, but it does suggest the observed result is unlikely to be a fluke, and consistent with a real effect.

Confidence Interval: How sure are we?

While the p-value tells us how surprising the result is under the assumption of no effect, the confidence interval (CI) tells us how precise the result is and how much uncertainty surrounds it.

A confidence interval gives us a range of values that likely contains the true effect in the population. Most studies report a 95% CI, which means if the study were repeated 100 times with different samples, the true effect would fall within that range about 95 times.

The width of the confidence interval provides important context. A narrow CI suggests more precision and more confidence in the estimate. This is often due to a larger sample size or more consistent data. A wide CI suggests less precision and more uncertainty, often due to smaller sample sizes or more variable data.

So looking at TheraFlame again, suppose our study reports that TheraFlame lowers CRP by 2 mg/L, with a 95% CI of 1.8 to 2.2. That’s a tight range, suggesting the researchers are fairly confident the true effect is close to 2 mg/L.

Now imagine that same result but with a 95% CI of -1 to +5. That’s a wide interval, and more importantly, it includes zero, meaning the supplement might not lower inflammation at all. In this case, the result isn’t statistically significant even if the p-value is borderline.

So while the p-value gives you a yes-or-no answer about statistical significance, the confidence interval shows you how much might be going on and how sure we are about it. A small p-value might tell you something is happening, but the confidence interval helps you understand how much is going on and how reliably.

Sample Size and Power: Why Bigger (Usually) Means Better

We’ve talked about how p-values tell us whether a result is statistically significant and how confidence intervals show how precise that result is. But there’s another key factor that influences both: sample size.

Simply put, the number of people in a study matters—a lot. Larger sample sizes generally lead to more precise estimates (narrower CIs), a greater ability to detect a true effect (higher statistical power), and more reliable, stable results.

Small studies can still be useful, but they come with more uncertainty. With fewer participants, it’s harder to tell if a result is real or just due to chance. A small sample may miss a real effect (called a Type II error), produce wide CIs, or give unreliable p-values that don’t replicate in future studies.

Let’s go back to TheraFlame. Suppose we’ve designed a small pilot study that includes just 12 participants. The researchers find that CRP levels dropped by 2 mg/L, with a p-value of 0.07 and a 95% CI of -1 to +5. That result isn’t statistically significant, and the wide CI includes both a possible effect and no effect, so it’s hard to say whether TheraFlame had any real impact.

Now imagine a second study with 3000 participants. The CRP drop is still 2 mg/L but now the p-value is 0.01 and the 95% CI is 1.8 to 2.2. This is a more convincing result as it’s statistically significant, precise, and less likely to be due to chance.

We can’t talk about sample size without introducing power, which is the probability that a study will detect a true effect, if one exists. Most studies aim for 80% power, which means there’s a 4 in 5 chance the study will catch a real effect, and only a 1 in 5 chance it will miss it (a Type II error).

Low-powered studies are more likely to miss real effects or produce unstable results. That doesn’t mean small studies are useless, but it does mean their findings should be interpreted with extra caution. A nonsignificant result in a small study doesn’t necessarily mean a treatment is ineffective—it might just mean the study wasn’t big enough to detect an effect.

And Finally, Statistical Significance ≠ Clinical Relevance

A study can show statistical significance, yet still fall short of having a meaningful impact in practice. Just because p<0.05 doesn’t mean the effect is important for your client. It simply means the result is unlikely to be due to chance.

For example, if TheraFlame lowers CRP by 1 mg/L with a p=0.01, that’s statistically significant. But is a 1-point drop in CRP helpful for your client with autoimmune issues, chronic inflammation, and high stress? Maybe, but that depends on clinical context, not just the number.

On the flip side, some results might not reach statistical significance, but that doesn’t always mean the intervention is ineffective. A small study, wide confidence interval, or early-stage trial might still offer useful insights, especially if it supports what you’re seeing in practice or aligns with our understanding of the mechanism. The key is to interpret these results thoughtfully, not dismiss them outright.

Bringing It All Together

You don’t need a PhD in statistics to make sense of the research. But learning a few key concepts, like what p-values and confidence intervals actually mean, and how sample size affects both, gives you a stronger foundation to ask better questions, spot overstated claims, and make decisions that are both data-informed and client-centered.

As a practitioner, you’re not expected to be fluent in statistics, but understanding the basics helps you recognize what’s meaningful, what’s uncertain, and how (or whether) it applies to your client.

FAQ

The following FAQs distill the key statistical concepts from this article to help practitioners interpret research with greater clarity and clinical discernment.

Why do functional practitioners need to understand basic research statistics?

Understanding core statistical concepts helps practitioners interpret evidence accurately and avoid being misled by overstated claims. Even a working knowledge of p-values, confidence intervals, and sample size improves data-informed, client-centered decision-making.

What is the null hypothesis in a clinical study?

The null hypothesis assumes there is no effect, difference, or relationship. Researchers begin by assuming the intervention does nothing and then test whether the data provide enough evidence to reject that assumption.

What does a p-value actually tell you?

A p-value estimates how likely the observed result would be if the null hypothesis were true. If p<0.05, the result is considered statistically significant, meaning it would be unlikely to occur by chance under a “no effect” assumption.

Does a statistically significant p-value prove that a treatment works?

No, statistical significance does not prove effectiveness. It only suggests the result is unlikely due to random chance; it does not confirm causation or clinical importance.

What is a confidence interval and why does it matter?

A confidence interval (CI) provides a range of values likely to contain the true effect in the population. A narrow CI suggests precision, while a wide CI signals uncertainty and less reliable estimates.

Why does it matter if a confidence interval includes zero?

If a confidence interval includes zero, the result may not be statistically significant. This means the true effect could be no difference at all, even if the point estimate appears meaningful.

How does sample size influence study results?

Larger sample sizes generally produce more precise estimates and more reliable p-values. Small studies often generate wider confidence intervals, unstable findings, and a higher risk of missing real effects.

What is statistical power in clinical research?

Statistical power is the probability that a study will detect a true effect if one exists. Most studies aim for 80% power, meaning there is a 4 in 5 chance of identifying a real effect.

What is a Type II error and why should practitioners care?

A Type II error occurs when a study fails to detect a true effect. This risk is higher in small or low-powered studies, meaning a nonsignificant result does not automatically indicate ineffectiveness.

Why doesn’t statistical significance always equal clinical relevance?

Statistical significance only indicates that a result is unlikely due to chance. Clinical relevance depends on whether the magnitude of effect meaningfully impacts patient outcomes in real-world practice.

How should practitioners interpret nonsignificant results?

Nonsignificant results should be interpreted cautiously rather than dismissed outright. Small sample size, wide confidence intervals, or early-stage research may limit statistical detection even when an effect exists.

What is the practical takeaway when reviewing new research?

Look beyond the abstract and examine the p-value, confidence interval, and sample size together. Evaluating these elements collectively helps determine whether findings are statistically sound, precise, and clinically meaningful.

ABOUT THE AUTHOR:

Paige Reagan
MRHP, FNTP

Paige has spent most of her career working in Research and Development in the areas of clinical research, regulatory affairs, and medical writing. She has a wide range of experience in the therapeutic areas of cardiovascular health, pulmonary arterial hypertension, diabetes, bone health, osteoarthritis/rheumatoid arthritis, and urology, among others. Her work has contributed to numerous regulatory approvals as well as publications in major medical journals such as the New England Journal of Medicine, Lancet, Circulation, and American Heart Journal.

Read more about Paige

Join Our Restorative Health Community

We combine curriculum, mentorship, and real‑world application to empower practitioners worldwide. Explore resources, connect with peers, and take the next step in your journey.

Study With Us
Find a Practitioner
Contact Admissions

Free Resources

Categories

How To Run A CGM Challenge That Gets Real Results September 15th 1:00 PM PST/4:00 PM EST.

X