Skip to content
TAtaimoorasghar.com

August 26, 2026 · 15 min read

A One-Point Difference Can Be Statistically Significant—but Does It Actually Matter?

A statistically significant one-point difference may be real without being large. Here is how to judge effect size, uncertainty, and practical meaning.

A one-point difference can be statistically significant without being large, clinically important, or meaningful to an individual person. Statistical significance answers whether the observed difference is difficult to explain by sampling variation under a specified null hypothesis. It does not, by itself, tell us whether the difference is substantial enough to influence symptoms, functioning, clinical decisions, or policy. That distinction became especially relevant in our recent study of depression among medical and non-medical university students in Pakistan, where median PHQ-9 scores differed by just one point even though the group comparison was statistically significant.

By Taimoor Asghar

Why statistical significance and meaningful difference are not the same thing

Researchers commonly summarize a comparison with a p-value. When that value falls below a predefined threshold such as 0.05, the result is often described as statistically significant. The temptation is then to translate “statistically significant” into “important.” That translation is not justified.

A p-value is primarily about compatibility between the observed data and a statistical model in which the null hypothesis is assumed. It is influenced not only by the size of the difference but also by sample size, variability, study design, and the statistical test being used. With a sufficiently large or precise dataset, even a small difference can produce a low p-value.

The reverse is also possible. A potentially important difference may fail to reach conventional statistical significance when the sample is small or estimates are imprecise. This is why interpreting research requires looking beyond whether p is above or below 0.05.

The distinction is particularly important when outcomes are questionnaire scores. A difference of one point on a multi-item symptom scale has a very different interpretation from a one-point difference on a scale with a narrow range or from a one-point change in an outcome directly linked to survival or disability.

A real example: PHQ-9 scores in medical and non-medical students

Our cross-sectional study included 602 undergraduate students from universities in Lahore, Pakistan: 424 medical students and 178 non-medical students. We evaluated depressive symptoms using the nine-item Patient Health Questionnaire, or PHQ-9, alongside the DASS-21 and several psychometric analyses. :contentReference[oaicite:0]{index=0}

The median PHQ-9 score among medical students was 9, compared with 10 among non-medical students. The unadjusted group difference was statistically significant. In regression analysis accounting for age, gender, and sleep duration, academic discipline remained associated with PHQ-9 scores: medical students scored approximately 1.43 points lower than non-medical students. :contentReference[oaicite:1]{index=1}

Readers interested in the complete methods and results can read the BMC Psychology study on PHQ-9 and DASS-21 symptom profiles among medical and non-medical students.

At first glance, those findings appear to establish a clear group difference. Statistically, there was evidence of one. But the next question is much more interesting: how much importance should we attach to a difference of roughly one to one-and-a-half PHQ-9 points?

The answer is more cautious than simply saying that one group was “more depressed.” The difference was detectable at the group level, yet academic discipline explained only a small proportion of overall variation. The primary PHQ-9 regression model had an adjusted R-squared of approximately 0.034, meaning that the variables in that model explained only a small fraction of the differences in PHQ-9 scores between individual students. :contentReference[oaicite:2]{index=2}

That combination—statistical significance alongside modest explanatory power—is exactly why magnitude and context must accompany p-values.

What does one PHQ-9 point actually represent?

The PHQ-9 contains nine questions covering depressive symptoms over the previous two weeks. Each item is scored from 0 to 3, producing a total score between 0 and 27. The original validation work showed that progressively higher scores were associated with greater depressive symptom severity and functional difficulty. :contentReference[oaicite:3]{index=3}

Because the total scale spans 28 possible integer scores, one point represents a relatively small movement on the scale. It could arise, for example, because a respondent reports one symptom occurring on “several days” rather than “not at all,” while every other response remains identical.

That does not make the point meaningless. Across hundreds or thousands of people, a consistent shift of one point may reveal a genuine population pattern. But it does mean that the interpretation of a group-average difference should not automatically be transferred to an individual student.

Group differences and individual changes answer different questions

Suppose University A has a mean PHQ-9 score one point higher than University B. That comparison concerns the average location of two distributions. It does not imply that every student at University A is one point worse, that most students could be distinguished based on their score, or that moving an individual person’s PHQ-9 by one point would necessarily represent a noticeable improvement or deterioration.

The distributions may overlap extensively. Many students in the nominally lower-scoring group may have higher scores than many students in the higher-scoring group.

This distinction becomes even more important when researchers discuss a minimal important difference: the magnitude of change considered meaningful in a particular context. Research evaluating PHQ-9 changes has produced estimates larger than a single point, and the exact threshold depends on the population, method, and purpose of measurement. One large methodological analysis, for example, examined patient-reported improvement and found that commonly used estimates around four points were within the range associated with meaningful perceived improvement, while also demonstrating that no universal cut-off perfectly separates people who feel better from those who do not. :contentReference[oaicite:4]{index=4}

Consequently, it would be inappropriate to take a one-point cross-sectional difference between two populations and describe it as though it were equivalent to a clinically meaningful within-person treatment response.

Why can a small difference become statistically significant?

Statistical testing considers the magnitude of an observed difference relative to its uncertainty. If measurements are reasonably precise and the sample contains enough participants, the uncertainty around a group estimate becomes smaller. A modest difference can therefore become distinguishable from zero.

Imagine two hypothetical studies comparing the same outcome. In the first, 20 people are enrolled in each group. In the second, 2,000 people are enrolled in each group. Both studies observe an average difference of one point. The larger study will generally estimate that difference much more precisely and may produce a substantially smaller p-value.

The effect itself has not become larger. Our confidence that the population difference is not exactly zero has changed.

This leads to a useful rule when reading research: a small p-value is not a measurement of effect size.

A result with p = 0.001 is not automatically more important than one with p = 0.04. The former may simply be estimated more precisely. Magnitude needs to be evaluated separately.

Five questions to ask after seeing a statistically significant result

1. How large is the difference in the original units?

Start with the actual number. If two groups differ by 1.0 PHQ-9 point, say so. Raw units often communicate information that disappears when a result is reduced to “significant” or “non-significant.”

Readers can then ask what one point means relative to the possible scale range, typical variation, severity categories, measurement error, and expected meaningful change.

2. What is the effect size?

Standardized effect sizes help describe how large a group difference is relative to variability within the sample. Depending on the analysis, researchers may report measures such as Cohen’s d, rank-based effect sizes, correlation coefficients, odds ratios, risk differences, or standardized regression coefficients.

No effect-size statistic should be interpreted mechanically, but it provides a dimension of information that a p-value cannot.

In our student study, the raw PHQ-9 medians were close despite a statistically detectable difference. That should immediately encourage a more measured interpretation than simply declaring one academic group psychologically healthier than the other. :contentReference[oaicite:5]{index=5}

3. What does the confidence interval allow?

A point estimate gives the best estimate generated by the model; a confidence interval communicates its uncertainty. Two studies could produce identical estimated differences but very different levels of precision.

Confidence intervals are particularly valuable because they help readers consider whether the data remain compatible with effects that would be trivial, moderate, or potentially important. A narrow interval around a small effect tells a different story from a wide interval spanning both negligible and substantial effects.

4. Is there an established meaningful-change threshold?

For some clinical outcomes, researchers have investigated minimal important differences, minimal clinically important differences, reliable change thresholds, or related benchmarks.

These benchmarks can help, but they need careful use. A threshold estimated for monitoring treatment response in adults with diagnosed depression cannot automatically be applied to a cross-sectional comparison between two university populations. The study design, population, baseline severity, purpose of testing, and anchor used to define improvement may all differ.

Meaningfulness is therefore contextual rather than a universal mathematical property of a score.

5. Does the predictor explain much of the variation?

A statistically significant predictor can still explain very little about why individuals differ from one another.

In our study, academic discipline was associated with depression and anxiety scores, but regression models explained only a small proportion of the overall variance. :contentReference[oaicite:6]{index=6}

This matters conceptually. Students are not defined by their degree programme. Mental health reflects many interacting influences, including personal history, social environment, economic pressures, sleep, relationships, academic circumstances, physical health, and other factors that may not have been captured in a particular dataset.

Statistically significant does not mean clinically significant

The phrase “clinically significant” is also frequently misunderstood. Statistical significance is generated by a statistical test. Clinical significance concerns whether the magnitude of an effect is relevant to symptoms, functioning, treatment, quality of life, or clinical decision-making.

The two concepts can occur in different combinations.

  • A result can be statistically significant but clinically small.
  • A result can be statistically significant and clinically important.
  • A potentially clinically important effect can be statistically inconclusive because the estimate is imprecise.
  • A statistically non-significant result does not prove that two groups are identical.

For mental-health questionnaires, this distinction is particularly important because scores are imperfect summaries of complex experiences. The PHQ-9 is useful for assessing depressive symptom severity, but a score is not equivalent to a complete clinical assessment. The original PHQ-9 validation study itself evaluated the questionnaire in relation to diagnostic assessment and functional outcomes rather than treating the numerical total as a diagnosis by itself. :contentReference[oaicite:7]{index=7}

Cross-sectional comparisons require another layer of caution

There is another reason not to overinterpret the medical versus non-medical difference in our study: it was cross-sectional.

We observed students at one period rather than randomly assigning people to academic disciplines and following what happened afterward. Consequently, the results identify associations but cannot establish that studying medicine reduces depression or that entering a non-medical programme increases it.

The two groups may differ in unmeasured characteristics. Selection into a degree programme is not random. Socioeconomic conditions, employment expectations, family circumstances, institutional environment, personality, academic pressure, access to support, and many other factors could influence both educational choices and psychological outcomes.

This is a general lesson for observational research: a statistically significant adjusted regression coefficient remains an association unless the research design and assumptions justify a causal interpretation.

A one-point difference can still matter at the population level

None of this means small effects should be dismissed.

Population health operates differently from individual clinical care. A small shift affecting a very large population can sometimes have meaningful consequences even when the average change for each person is modest. Public-health researchers therefore care about effect magnitude, prevalence, distribution, feasibility, cost, and the number of people exposed.

A small average difference can also direct attention toward an overlooked group. In our study, the direction of the result challenged the assumption that medical students should automatically be treated as the university population with the greatest mental-health burden. Non-medical students reported somewhat higher depression and anxiety scores in this sample, while the overall differences remained modest and academic discipline explained relatively little variance. :contentReference[oaicite:8]{index=8}

The policy implication is therefore not that universities should move resources from medical students to non-medical students. A more defensible interpretation is that mental-health strategies should not rely on academic discipline as a simple proxy for individual need.

Thresholds can make tiny numerical differences look dramatic

Another source of confusion appears when continuous scores are converted into categories.

Consider two students with hypothetical PHQ-9 scores of 9 and 10. Numerically, they differ by one point. Yet conventional severity categories place 9 in the mild range and 10 in the moderate range. The categorical labels can make the difference appear more dramatic than the underlying score change.

Categories are useful for communication and structured interpretation, but biological and psychological states do not suddenly transform when a questionnaire moves across an arbitrary integer boundary. Someone scoring 10 is not necessarily fundamentally different from someone scoring 9.

Thresholds should therefore support interpretation rather than replace it. Scores, symptom patterns, functional impairment, history, duration, context, and clinical assessment all matter.

Look at distributions, not just averages

Group averages can conceal substantial heterogeneity.

In our study, the PHQ-9 severity distributions provided more context than the one-point median difference alone. Non-medical students were less often in the minimal category and somewhat more often represented in higher symptom categories. At the same time, students from both disciplines appeared throughout the severity spectrum. :contentReference[oaicite:9]{index=9}

This is why good reporting often combines several layers of evidence:

  • raw means or medians;
  • measures of variability;
  • effect sizes;
  • confidence intervals;
  • severity distributions when clinically appropriate;
  • adjusted analyses;
  • model explanatory power;
  • and study limitations.

Each contributes something different. No single statistic tells the entire story.

What researchers should write instead of “Group A was significantly worse”

Language matters because readers frequently interpret “significantly higher” as “substantially higher.” More informative reporting would state the direction, magnitude, uncertainty, and context together.

For example, instead of saying, “Non-medical students had significantly greater depression,” a fuller interpretation would explain that non-medical students had modestly higher PHQ-9 scores in this sample, that the association persisted after adjustment, and that academic discipline accounted for only a small proportion of overall symptom variation.

That wording is less dramatic but more scientifically informative.

The same principle applies far beyond mental-health research. Blood pressure, pain scores, quality-of-life scales, laboratory values, test scores, reaction times, and dozens of other outcomes can produce statistically significant differences whose practical importance must be evaluated separately.

What readers should remember when interpreting p-values

The most useful habit is to stop treating statistical significance as the final result. It is one part of the evidence.

When a paper reports p < 0.05, immediately ask: “How big was the difference?” Then ask how precisely it was estimated, how much distributions overlap, whether the outcome has a meaningful-change benchmark, whether the study was observational or experimental, and whether the effect is large enough to influence decisions.

For our Pakistani university sample, the answer is nuanced. There was statistical evidence that depressive symptom scores differed between academic groups. The direction was notable because non-medical students scored somewhat higher than medical students. But the magnitude was modest, the regression models explained relatively little overall variation, and the cross-sectional design prevents causal conclusions. :contentReference[oaicite:10]{index=10}

That nuance is not a weakness. It is what responsible statistical interpretation looks like.

The broader lesson

A one-point difference can be scientifically interesting without being individually transformative. Statistical significance can tell us that a pattern is unlikely to be adequately summarized as exactly zero under the assumptions of the model. It cannot tell us, on its own, how much the pattern matters.

To answer that second question, we need effect sizes, confidence intervals, measurement properties, meaningful-change benchmarks, distributions, study design, external evidence, and real-world context.

Researchers should therefore resist both extremes: dismissing every small effect as irrelevant and exaggerating every low p-value as important. The better question is not simply, “Was it statistically significant?” It is, “How large is the effect, how certain are we about it, and what does it mean in the setting where the result will actually be used?”

Medical disclaimer: This article is intended for educational discussion of research methods and mental-health measurement. PHQ-9 and DASS-21 scores are screening or symptom-severity measures and should not be used alone to diagnose an individual or make treatment decisions. Anyone concerned about their mental health should seek assessment from an appropriately qualified healthcare professional.

Key takeaways

  • Statistical significance does not measure the size or practical importance of an effect.
  • In the source study, medical and non-medical students differed by only one point in median PHQ-9 score despite a statistically significant comparison.
  • A group-level difference should not be interpreted as equivalent to a clinically meaningful change within an individual.
  • Effect sizes, confidence intervals, score distributions, explanatory power, and study design should be interpreted alongside p-values.
  • Cross-sectional associations cannot establish that academic discipline causes differences in depression or anxiety symptoms.
  • Small population effects can still be useful when they challenge assumptions or influence decisions affecting large groups.

Frequently asked questions

Can a one-point difference really be statistically significant?
Yes. Statistical significance depends on the difference relative to its uncertainty, which is affected by sample size and variability. A small difference can therefore produce a low p-value when it is estimated precisely.
Does statistical significance mean a difference is clinically important?
No. Statistical significance and clinical importance answer different questions. A p-value evaluates evidence against a statistical null hypothesis, while clinical importance concerns whether an effect is large enough to matter for symptoms, functioning, treatment, or decisions.
Is a one-point PHQ-9 difference clinically meaningful?
A one-point group difference should not automatically be interpreted as a clinically meaningful individual change. Studies of meaningful PHQ-9 change generally evaluate larger changes, and the appropriate benchmark varies by population, setting, and purpose.
Why does sample size affect statistical significance?
Larger samples generally estimate population differences more precisely. As uncertainty decreases, even a modest effect can become statistically distinguishable from zero without becoming any larger in practical terms.
What should I examine besides a p-value?
Look at the raw difference, effect size, confidence interval, variability, distribution of scores, meaningful-change thresholds when appropriate, model explanatory power, study design, and whether the result is relevant to real-world decisions.
Did the Pakistani student study prove that being a non-medical student causes worse depression?
No. The study was cross-sectional and identified an association between academic discipline and symptom scores. It cannot establish that academic discipline caused the observed difference.

References

  1. Asghar T, Hassan A, Sahar I, et al. Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students: a cross-sectional study in Pakistan. BMC Psychology. 2026. https://doi.org/10.1186/s40359-026-05332-5
  2. Kroenke K, Spitzer RL, Williams JBW. The PHQ-9: Validity of a Brief Depression Severity Measure. Journal of General Internal Medicine. 2001;16(9):606-613. https://pmc.ncbi.nlm.nih.gov/articles/PMC1495268/
  3. Bauer-Staeb C, Kounali DZ, Welton NJ, et al. Effective dose 50 method as the minimal clinically important difference: Evidence from depression and anxiety patient-reported outcomes. Journal of Clinical Epidemiology. 2021. https://pmc.ncbi.nlm.nih.gov/articles/PMC8485844/