Skip to content
TAtaimoorasghar.com

August 20, 2026 · 16 min read

PHQ-9 and DASS-21 in University Students: What Our Pakistani Study Found

Our Pakistani study of 602 university students examined PHQ-9 and DASS-21 scores, reliability, symptom patterns, and differences by discipline.

By Taimoor Asghar

In our study of 602 university students in Lahore, Pakistan, the PHQ-9 and DASS-21 both performed adequately as measures of psychological symptoms, but they also revealed an important caution: questionnaire scores are more informative when we look beyond a single total or diagnostic-looking cutoff. Non-medical students reported somewhat higher depression and anxiety scores than medical students, female students reported higher symptom levels across the measured domains, and several cognitive and self-worth-related symptoms emerged as particularly informative. At the same time, academic discipline explained only a small fraction of the overall variation in mental-health scores.

The findings matter because the PHQ-9 and DASS-21 are widely used in student mental-health research. A questionnaire can be reliable and still require careful interpretation. It can detect differences between groups without proving that belonging to one group causes poorer mental health. And a statistically significant result does not automatically mean that a single characteristic, such as being a medical or non-medical student, is the dominant explanation for an individual’s symptoms.

This article explains what we found, what the psychometric analyses tell us about the PHQ-9 and DASS-21 in Pakistani university students, and how these results should—and should not—be interpreted.

Why Study the PHQ-9 and DASS-21 in University Students?

University is a period when young adults often face overlapping academic, financial, social, family, and career pressures. Mental-health surveys therefore frequently use short self-report instruments to estimate the burden and severity of symptoms without requiring a lengthy clinical assessment for every participant.

Two commonly used instruments are the Patient Health Questionnaire-9, or PHQ-9, and the 21-item Depression Anxiety Stress Scales, or DASS-21. They overlap in some of the psychological experiences they capture, but they were designed differently and should not be treated as interchangeable diagnostic tests.

What does the PHQ-9 measure?

The PHQ-9 contains nine items corresponding closely to core depressive symptoms. Respondents indicate how frequently they have been bothered by each symptom during the preceding two weeks. The total score ranges from 0 to 27, with higher scores reflecting a greater burden of depressive symptoms.

Although thresholds are often used to categorize symptom severity or identify people who may benefit from further assessment, the PHQ-9 is fundamentally a screening and severity instrument. A score alone does not establish a psychiatric diagnosis. Clinical context, duration, impairment, differential diagnoses, medical conditions, substance use, and safety concerns may all need separate evaluation.

What does the DASS-21 measure?

The DASS-21 contains three seven-item domains intended to measure depression, anxiety, and stress. This makes it attractive for university surveys because researchers can examine several dimensions of psychological distress within one relatively short questionnaire.

The distinction between these domains is conceptually useful. Depression-related items tend to emphasize low mood, reduced positive affect, and hopelessness; anxiety items include manifestations of anxious arousal and fear; and stress items capture tension, irritability, difficulty relaxing, and related experiences. However, real-world psychological symptoms overlap. One of the aims of psychometric analysis is therefore to test whether the proposed domains are sufficiently distinguishable in the population being studied.

What Our Pakistani PHQ-9 and DASS-21 Study Examined

Our cross-sectional study included 602 undergraduate university students in Lahore, Pakistan: 424 medical students and 178 non-medical students. We used both the PHQ-9 and DASS-21 and examined considerably more than simple average scores.

The analyses included descriptive statistics, comparisons between medical and non-medical students, multivariable regression, confirmatory factor analysis, reliability testing, item response theory, measurement invariance assessment, and symptom network analysis. The published BMC Psychology study on the PHQ-9 and DASS-21 among Pakistani university students provides the complete methodological and statistical report.

This broader analytical approach was important. If a questionnaire is going to be used to compare groups, it is useful to know not only whether the groups have different total scores but also whether the instrument behaves reasonably within the sample and whether it appears to measure comparable constructs across those groups.

Finding 1: Non-Medical Students Had Higher Depression Scores

One of the clearest findings challenged the common assumption that medical students must necessarily report the greatest psychological burden.

The median PHQ-9 score among non-medical students was 10, compared with 9 among medical students. In adjusted regression analyses, non-medical students also had higher PHQ-9 scores and higher DASS-21 depression and anxiety scores. The difference in DASS-21 stress was not statistically significant after adjustment.

This does not mean that medical education protects against depression or anxiety. Nor does it establish that non-medical courses cause poorer mental health. The study was cross-sectional, so exposure and outcome were measured at the same general point in time. It cannot establish temporal or causal relationships.

The more useful interpretation is that psychological distress was not confined to the group traditionally presumed to be at highest risk. University mental-health policies that concentrate almost exclusively on medical students may therefore overlook substantial needs elsewhere on campus.

The effect of discipline was statistically significant but small

This distinction is especially important. The regression models explained only a small proportion of overall variation in symptom scores, with R-squared values of approximately 0.029 to 0.041 across models.

In practical terms, knowing whether somebody was a medical or non-medical student provided only limited information about that person’s overall symptom burden. Many other measured and unmeasured influences likely contribute to mental health, including individual vulnerability, social relationships, finances, academic experiences, sleep, previous mental-health problems, family circumstances, and environmental factors.

This is why statistically significant group comparisons should not be converted into stereotypes. A university cannot reliably identify who needs support simply by looking at a student’s academic discipline.

Finding 2: Female Students Reported Higher Symptoms Across Domains

Female gender was associated with higher scores across the PHQ-9 and the three DASS-21 domains in the adjusted analyses, with the reported associations reaching statistical significance at p<0.001.

This finding supports the value of examining demographic differences when planning student services, but it should again be interpreted as a population-level association rather than a deterministic rule. It does not imply that every female student will experience high psychological distress or that male students have little need for mental-health support.

Universities should ideally combine broad, accessible services with approaches that recognize groups showing elevated average symptom burden. Gender-sensitive support may be useful, but screening and referral pathways should remain available to students regardless of gender.

Finding 3: Sleep Was Related to PHQ-9 Depression Scores

Sleep also emerged as a relevant correlate. Each additional hour of reported sleep was associated with a 0.37-point lower PHQ-9 score in the adjusted model.

The direction of the relationship is clinically plausible, but the design does not allow us to determine which came first. Poor sleep may contribute to emotional difficulties, depressive symptoms may interfere with sleep, or both may arise from other factors such as academic pressure, irregular schedules, physical illness, substance use, or stressful life circumstances.

That makes sleep a useful area for further study rather than a simple causal explanation. Campus mental-health strategies may reasonably include sleep education and healthier academic scheduling, but students experiencing persistent depressive symptoms should not be told that obtaining a particular number of sleep hours is a substitute for appropriate assessment or treatment.

Finding 4: Previous Depression Treatment Was a Strong Predictor

Previous treatment for depression showed the strongest association with PHQ-9 score among the predictors reported in the study, with an estimated coefficient of 4.15 points.

This makes intuitive sense: a history of treatment may identify students who previously experienced clinically significant symptoms or who remain vulnerable to recurrence. However, treatment history should not be interpreted as causing higher current scores. Rather, it can function as an indicator of prior mental-health burden.

For universities, this finding reinforces the need for continuity of care. Students arriving with an established mental-health history may benefit from clear information about how to access local counselling, psychiatric services, primary care, medication review, or crisis support where appropriate.

How Reliable Were the PHQ-9 and DASS-21?

Reliability is one of the basic questions researchers ask when evaluating a psychological scale. In our study, Cronbach’s alpha values across the examined scales ranged from 0.82 to 0.88.

Values in this range indicate good internal consistency for research use in this sample. In simple terms, items intended to assess related aspects of psychological distress showed a reasonable degree of coherence.

However, reliability is not synonymous with validity. A questionnaire can produce internally consistent responses without perfectly distinguishing every psychological construct it is intended to measure. That became particularly relevant when we examined the DASS-21 factor structure.

What the DASS-21 Factor Analysis Showed

Confirmatory factor analysis was used to examine how well the proposed DASS-21 structure fitted the observed responses. The reported fit statistics were generally adequate: the Comparative Fit Index was 0.954, Tucker-Lewis Index 0.948, Root Mean Square Error of Approximation 0.052, and Standardized Root Mean Square Residual 0.031.

At first glance, that supports the expected structure. But another result deserves equal attention: correlations between the latent DASS-21 factors were extremely high. The depression-stress correlation was 0.939, while the anxiety-stress correlation was 0.949.

Why do very high factor correlations matter?

If three supposedly distinct psychological domains are almost perfectly correlated statistically, it becomes harder to argue that each subscale is capturing a sharply independent construct within that particular population.

This does not make the DASS-21 useless. Its total and subscale scores can still provide meaningful information. Instead, the finding suggests caution when interpreting relatively small differences between its depression, anxiety, and stress domains as though they represent completely separate psychological processes.

University students may experience psychological distress in a highly interconnected way. A student struggling with persistent tension may also report sleep disruption, low motivation, worry, concentration problems, and depressed mood. Statistical overlap between scales can reflect this clinical and experiential overlap as well as properties of the questionnaire itself.

Further validation work in diverse Pakistani samples would help determine whether these very high correlations consistently appear across universities, languages, regions, and educational settings.

Measurement Invariance: Why It Matters for Medical vs Non-Medical Comparisons

Our analyses supported measurement invariance across academic discipline. This is an important technical result because comparisons between groups are difficult to interpret if a questionnaire functions fundamentally differently in each group.

Suppose, for example, that medical and non-medical students interpreted the same item in systematically different ways unrelated to their true symptom level. A difference in total score might then partly reflect measurement behavior rather than an actual difference in the underlying construct.

Evidence supporting measurement invariance gives greater confidence that the scales can be meaningfully compared across the medical and non-medical groups studied. It does not prove perfect measurement equivalence in every possible context, but it strengthens the interpretation of the observed group comparisons.

Which PHQ-9 and DASS-21 Symptoms Were Most Informative?

Another distinctive part of the study was the use of item response theory. Rather than treating every questionnaire item as equally informative, item response theory examines how strongly individual items distinguish between people at different levels of the underlying symptom trait.

Self-worth, concentration, feeling down, and appetite-related items showed relatively high discrimination in the item response analyses. The PHQ-9 interest or anhedonia item showed the lowest discrimination, with an estimated discrimination parameter of 0.468.

This does not mean anhedonia is clinically unimportant. Loss of interest remains a core depressive symptom. Psychometric discrimination has a narrower meaning: within this sample and statistical model, that item differentiated levels of the underlying measured trait less strongly than several other items.

This distinction matters because questionnaire research should not be translated mechanically into clinical priorities. An item may have lower statistical discrimination while remaining highly important during an individual clinical assessment.

What Symptom Network Analysis Added

Traditional questionnaires often assume that symptoms are indicators of an underlying disorder or latent construct. Network analysis offers another perspective by examining symptoms as an interconnected system.

In our network analysis, self-worth, concentration, and downheartedness emerged among the most central symptoms. This broadly complemented the item-response findings, suggesting that cognitive and self-evaluative symptoms may occupy particularly informative positions within the symptom pattern observed in these students.

However, network centrality should not be interpreted as proof that targeting one central symptom will necessarily improve all others. The reported stability coefficients ranged from 0.31 to 0.44, which argues for caution when ranking nodes too precisely. Network results are useful for generating hypotheses, but replication in independent samples is important before drawing strong intervention conclusions.

PHQ-9 Versus DASS-21: Which Should Universities Use?

There is no universal answer because the instruments serve somewhat different purposes.

  • Use the PHQ-9 when depression is the primary construct of interest. It is concise, focuses specifically on depressive symptoms, and is commonly used in both research and clinical screening contexts.
  • Use the DASS-21 when the goal is to examine broader psychological distress. It provides depression, anxiety, and stress domains in a single instrument.
  • Do not treat either questionnaire as a stand-alone psychiatric diagnosis. Elevated scores indicate symptoms that may warrant attention or further assessment.
  • Consider the population. Psychometric performance established in one language, university, region, or country should not automatically be assumed to generalize perfectly to another.
  • Have a response pathway before screening. Collecting information about severe psychological distress without having an appropriate referral or safety process can create ethical and practical problems.

What These Findings Mean for Pakistani Universities

The broader message is that campus mental-health policy should move beyond the assumption that one faculty or professional course uniquely defines risk.

In our sample, non-medical students reported higher average depression and anxiety scores after adjustment, yet academic discipline accounted for only a limited share of overall variation. That combination is important. It suggests that universities should recognize group-level differences without allowing those differences to dictate access to support.

A more sensible model would include accessible mental-health information for all students, confidential routes to professional assessment, faculty and staff awareness of warning signs, appropriate referral pathways, support for students with previous mental-health treatment, and attention to modifiable aspects of student life such as disruptive schedules and sleep health.

Screening tools can support these systems, but screening without services is not a complete mental-health strategy.

Important Limitations of the Study

Several limitations should shape interpretation of the findings.

First, the research was cross-sectional. Associations between discipline, gender, sleep, treatment history, and psychological symptoms cannot establish causation or the direction of an effect.

Second, the study relied on self-reported questionnaires. PHQ-9 and DASS-21 scores represent reported symptoms, not clinician-confirmed diagnoses.

Third, the participants were university students recruited in Lahore, Pakistan. Results should not automatically be generalized to every Pakistani university, all age groups, or populations outside higher education.

Fourth, although the models detected statistically significant predictors, their R-squared values were low. Most variation in symptom scores remained unexplained by the included predictors.

Finally, the very high correlations between DASS-21 latent factors raise an important psychometric question about how clearly depression, anxiety, and stress are differentiated by the instrument in this population. That finding deserves replication rather than overinterpretation from one dataset.

The Main Lesson From Our PHQ-9 and DASS-21 Study

The strongest lesson is not that one group of students is simply more or less mentally healthy than another. It is that university psychological distress is multidimensional, widely distributed, and poorly captured by stereotypes.

The PHQ-9 and DASS-21 demonstrated good internal consistency and generally adequate psychometric performance in our Pakistani sample. The instruments also highlighted meaningful differences and symptom patterns. Non-medical students reported somewhat higher depression and anxiety symptoms, women reported higher scores across measured domains, greater reported sleep duration was associated with lower PHQ-9 scores, and previous depression treatment was strongly associated with current symptom burden.

At the item and network levels, self-worth, concentration, and mood-related symptoms appeared particularly informative. At the same time, the very high correlations among DASS-21 factors remind us that depression, anxiety, and stress may not separate neatly when measured in real students experiencing overlapping forms of distress.

For researchers, the findings support looking beyond Cronbach’s alpha and total scores toward factor structure, measurement equivalence, item characteristics, and symptom relationships. For universities, the practical implication is broader: mental-health services should be designed for the whole student population rather than based on assumptions about which academic discipline ought to be most distressed.

Frequently Asked Questions

Can the PHQ-9 diagnose depression in a university student?

No. The PHQ-9 is a validated symptom-screening and severity measure, but an individual score does not by itself establish a clinical diagnosis of a depressive disorder. Diagnosis requires appropriate clinical assessment and consideration of the person’s broader circumstances.

Does a high DASS-21 score mean someone has depression or an anxiety disorder?

Not necessarily. DASS-21 scores describe levels of self-reported depression-, anxiety-, and stress-related symptoms. They should not be interpreted as equivalent to a formal psychiatric diagnosis.

Did medical students have worse mental health in this study?

No. In this sample, non-medical students had higher PHQ-9 scores and higher adjusted DASS-21 depression and anxiety scores. However, academic discipline explained only a small proportion of overall symptom variation, so the result should not be used to stereotype either group.

Were the PHQ-9 and DASS-21 reliable in Pakistani university students?

Internal consistency was good in this study, with Cronbach’s alpha values between 0.82 and 0.88. The DASS-21 also showed adequate confirmatory factor-analysis fit, although very high correlations between its latent factors raised questions about how distinct the depression, anxiety, and stress subscales were in this sample.

What symptoms appeared most informative?

Self-worth, concentration, feeling down, and appetite-related items showed relatively high discrimination in item-response analyses. Self-worth, concentration, and downheartedness were also prominent in the symptom network. These are research findings and should not be used to reduce an individual clinical assessment to only a few symptoms.

Medical disclaimer: This article is for educational and research communication purposes and is not a substitute for individualized medical or mental-health assessment. PHQ-9 and DASS-21 scores should not be used alone to diagnose or exclude a psychiatric condition. Anyone experiencing persistent or severe psychological symptoms, significant functional impairment, thoughts of self-harm, or concerns about personal safety should seek timely assessment from an appropriately qualified healthcare or mental-health professional.

Key takeaways

  • Both the PHQ-9 and DASS-21 showed good internal consistency in the 602-student Pakistani sample, with Cronbach's alpha values ranging from 0.82 to 0.88.
  • Non-medical students reported higher depression and anxiety scores than medical students, but academic discipline explained only a small proportion of the overall variation in symptoms.
  • Female students had higher adjusted scores across the PHQ-9 and DASS-21 domains, while additional reported sleep was associated with lower PHQ-9 scores.
  • The DASS-21 showed adequate confirmatory factor-analysis fit, but very high correlations between depression, anxiety, and stress factors raised questions about how distinct the subscales were in this population.
  • Self-worth, concentration, and mood-related symptoms emerged as particularly informative in item-response and symptom-network analyses.
  • PHQ-9 and DASS-21 scores are screening and research measures rather than stand-alone psychiatric diagnoses.

Frequently asked questions

Can the PHQ-9 diagnose depression in a university student?
No. The PHQ-9 measures depressive symptom severity and can support screening, but a score alone does not establish a clinical diagnosis. Appropriate clinical assessment is required.
Does a high DASS-21 score mean someone has a psychiatric disorder?
Not necessarily. The DASS-21 measures self-reported depression-, anxiety-, and stress-related symptoms. Its scores should not be treated as equivalent to formal psychiatric diagnoses.
Did medical students report more depression than non-medical students?
No. In this Pakistani sample, non-medical students had a median PHQ-9 score of 10 compared with 9 among medical students and also had higher adjusted depression and anxiety scores. Academic discipline nevertheless explained only a small proportion of overall variation.
Were the PHQ-9 and DASS-21 reliable in the Pakistani study?
Yes. Cronbach’s alpha values ranged from 0.82 to 0.88, indicating good internal consistency in the sample. The DASS-21 showed adequate factor-model fit, although very high correlations among its latent factors suggested caution when treating its three domains as sharply distinct.
Which symptoms were especially informative in the study?
Self-worth, concentration, feeling down, and appetite-related items showed relatively high discrimination in item-response analyses, while self-worth, concentration, and downheartedness were prominent in the symptom network.

References

  1. Asghar T, Hassan A, Sahar I, Tahir M, Shahid B, Komal K. Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students: a cross-sectional study in Pakistan. BMC Psychology. 2026. https://doi.org/10.1186/s40359-026-05332-5