Skip to content
TAtaimoorasghar.com

August 20, 2026 · 15 min read

PHQ-9 vs DASS-21: How Well Do They Measure Student Mental Health?

PHQ-9 and DASS-21 both performed well in our Pakistani student study, but they measure different aspects of mental health and are not interchangeable.

The PHQ-9 and DASS-21 can both provide useful information about student mental health, but they answer different questions. The PHQ-9 is a focused nine-item measure of depressive symptoms, while the DASS-21 measures three related domains: depression, anxiety, and stress. In our study of 602 university students in Lahore, Pakistan, both instruments showed good internal consistency and meaningful psychometric performance. However, very strong correlations among the DASS-21’s latent depression, anxiety, and stress factors suggested that these subscales were not as sharply separated in this student population as their labels might imply.

By Taimoor Asghar

Why compare the PHQ-9 and DASS-21?

Mental-health research among university students often depends on self-report questionnaires. These instruments make it possible to collect standardized information from hundreds or thousands of participants without conducting a full clinical interview with every student. The choice of questionnaire, however, affects what researchers can conclude.

The Patient Health Questionnaire-9, usually called the PHQ-9, concentrates on depressive symptoms. The Depression Anxiety Stress Scales-21, or DASS-21, covers a broader spectrum of psychological distress by dividing 21 questions into depression, anxiety, and stress subscales.

That difference creates a practical question for student mental-health research: is it better to use a focused depression measure or a broader multidimensional measure?

There is no universal winner. Instrument selection should depend on the research question, population, psychometric performance, and intended interpretation. Our recently published BMC Psychology study offered an opportunity to examine both instruments simultaneously in the same group of Pakistani university students. The full study, our BMC Psychology investigation of PHQ-9 and DASS-21 psychometric properties among Pakistani university students, included reliability testing, confirmatory factor analysis, item response theory, regression modelling, measurement invariance testing, and symptom network analysis.

PHQ-9 vs DASS-21 at a glance

FeaturePHQ-9DASS-21
Number of items921
Main purposeAssessment of depressive symptom severityAssessment of depression, anxiety, and stress symptoms
Number of symptom domainsOne primary depression scaleThree seven-item subscales
Typical useDepression screening and severity measurementBroader assessment of negative emotional states
Clinical diagnosis?No; screening or severity information must be interpreted clinicallyNo; it does not independently establish a psychiatric diagnosis
Burden on respondentsVery shortLonger but still relatively brief

The distinction in the final row matters when researchers are planning large university surveys. Nine additional or unnecessary questions can reduce completion quality when questionnaires are already lengthy. Conversely, choosing the shorter PHQ-9 means giving up direct measurement of anxiety and stress.

What exactly does the PHQ-9 measure?

The PHQ-9 was developed as a brief measure of depressive symptoms. Its nine symptom questions correspond closely to the major symptom domains used in the diagnostic framework for depressive disorders, including depressed mood, loss of interest, sleep disturbance, fatigue, appetite changes, concentration difficulties, psychomotor changes, feelings of worthlessness or failure, and thoughts related to death or self-harm.

Responses refer to symptoms experienced over the previous two weeks. Each item contributes to a total score from 0 to 27, with higher scores indicating greater depressive symptom severity.

The original validation study by Kroenke, Spitzer, and Williams found that the PHQ-9 was useful as both a measure of depressive symptom severity and a screening instrument. That does not mean a PHQ-9 score by itself proves that someone has major depressive disorder. Diagnosis requires appropriate clinical assessment, including consideration of impairment, alternative explanations, medical conditions, substance use, psychiatric history, and other contextual factors.

Why the PHQ-9 works well in student research

The PHQ-9 has several practical advantages in university populations. It is short, straightforward to score, widely studied, and produces a single depression severity measure that is relatively easy to analyse statistically. Researchers can examine mean or median symptom scores, associations with demographic or behavioral factors, changes over time, or the proportion of participants crossing a predefined screening threshold.

Its brevity is particularly valuable when depression represents only one part of a larger questionnaire containing variables such as sleep, academic characteristics, financial circumstances, social support, lifestyle, or health behaviors.

The trade-off is equally clear: the PHQ-9 is primarily about depression. It cannot substitute for a dedicated measure of anxiety or perceived stress.

What does the DASS-21 measure?

The DASS-21 takes a different approach. It contains 21 questions divided into three seven-item subscales designed to characterize depression, anxiety, and stress. Rather than attempting to diagnose specific psychiatric disorders, the instrument measures related dimensions of negative emotional experience.

The depression scale emphasizes experiences such as low positive affect, hopelessness, reduced enthusiasm, and feelings of worthlessness. The anxiety component captures symptoms associated with physiological arousal, fear, and anxious experience. The stress component reflects states such as difficulty relaxing, irritability, impatience, and tension.

This makes the DASS-21 attractive for university studies where investigators want a broader description of psychological distress instead of depression alone. Academic pressure, examinations, uncertainty about employment, financial problems, family expectations, changing social relationships, and sleep disruption may manifest through overlapping depressive, anxious, and stress-related experiences.

What our Pakistani university student study found

Our cross-sectional study included 602 undergraduate university students in Lahore: 424 medical students and 178 non-medical students. Rather than examining only symptom prevalence or group differences, we also investigated whether the instruments behaved psychometrically as expected in this population.

Several findings are especially relevant when deciding between the PHQ-9 and DASS-21.

1. Internal consistency was good for both instruments

Cronbach’s alpha values across the scales ranged from 0.82 to 0.88. In practical terms, the items within the PHQ-9 and DASS-21 domains showed good internal consistency in this sample.

Reliability is necessary because researchers want items within a scale to provide reasonably coherent information about the underlying construct. However, a high alpha should never be interpreted as proof that a questionnaire is valid, one-dimensional, diagnostically accurate, or suitable for every cultural setting. Reliability is only one component of psychometric evaluation.

2. The DASS-21 three-factor model showed acceptable overall fit

Confirmatory factor analysis of the DASS-21 produced a comparative fit index of 0.954, Tucker-Lewis index of 0.948, and root mean square error of approximation of 0.052. Taken together, these findings indicated adequate fit of the specified factor structure in the study population.

At first glance, that appears reassuring for treating depression, anxiety, and stress as related but distinguishable constructs. Another result, however, deserves equal attention.

3. DASS-21 depression, anxiety, and stress were extremely strongly correlated

The latent correlation between depression and stress was 0.939, while the correlation between anxiety and stress was 0.949. Correlations approaching one raise an important psychometric question: are the subscales capturing clearly distinct constructs, or are they partly reflecting a more general dimension of psychological distress?

This does not make the DASS-21 unusable. The overall factor model still demonstrated adequate fit, and the instrument showed good reliability. Instead, it means researchers should be cautious about interpreting small differences among DASS depression, anxiety, and stress scores as though each represents a completely independent psychological phenomenon.

That caution may be particularly relevant in student populations, where symptoms frequently overlap. Poor sleep may coexist with irritability, reduced concentration, fatigue, worry, low motivation, and feelings of being overwhelmed. Statistical labels can separate these experiences into domains more cleanly than they are actually experienced by students.

4. The PHQ-9 provided a more focused depression measure

The PHQ-9 has a narrower objective. That can be a strength rather than a limitation when the primary research question concerns depression. Instead of attempting to distinguish among several highly correlated forms of distress, it measures a defined group of depressive symptoms using only nine items.

Our results showed that PHQ-9 scores detected meaningful differences within the student sample. Non-medical students had a median PHQ-9 score of 10 compared with 9 among medical students. In adjusted analysis, academic discipline remained associated with PHQ-9 scores, although discipline accounted for only a small amount of the overall variability in symptoms.

Female gender was associated with higher symptom scores across the measured domains, and each additional hour of sleep was associated with a lower PHQ-9 score. Previous treatment for depression was a particularly strong predictor of PHQ-9 symptom severity. Because the study was cross-sectional, these associations should not be interpreted as proof of cause and effect.

5. Some symptoms were more informative than others

Item response theory allowed us to look beyond total scores and examine how much information individual symptoms contributed. Self-worth, concentration, feeling down, and appetite-related items demonstrated comparatively high discrimination, while loss of interest or anhedonia showed the lowest discrimination parameter in the reported analysis.

This illustrates why psychometric research can reveal information that simple prevalence estimates cannot. Two questionnaires may both produce reliable total scores while individual questions differ substantially in how effectively they distinguish between respondents at different levels of the underlying construct.

6. Symptom network analysis highlighted cognitive and self-evaluative symptoms

Network analysis identified self-worth, concentration, and downheartedness among the most central nodes. Such findings can help researchers understand how symptoms relate to one another instead of viewing mental health solely through a single total score.

Network centrality should nevertheless be interpreted cautiously. Stability coefficients in the study ranged from 0.31 to 0.44, so the results are better viewed as informative exploratory evidence than as a definitive ranking of which student symptoms matter most.

Does the DASS-21 measure three genuinely separate problems?

This is one of the most interesting questions raised by our findings. The DASS-21 was explicitly designed to represent depression, anxiety, and stress, and a three-factor solution can fit observed data well. Yet good model fit does not automatically mean those factors are highly distinct.

If two latent factors correlate at around 0.94 or 0.95, they share a very large amount of statistical variance. Researchers therefore need to distinguish between two ideas: whether three factors can be modelled and whether those factors are strongly discriminable from one another.

In our Pakistani student sample, the answer to the first question was broadly favorable. The answer to the second was more complicated.

This finding also illustrates why validation in the actual population of interest matters. A questionnaire that performs well in one country, language, age group, or clinical setting should not automatically be assumed to have an identical structure in another population. Cultural differences in describing distress, translation, educational background, stigma, response patterns, and symptom interpretation can influence psychometric results.

Which tool should student researchers choose?

Choose the PHQ-9 when depression is the central outcome

If the main question is whether depressive symptom burden differs between groups or is associated with sleep, academic factors, social circumstances, or another exposure, the PHQ-9 offers a compact and well-established outcome measure.

It is also appealing for repeated measurement because asking nine items creates relatively little respondent burden. For longitudinal surveys or intervention studies, that efficiency may improve participation.

Choose the DASS-21 when the research question covers broader distress

If investigators genuinely need information about depressive, anxious, and stress-related symptoms, the DASS-21 collects all three within one questionnaire. This can be more efficient than combining several entirely separate instruments.

The researcher should still examine the correlations between the three domains and, ideally, test whether the assumed factor structure works in the study population. Reporting three named subscales without considering their discriminant validity can create a misleading impression of precision.

Use both when the psychometric comparison itself is valuable

There are situations where administering both instruments is justified. Our study is one example. Using the PHQ-9 and DASS-21 simultaneously made it possible to compare a targeted depressive-symptom measure with a broader distress instrument and investigate their properties through several complementary statistical methods.

For routine surveys, however, adding both questionnaires merely because they are popular can produce unnecessary redundancy. Every survey item should have a clear analytical purpose.

PHQ-9 and DASS-21 scores are not diagnoses

This distinction is fundamental. A questionnaire score is evidence about reported symptoms, not a complete psychiatric evaluation.

A student may score highly because of an emerging depressive disorder, an anxiety disorder, acute stress, grief, sleep deprivation, a medical problem, substance-related effects, major personal circumstances, or several overlapping factors. Conversely, a low questionnaire score does not guarantee the absence of clinically important difficulties.

The PHQ-9 is closely linked to depressive symptom criteria and has established usefulness as a screening and severity instrument, but a positive screen requires appropriate evaluation when used clinically. The DASS-21 was designed to characterize dimensions of depression, anxiety, and stress rather than independently diagnose DSM or ICD disorders.

This becomes especially important when prevalence is reported in research. Statements such as “X percent of students were depressed” can overstate what a self-report screening questionnaire actually establishes. More precise language is often “X percent had scores above the specified screening threshold” or “X percent reported depressive symptoms within a particular severity range.”

Why measurement quality matters for university mental-health policy

Questionnaire choice may seem like a technical methodological issue, but it has practical consequences. Universities may use research findings to decide whether counselling services need expansion, which student groups require attention, whether sleep or academic support should be targeted, and whether mental-health problems appear concentrated within a particular faculty.

If the instrument does not perform well in the target population, apparent differences between groups may partly reflect measurement rather than genuine differences in mental health.

Our study also cautions against assuming that medical students automatically represent the group with the greatest mental-health burden. In this sample, non-medical students reported higher depression and anxiety scores after adjustment, although academic discipline explained only a small proportion of the overall variance. The larger message is that student mental health cannot be reduced to degree programme alone.

Good measurement therefore requires more than selecting a famous questionnaire. Researchers should consider reliability, construct validity, factor structure, measurement invariance, individual-item performance, cultural context, and the exact question the study is attempting to answer.

Limitations of comparing PHQ-9 and DASS-21 in our study

Our findings should be interpreted within the design of the study. The data were cross-sectional, meaning exposure and symptom information were collected at one point in time. Associations with sleep, gender, academic discipline, or previous treatment therefore cannot establish causal relationships.

The participants were undergraduate students recruited from universities in Lahore, so results should not automatically be generalized to every university student in Pakistan or to students in other countries. Self-report questionnaires are also influenced by recall, willingness to disclose symptoms, language interpretation, and individual response style.

Psychometric performance can vary between populations. The very strong DASS-21 factor correlations observed here are particularly important to investigate in independent Pakistani samples before drawing broad conclusions about the instrument’s dimensional structure.

Finally, psychometric statistics answer different questions. Cronbach’s alpha evaluates internal consistency; confirmatory factor analysis evaluates a hypothesized measurement structure; item response theory investigates item characteristics; network analysis explores relationships among symptoms; and measurement invariance asks whether measurement functions comparably across groups. No single statistic proves that an instrument is universally valid.

The bottom line: PHQ-9 vs DASS-21

For student mental-health research, the PHQ-9 and DASS-21 should be viewed as complementary tools rather than competitors. The PHQ-9 is particularly useful when depression is the primary outcome and survey efficiency matters. The DASS-21 provides broader coverage when investigators need depression, anxiety, and stress information within a single measure.

In our 602-student Pakistani sample, both showed good internal consistency, and the DASS-21 demonstrated adequate overall confirmatory factor-analysis fit. At the same time, its depression, anxiety, and stress factors were extremely strongly correlated, suggesting that researchers should be cautious about treating the three subscales as completely distinct constructs in this population.

The most appropriate questionnaire is therefore not simply the one with more questions or more subscales. It is the instrument that best matches the research question, demonstrates acceptable performance in the target population, and is interpreted within its psychometric and clinical limitations.

Medical disclaimer: This article is for educational and research information only. The PHQ-9 and DASS-21 are self-report assessment instruments and are not substitutes for an individualized evaluation by a qualified healthcare professional. Anyone experiencing significant psychological distress, thoughts of self-harm, or concerns about their mental health should seek appropriate professional assessment and urgent assistance when necessary.

Key takeaways

  • The PHQ-9 is a focused nine-item measure of depressive symptoms, while the DASS-21 assesses depression, anxiety, and stress across 21 items.
  • In our study of 602 Pakistani university students, internal consistency across the PHQ-9 and DASS-21 measures was good, with Cronbach's alpha values ranging from 0.82 to 0.88.
  • The DASS-21 three-factor model showed adequate overall fit, but depression, anxiety, and stress latent factors were extremely strongly correlated.
  • High DASS-21 subscale correlations mean researchers should be cautious about treating depression, anxiety, and stress as completely independent constructs in this population.
  • The PHQ-9 is generally more efficient when depression is the primary outcome, while the DASS-21 is useful when broader psychological distress is central to the research question.
  • Neither the PHQ-9 nor DASS-21 should be interpreted as a substitute for an individualized clinical diagnosis.

Frequently asked questions

What is the main difference between the PHQ-9 and DASS-21?
The PHQ-9 is a nine-item instrument focused specifically on depressive symptoms, whereas the DASS-21 contains 21 items divided into depression, anxiety, and stress subscales. The PHQ-9 is therefore more focused and shorter, while the DASS-21 provides broader information about psychological distress.
Which performed better in the Pakistani university student study?
Both instruments demonstrated good internal consistency, with Cronbach’s alpha values across the measures ranging from 0.82 to 0.88. The DASS-21 also showed adequate confirmatory factor-analysis fit, although its depression, anxiety, and stress latent factors were extremely highly correlated. The findings therefore support the usefulness of both measures while raising caution about how distinctly the DASS-21 subscales should be interpreted in this population.
Can the PHQ-9 diagnose depression in a university student?
No questionnaire score should be treated as a stand-alone diagnosis. The PHQ-9 is a validated depression screening and symptom-severity measure, but diagnosis requires appropriate clinical assessment and consideration of the person’s symptoms, functioning, history, medical circumstances, and alternative explanations.
Does the DASS-21 diagnose depression, anxiety disorders, or stress disorders?
No. The DASS-21 measures dimensions of depression, anxiety, and stress-related symptoms. Its scores can describe psychological distress and are useful in research and assessment, but they do not independently establish a psychiatric diagnosis.
Why do high correlations between DASS-21 subscales matter?
Very high correlations suggest that the depression, anxiety, and stress subscales may share substantial underlying variance. In the Pakistani student study, depression-stress and anxiety-stress latent correlations were approximately 0.94 and 0.95. This means researchers should be cautious about interpreting the three domains as completely separate psychological constructs in that population.
Should a student mental-health study use both PHQ-9 and DASS-21?
Using both can be justified when the research question specifically requires comparison of depression measurement with broader psychological distress or when psychometric properties are being investigated. For many routine surveys, however, using both may be unnecessarily repetitive, and instrument selection should follow the study’s primary objectives.

References

  1. Asghar T, Hassan A, Sahar I, Tahir M, Shahid B, Komal K. Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students: a cross-sectional study in Pakistan. BMC Psychology. 2026. https://doi.org/10.1186/s40359-026-05332-5
  2. Kroenke K, Spitzer RL, Williams JBW. The PHQ-9: Validity of a Brief Depression Severity Measure. Journal of General Internal Medicine. 2001;16(9):606-613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x
  3. Lovibond SH, Lovibond PF. Manual for the Depression Anxiety Stress Scales. 2nd ed. Sydney: Psychology Foundation; 1995. https://www2.psy.unsw.edu.au/groups/dass/