Skip to content
TAtaimoorasghar.com

August 17, 2026 · 15 min read

Can PHQ-9 and DASS-21 Tell Us More Than a Depression Score?

PHQ-9 and DASS-21 can reveal more than symptom totals, including item performance, symptom patterns, scale structure, and differences between groups.

Yes. The PHQ-9 and DASS-21 can tell researchers considerably more than whether a depression score is high or low. When their individual items and underlying measurement properties are examined, these questionnaires can reveal which symptoms carry the most information, how depression overlaps with anxiety and stress, whether scales behave similarly across groups, and which symptoms occupy prominent positions within a broader symptom network. A recent study of 602 university students in Lahore, Pakistan illustrates how these familiar questionnaires can be used as psychometric tools rather than treated simply as score-generating checklists.

By Taimoor Asghar

Why a total depression score tells only part of the story

Mental-health questionnaires are commonly summarized into a single number. A student completes nine PHQ-9 items, the responses are added together, and the resulting score represents overall depressive symptom severity. DASS-21 responses can similarly be combined into depression, anxiety, and stress subscale scores.

This approach is useful because total scores provide a concise way to compare individuals or groups. They can support screening, epidemiological research, outcome monitoring, and statistical analysis. The original PHQ-9 validation study established the instrument as a brief measure of depressive symptom severity, with nine items corresponding to major depressive symptoms and a possible total score from 0 to 27.

However, two people with the same total score do not necessarily have the same symptom profile. One student might report low mood, concentration difficulties, impaired self-worth, and sleep disturbance. Another may reach a similar score through reduced interest, fatigue, appetite changes, and psychomotor symptoms. Their totals may match even though the experiences contributing to those totals differ substantially.

That distinction matters in research. If investigators examine only the summed score, they may miss information contained in the individual items, relationships between symptoms, or the structure of the questionnaire itself.

What the PHQ-9 and DASS-21 actually measure

The PHQ-9

The Patient Health Questionnaire-9 is a nine-item self-report instrument designed around depressive symptoms. Respondents indicate how frequently each symptom has affected them during the preceding two weeks. The items cover areas such as depressed mood, reduced interest or pleasure, sleep disturbance, fatigue, appetite changes, feelings of worthlessness or failure, concentration problems, psychomotor changes, and thoughts related to death or self-harm.

The PHQ-9 is widely used as a depression screening and severity measure. It should not, however, be interpreted as an automatic diagnosis based solely on a numerical result. Clinical diagnosis requires appropriate professional assessment, consideration of functional impairment, context, differential diagnoses, and other relevant information.

The DASS-21

The Depression Anxiety Stress Scales-21 contains 21 items divided into three seven-item domains: depression, anxiety, and stress. According to the developers of the DASS, the instrument was created to measure related negative emotional states rather than merely produce a single general distress score.

The depression component includes experiences related to dysphoria, hopelessness, self-deprecation, loss of interest, anhedonia, and inertia. The anxiety scale includes features such as autonomic arousal and anxious affect, while the stress scale focuses more on tension, difficulty relaxing, irritability, agitation, and persistent arousal.

Using PHQ-9 and DASS-21 together therefore creates an opportunity to examine depressive symptoms from two related but non-identical measurement systems while also evaluating anxiety and stress.

What our Pakistani student study added beyond total scores

Our 2026 BMC Psychology study, Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students, included 602 undergraduate students from universities in Lahore, Pakistan. Of these, 424 were medical students and 178 were studying non-medical disciplines.

The study did compare conventional questionnaire scores. Non-medical students had a median PHQ-9 score of 10 compared with 9 among medical students. In adjusted analyses, non-medical students also had higher PHQ-9, DASS depression, and DASS anxiety scores, while the adjusted difference for DASS stress did not reach statistical significance.

Those comparisons were useful, but they represented only one part of the analysis. The study also used confirmatory factor analysis, reliability analysis, item response theory, measurement invariance testing, multivariable regression, and symptom network analysis. These methods allowed us to ask questions that cannot be answered by simply calculating an average PHQ-9 score.

1. Reliability asks whether the items work together coherently

A questionnaire is more than a collection of questions. Its items are intended to measure an underlying construct with sufficient consistency. Reliability analysis examines whether that assumption is reasonably supported by the data.

In the Pakistani student sample, Cronbach’s alpha values across the evaluated scales ranged from 0.82 to 0.88. These results supported adequate internal consistency in this dataset.

That does not mean every item is interchangeable or that reliability alone proves validity. A questionnaire can have strong internal consistency while still failing to capture the intended construct appropriately in a particular population. Conversely, extremely high similarity among items can sometimes indicate redundancy rather than superior measurement.

Reliability should therefore be considered one component of a broader psychometric evaluation rather than a final verdict on questionnaire quality.

2. Factor analysis can test what lies beneath the scores

The DASS-21 is intended to represent depression, anxiety, and stress as related but distinguishable dimensions. Confirmatory factor analysis provides a way to examine whether observed questionnaire responses are compatible with that proposed structure.

In our study, the DASS-21 model demonstrated adequate overall fit, with a comparative fit index of 0.954, Tucker-Lewis index of 0.948, and root mean square error of approximation of 0.052.

The more interesting finding appeared when the relationships between the latent factors were examined. Depression and stress correlated at 0.939, while anxiety and stress correlated at 0.949.

Those correlations are extremely high. They do not mean that depression, anxiety, and stress are literally identical constructs. They do, however, raise an important measurement question: how distinctly were these dimensions operating in this particular sample?

This is an example of information that disappears when researchers report only three subscale means. A set of apparently well-behaved total scores may still contain substantial overlap at the latent level.

Why overlap matters

Depression, anxiety, and stress frequently co-occur. Students experiencing substantial psychological distress may endorse symptoms belonging to several domains simultaneously. Some emotional and physiological experiences may also share common underlying processes.

For researchers, very strong factor correlations mean that differences between subscales should be interpreted carefully. Researchers should avoid assuming that numerical separation automatically proves psychological separation.

The finding also demonstrates why psychometric properties should ideally be assessed in the population where a scale is being used. Evidence from one country, language, clinical setting, or age group cannot automatically establish identical measurement behavior elsewhere.

3. Item response theory asks which questions are most informative

Traditional scoring usually gives each questionnaire item a similar role in the final total. Item response theory approaches the instrument differently. Instead of asking only how much each response adds to a sum, IRT can examine how individual items behave across levels of an underlying trait.

One important IRT concept is discrimination. In simplified terms, an item with stronger discrimination is better able to distinguish between respondents with different levels of the latent characteristic being measured.

In our analysis, items involving self-worth, concentration, feeling down, and appetite showed particularly strong discrimination. The interest or anhedonia item showed the lowest discrimination parameter in the reported PHQ-9 item analysis, with a value of 0.468.

This does not mean that anhedonia is clinically unimportant. Loss of interest or pleasure is a central feature of depressive disorders. IRT findings are measurement findings within a particular dataset; they should not be converted into statements about which symptoms matter most for an individual patient.

Instead, the result tells researchers that different questionnaire items may contribute different amounts of statistical information in a given population.

Why item-level information is useful

Item-level analyses can help researchers explore several questions:

  • Which symptoms best distinguish students across different levels of depressive symptom burden?
  • Are some items mainly informative at lower or higher portions of the symptom spectrum?
  • Do particular questions contribute relatively little information in a specific sample?
  • Could cultural, linguistic, educational, or contextual factors influence the way certain items are understood?
  • Do the same items perform similarly in different student groups?

These questions move the discussion beyond whether the mean PHQ-9 score is 8, 9, or 10.

4. Symptom network analysis changes the unit of attention

Another way to look beyond total scores is symptom network analysis. Traditional latent-variable approaches often conceptualize observed symptoms as manifestations of an underlying disorder. Network approaches instead examine statistical relationships among the symptoms themselves.

In a symptom network, individual questionnaire items can be represented as nodes, while estimated relationships between symptoms are represented as edges. Researchers can then examine whether particular symptoms occupy relatively central positions within the network.

In our student sample, self-worth, concentration, and downheartedness emerged among the most central nodes. Interestingly, these findings overlapped with aspects of the IRT results, where self-worth, concentration, and feeling down also demonstrated comparatively strong discrimination.

This convergence is intriguing because two different analytical approaches highlighted similar cognitive and affective symptoms. Still, it should be interpreted cautiously. A central network node is not automatically a causal driver of depression, and a highly discriminating item is not automatically the best therapeutic target.

The network stability coefficients in the study ranged from 0.31 to 0.44, reinforcing the need for restraint when interpreting the relative ranking of central symptoms. Network analysis can generate useful hypotheses about symptom organization, but cross-sectional networks cannot establish that changing one symptom will cause downstream improvements in others.

5. Measurement invariance asks whether group comparisons are fair

Researchers frequently compare questionnaire scores between groups: men and women, younger and older participants, clinical and non-clinical populations, or, as in our study, medical and non-medical students.

But a basic problem arises before any mean comparison is interpreted: does the questionnaire measure the construct in a sufficiently comparable way across those groups?

Measurement invariance analysis addresses that question. If a scale functions very differently between two groups, an observed score difference could partly reflect measurement differences rather than a true difference in the underlying construct.

Our analyses supported measurement invariance across academic discipline. This provided additional justification for comparing the relevant questionnaire constructs between medical and non-medical students.

This is a valuable example of something total scores cannot tell us by themselves. Two groups may differ by several points on a questionnaire, but without evaluating measurement equivalence, researchers have less evidence that the numbers carry the same meaning in both groups.

6. Scores can be linked with demographic and behavioral factors

Questionnaire data become more informative when examined alongside participant characteristics. Multivariable models in the study assessed associations between symptom scores and variables such as academic discipline, gender, sleep duration, and previous depression treatment.

Female gender was associated with higher scores across the evaluated symptom domains in the study. Each additional reported hour of sleep was associated with a 0.37-point lower PHQ-9 score, while previous depression treatment was the strongest predictor reported in the PHQ-9 model, with a coefficient of 4.15.

These are associations rather than demonstrations of cause and effect. The study was cross-sectional, meaning variables were measured at approximately the same point in time. We therefore cannot conclude from these results that increasing sleep by one hour would directly reduce an individual’s PHQ-9 score by 0.37 points, nor can we determine the temporal direction of all observed relationships.

The low model R-squared values, approximately 0.029 to 0.041 across the reported models, are also informative. They indicate that the measured predictors accounted for only a small fraction of overall variation in symptom scores.

That finding discourages overly simple explanations of student mental health. Academic discipline, gender, sleep, and treatment history may be statistically associated with symptom levels, but much of the variation remains related to factors not captured by those models.

A high-scoring symptom is not necessarily the most informative symptom

One of the most useful conceptual lessons from psychometric analysis is the difference between symptom frequency, symptom severity, statistical discrimination, and network centrality.

A symptom can be frequently endorsed without being especially good at distinguishing between different levels of overall depressive burden. Another symptom may be less frequent but statistically more informative. A third may occupy a highly connected position within a symptom network.

These properties answer different questions:

  • Average item score: How commonly or strongly was the symptom reported?
  • Discrimination: How effectively did the item differentiate respondents across levels of the measured trait?
  • Factor loading: How strongly was the item related to the latent factor in the specified measurement model?
  • Network centrality: How prominently was the symptom positioned within the estimated symptom network?
  • Total score contribution: How much did the response contribute numerically to the summed scale?

None of these concepts should be substituted uncritically for another.

What does this mean for university mental-health research?

University mental-health studies frequently focus on prevalence estimates or score differences. Those are valuable questions, but researchers can extract substantially more information from the same questionnaire data when study design and sample size permit.

A more comprehensive measurement strategy can investigate whether a questionnaire is reliable, whether its proposed dimensions fit the observed responses, which individual items are most informative, whether the measure behaves similarly across groups, and how symptoms relate to one another.

This can be particularly valuable in settings where commonly used psychological instruments were originally developed in different cultural or linguistic environments. Rather than assuming a scale behaves identically everywhere, researchers can evaluate that assumption empirically.

Our findings also challenge the tendency to treat medical students as the automatic reference group for university mental-health concern. In this Lahore sample, non-medical students reported higher depressive and anxiety scores after adjustment, while academic discipline explained relatively little overall variance. University mental-health strategies therefore should not be restricted to one faculty simply because that group traditionally receives more research attention.

What these analyses cannot tell us

More sophisticated statistics do not remove the limitations of the underlying data. Several boundaries remain important.

They do not establish a clinical diagnosis

PHQ-9 and DASS-21 are self-report measures. They can quantify symptoms and support screening or research, but questionnaire scores alone should not replace a comprehensive clinical assessment.

Cross-sectional analysis cannot establish causality

Associations involving sleep, academic discipline, gender, previous treatment, or individual symptoms cannot demonstrate which factor caused another when all are measured cross-sectionally.

Central symptoms are not automatically intervention targets

Network centrality can describe the structure of estimated symptom relationships. It does not prove that treating the most central node will produce the greatest improvement.

Psychometric findings are population-dependent

An item that discriminates strongly in Pakistani university students may not behave identically in another country, age group, language, or clinical population. Replication remains essential.

Total scores still have value

Looking beyond scores does not make total scores obsolete. Summed scales remain practical, interpretable, and useful. The better conclusion is that total scores and item-level psychometric analyses answer different questions and can complement one another.

The broader lesson: questionnaires are datasets, not just calculators

It is tempting to think of the PHQ-9 as nine questions that produce one depression number and the DASS-21 as 21 questions that produce three numbers. Psychometrically, however, every response contains additional information.

The pattern of responses can help researchers evaluate reliability. Relationships among items can test the proposed factor structure. Item response theory can identify questions that provide different amounts of information. Invariance testing can examine whether comparisons across groups are defensible. Network analysis can describe how symptoms are statistically interconnected. Regression models can explore how questionnaire outcomes relate to demographic, behavioral, and historical factors.

When these approaches are combined thoughtfully, questionnaires become tools for investigating the structure of psychological symptoms rather than merely devices for assigning severity categories.

What should readers take from this study?

The clearest conclusion is that depression measurement does not need to stop at the total score. In the Pakistani university sample, PHQ-9 and DASS-21 analyses provided evidence about reliability, dimensional structure, group comparability, item discrimination, and symptom connectivity in addition to conventional comparisons of depressive, anxiety, and stress scores.

Self-worth, concentration, and low-mood-related symptoms repeatedly emerged as informative features across item-response and network analyses. The DASS-21 showed adequate overall factor-model fit, yet the very high correlations among its latent dimensions suggested substantial overlap among depression, anxiety, and stress in this population. Measurement invariance supported comparisons between medical and non-medical students, while the modest explanatory power of regression models showed that academic discipline and the measured covariates captured only a small portion of the complexity underlying student symptoms.

That is considerably more information than a single depression score can provide.

Medical disclaimer: This article is for educational and research-information purposes only. PHQ-9 and DASS-21 scores should not be used by readers to diagnose themselves or another person. Mental-health symptoms, functional impairment, or thoughts of self-harm require appropriate assessment by a qualified healthcare professional. Anyone experiencing immediate danger or a mental-health emergency should seek urgent professional assistance through appropriate local services.

Key takeaways

  • PHQ-9 and DASS-21 contain useful item-level and structural information that is lost when researchers examine only total scores.
  • In a 602-student Pakistani sample, both instruments showed adequate internal consistency, with Cronbach's alpha values ranging from 0.82 to 0.88.
  • DASS-21 factor-model fit was adequate, but very high correlations among depression, anxiety and stress factors suggested substantial overlap between the constructs in this population.
  • Self-worth, concentration and low-mood-related symptoms emerged as informative across item response and symptom network analyses.
  • Measurement invariance supported comparisons between medical and non-medical students, while low regression R-squared values showed that academic discipline and measured covariates explained only a small proportion of symptom variation.
  • Psychometric and network findings describe patterns in a dataset and should not be interpreted as diagnoses, proof of causation or automatic treatment priorities.

Frequently asked questions

Can PHQ-9 tell researchers more than a total depression score?
Yes. Researchers can analyze individual PHQ-9 items using methods such as factor analysis, item response theory and symptom network analysis to examine how symptoms behave, which items provide more statistical information, and how symptoms relate to one another.
What is the main difference between PHQ-9 and DASS-21?
PHQ-9 primarily measures depressive symptoms, whereas DASS-21 contains separate scales intended to measure depression, anxiety and stress. The instruments therefore provide overlapping but different information about psychological symptoms.
Which PHQ-9 symptoms were especially informative in the Pakistani university study?
Item response analyses highlighted self-worth, concentration, feeling down and appetite items as having relatively strong discrimination. Self-worth, concentration and downheartedness also appeared among the more central symptoms in network analysis. These statistical findings should not be interpreted as proving that those symptoms are clinically more important for every individual.
Did the DASS-21 clearly separate depression, anxiety and stress in the study?
The three-factor model showed adequate overall fit, but latent correlations were extremely high, including 0.939 between depression and stress and 0.949 between anxiety and stress. This suggested substantial overlap among the dimensions in this particular student sample.
Can a PHQ-9 or DASS-21 score diagnose depression or another mental-health disorder?
No questionnaire score should be treated as a stand-alone clinical diagnosis. These instruments can support screening, symptom measurement and research, while diagnosis requires appropriate clinical assessment and consideration of the person’s broader circumstances.

References

  1. Asghar T, Hassan A, Sahar I, et al. Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students: a cross-sectional study in Pakistan. BMC Psychology. 2026. https://doi.org/10.1186/s40359-026-05332-5
  2. Kroenke K, Spitzer RL, Williams JBW. The PHQ-9: Validity of a Brief Depression Severity Measure. Journal of General Internal Medicine. 2001;16:606-613. https://pmc.ncbi.nlm.nih.gov/articles/PMC1495268/
  3. Lovibond SH, Lovibond PF. Manual for the Depression Anxiety Stress Scales. 2nd ed. Sydney: Psychology Foundation; 1995. https://www2.psy.unsw.edu.au/dass/
  4. Henry JD, Crawford JR. The short-form version of the Depression Anxiety Stress Scales (DASS-21): Construct validity and normative data in a large non-clinical sample. British Journal of Clinical Psychology. 2005;44:227-239. https://pubmed.ncbi.nlm.nih.gov/16004657/