Cronbach’s alpha can tell us whether items on a questionnaire tend to move together, but it cannot tell us whether a proposed factor structure is convincing, which individual symptoms are most informative, whether a scale measures equally well across groups, or how symptoms connect with one another. In our study of 602 university students in Lahore, Pakistan, we therefore evaluated the PHQ-9 and DASS-21 using a broader psychometric framework combining reliability analysis, confirmatory factor analysis (CFA), measurement invariance testing, item response theory (IRT), and symptom network analysis. The result was a much more detailed picture of how these widely used mental-health measures behaved in this student population. :contentReference[oaicite:0]{index=0}
By Taimoor Asghar
Why Cronbach’s alpha is only the beginning
Cronbach’s alpha remains one of the most familiar statistics in questionnaire research. At a basic level, it summarizes internal consistency: whether responses to items intended to measure a common construct are sufficiently interrelated. A scale with very low internal consistency raises an obvious concern because its items may not be functioning coherently.
But a satisfactory alpha does not prove that a questionnaire measures what researchers think it measures. It does not establish that depression, anxiety and stress are empirically distinguishable. It does not reveal whether one symptom is far more informative than another. It does not show where along the severity continuum a scale measures most precisely. And it does not tell us which symptoms might occupy central or bridging positions in a wider system of psychological distress.
That distinction mattered in our analysis. The PHQ-9 had a Cronbach’s alpha of 0.823, while the DASS-21 depression, anxiety and stress subscales had alpha values of 0.879, 0.854 and 0.858, respectively. Those coefficients supported good internal consistency. Yet the analyses that followed uncovered questions and patterns that alpha alone could never reveal. :contentReference[oaicite:1]{index=1}
What we tested in the Pakistani student sample
The study included 602 undergraduate students recruited from universities in Lahore, including 424 medical students and 178 students from non-medical disciplines. Both the nine-item Patient Health Questionnaire and the 21-item Depression Anxiety Stress Scales were administered. The study was cross-sectional, so its analyses describe measurement properties and associations observed at one point in time rather than demonstrating causal relationships. :contentReference[oaicite:2]{index=2}
The full study, Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students, was published in BMC Psychology in August 2026. The psychometric component was designed to answer several different questions rather than treating reliability as a single-number problem.
- Do the DASS-21 items behave according to the expected depression, anxiety and stress factor structure?
- Can scores be compared meaningfully between medical and non-medical students?
- Which PHQ-9 symptoms discriminate most strongly between students at different levels of latent depression?
- At which severity levels is the PHQ-9 most informative?
- How do PHQ-9 and DASS-21 symptoms connect when examined simultaneously as a network?
Each method addressed a different part of the measurement problem.
Confirmatory factor analysis: does the DASS-21 structure fit the data?
Confirmatory factor analysis begins with a theory about structure. For the DASS-21, the conventional model proposes three correlated latent dimensions: depression, anxiety and stress. Each of the 21 observed questionnaire items is assigned to one of those domains.
Instead of simply asking whether items within each subscale correlate, CFA asks whether the entire pattern of relationships is reasonably consistent with the hypothesized measurement model. We tested the three-factor DASS-21 structure using robust maximum-likelihood estimation and evaluated model performance using established fit indices including the Comparative Fit Index, Tucker-Lewis Index, root mean square error of approximation and standardized root mean square residual. :contentReference[oaicite:3]{index=3}
The three-factor model fitted reasonably well
The DASS-21 three-factor model produced a CFI of 0.954, TLI of 0.948, RMSEA of 0.052 with a 90% confidence interval from 0.046 to 0.058, and an SRMR of 0.031. Most standardized item loadings exceeded 0.60, although some were lower. These results supported an adequately fitting three-factor representation of the data. :contentReference[oaicite:4]{index=4}
If we had stopped with alpha, the conclusion would simply have been that all three DASS-21 subscales had good internal consistency. CFA added an important second layer: the individual items generally loaded on the intended depression, anxiety and stress factors, and the overall hypothesized model fitted the observed response structure reasonably well.
But the factors were extremely strongly correlated
CFA also revealed a less reassuring feature. The latent correlation between depression and stress was 0.939, while the correlation between anxiety and stress was 0.949. These are extremely high relationships between supposedly distinct latent constructs. :contentReference[oaicite:5]{index=5}
This does not mean that the DASS-21 is unusable or that the three-factor solution should automatically be discarded. It does mean that researchers should be cautious about interpreting depression, anxiety and stress scores as entirely independent psychological dimensions in this population. A model can fit adequately while its latent variables remain very strongly intertwined.
That finding illustrates exactly why reliability coefficients should not be treated as evidence of dimensional validity. Alpha told us that each set of items was internally coherent. CFA showed us both that the intended item structure was defensible and that the boundaries between the resulting latent domains were unusually narrow.
Measurement invariance: are group comparisons measuring the same thing?
Once researchers begin comparing questionnaire scores between groups, another question becomes critical: does the instrument operate similarly in those groups? A difference in average scores is easier to interpret when the scale has comparable measurement properties on both sides of the comparison.
We therefore tested DASS-21 measurement invariance between medical and non-medical students. Configural invariance asks whether the broad factor pattern is similar. Metric invariance tests whether factor loadings can be considered equivalent. Scalar invariance goes further by examining whether item intercept-related parameters are sufficiently comparable to support interpretation of group-level mean differences.
Configural, metric and scalar invariance were supported. Changes in model fit were minimal: moving from the configural to metric model produced a change in CFI of −0.0006 and change in RMSEA of −0.0010; the transition from metric to scalar invariance produced a change in CFI of −0.0005 and RMSEA of −0.0010. These changes remained within the prespecified thresholds used in the analysis. :contentReference[oaicite:6]{index=6}
This finding was especially useful because the study compared symptom levels across academic disciplines. It provided evidence that observed differences were not simply an obvious consequence of the DASS-21 having radically different measurement structures in medical versus non-medical students.
Item response theory: not every PHQ-9 symptom contributes equally
Classical reliability analysis treats a questionnaire largely at the total-score level. Item response theory turns attention toward individual questions and asks how each item behaves across an underlying latent trait.
For the PHQ-9, we fitted a graded response model appropriate for ordered response categories. Each item received a discrimination parameter and a series of thresholds. Discrimination indicates how sharply an item differentiates between respondents with different levels of the underlying depression trait. Thresholds describe where on that latent continuum respondents become increasingly likely to endorse higher response categories. :contentReference[oaicite:7]{index=7}
Self-worth and concentration were especially informative
The PHQ-9 items were not psychometrically interchangeable. Self-worth had a discrimination parameter of approximately 1.99, concentration approximately 1.82, appetite approximately 1.78 and feeling down approximately 1.90. By contrast, the loss-of-interest or anhedonia item had the lowest discrimination parameter, approximately 0.47. :contentReference[oaicite:8]{index=8}
This does not mean that anhedonia is clinically unimportant. Clinical importance and psychometric discrimination are different concepts. Anhedonia is a core depressive symptom, but in this particular sample its response pattern did less to distinguish positions along the estimated latent depression continuum than several cognitive and somatic symptoms.
That distinction is valuable for anyone interpreting questionnaire data. A nine-item scale generates one total score, but those nine items may contribute very differently to statistical measurement precision.
Where did the PHQ-9 provide the most information?
The test information curve showed that the PHQ-9 was most precise from approximately theta 0 to +2 on the latent depression continuum. In the study’s interpretation, this corresponded broadly to the mild-to-moderate range of depressive symptom severity. Standard error was lowest across the same region, meaning measurement precision was greatest there. :contentReference[oaicite:9]{index=9}
This is one of the advantages of IRT over a single reliability coefficient. Alpha gives one summary value for a sample. An information curve shows that measurement precision can vary depending on where a respondent lies on the underlying trait.
IRT also identified imperfections
Advanced psychometric analysis should not be used only to generate reassuring statistics. Item-level fit testing identified significant misfit for the self-worth and self-harm items in the graded response model. The remaining seven PHQ-9 items showed adequate item-level fit under the reported testing procedure. Some residual associations also suggested minor local dependence between particular items. :contentReference[oaicite:10]{index=10}
The self-worth and self-harm items were retained because they are integral parts of the standard PHQ-9 and because removing items simply to improve statistical fit can undermine comparability and content validity. The more appropriate conclusion is that an otherwise useful model still had item-level limitations that deserve consideration and replication.
Network analysis: what happens when symptoms are treated as connected nodes?
CFA and IRT generally operate within latent-variable frameworks: observed item responses are understood partly through underlying constructs such as depression. Network analysis asks a different question. Instead of assuming that correlations between symptoms arise solely because they reflect a common latent disorder, symptoms can be represented as interconnected nodes within a statistical network.
Our analysis combined all nine PHQ-9 items and all 21 DASS-21 items, producing a 30-node symptom network. A Gaussian graphical model was estimated using graphical LASSO regularization with extended Bayesian information criterion model selection. In this type of network, an edge represents an estimated conditional association between two symptoms after accounting for the remaining nodes. :contentReference[oaicite:11]{index=11}
Before estimating the final network, a redundancy check was conducted to look for highly overlapping symptom pairs. No pair exceeded the study’s specified redundancy threshold, so all 30 items were retained. :contentReference[oaicite:12]{index=12}
Central symptoms highlighted cognitive and self-worth features
The resulting network showed recognizable clustering within depression, anxiety and stress domains. Strength centrality was particularly high for DASS self-worth, feeling down, having nothing to look forward to and PHQ-9 concentration. Betweenness centrality highlighted PHQ-9 self-worth, a DASS meaninglessness item and PHQ-9 sleep disturbance as potential bridges between parts of the network. :contentReference[oaicite:13]{index=13}
The convergence is noteworthy. IRT identified self-worth, concentration and feeling down among the more discriminating PHQ-9 symptoms, while network analysis independently placed self-worth, concentration and related depressive-cognitive symptoms in structurally prominent positions.
However, centrality should not be confused with proof of a causal treatment target. A cross-sectional network cannot demonstrate that changing one central symptom will cause improvement in all connected symptoms. Centrality describes the estimated structure of the observed data. Testing whether a symptom is truly an intervention target would require longitudinal or experimental evidence.
How stable was the network?
Network diagrams can look compelling even when their estimated structures are unstable, so robustness assessment matters. The study used bootstrapping and case-dropping procedures. Correlation-stability coefficients were approximately 0.44 for strength, 0.31 for betweenness and 0.38 for closeness. These exceeded the analysis threshold of 0.25, although they also indicate that centrality estimates should be interpreted with appropriate caution rather than as exact rankings. :contentReference[oaicite:14]{index=14}
What each method contributed
Using several psychometric approaches on the same dataset allowed each method to answer a question that the others could not.
| Method | Main question | What it added |
|---|---|---|
| Cronbach’s alpha | Do the scale items show internal consistency? | PHQ-9 and DASS-21 subscales showed good internal consistency. |
| CFA | Does the hypothesized DASS-21 factor structure fit? | The three-factor model fitted adequately, while factor correlations revealed substantial overlap. |
| Measurement invariance | Does the DASS-21 operate comparably across disciplines? | Configural, metric and scalar invariance were supported for medical and non-medical groups. |
| IRT | Which PHQ-9 items discriminate best, and where is measurement precise? | Self-worth, concentration, feeling down and appetite were highly informative; precision peaked around mild-to-moderate latent severity. |
| Network analysis | How are symptoms conditionally connected? | Self-worth, concentration and related depressive symptoms emerged as prominent nodes or bridges. |
The important point is not that one technique is superior. These approaches operate at different levels of the measurement problem. Internal consistency, structural validity, item functioning and symptom connectivity are related questions, but they are not the same question.
Why the very high DASS-21 factor correlations matter
One of the most interesting findings was the contrast between good conventional psychometric indicators and extremely high latent correlations among DASS-21 domains. The instrument could show good alpha values, satisfactory factor loadings and acceptable global CFA fit while depression, anxiety and stress remained very strongly related at the latent level.
For applied research, that means a statistically significant difference on one DASS-21 subscale should not automatically be interpreted as evidence for an isolated psychological process. The constructs may remain useful descriptively, but their substantial shared variance deserves explicit recognition, particularly in populations experiencing broad psychological distress.
It also demonstrates why psychometric validation should not become a checklist in which researchers report alpha above 0.70 and declare an instrument validated. Measurement evidence is multidimensional. An instrument can perform well according to one criterion and simultaneously raise substantive questions according to another.
Why the IRT and network findings should not be overinterpreted
Findings involving individual symptoms can be tempting to convert immediately into clinical conclusions. That would go beyond what this study can establish. The sample consisted of university students rather than a clinical population, and the design was cross-sectional. IRT parameters are sample-dependent estimates, and network centrality can vary across populations, analytical decisions and sampling conditions.
The self-worth and concentration findings are therefore better viewed as hypotheses worth replicating rather than a new diagnostic hierarchy for depression. Likewise, a central network node is not automatically the symptom that a clinician should treat first.
The network stability results were acceptable according to the study’s stated criterion but were not so high that minor ordering differences between central symptoms should be treated as definitive. Replication in independent Pakistani samples, clinical populations and longitudinal datasets would strengthen interpretation.
What this means for future questionnaire research
For researchers using the PHQ-9, DASS-21 or similar scales, the broader lesson is methodological. Reliability should be treated as one component of measurement evaluation rather than the final destination.
A stronger validation strategy might begin with internal consistency, examine factor structure, test whether group comparisons are psychometrically defensible, investigate item-level behavior and then explore relationships among symptoms when the research question justifies it. Not every study needs every technique, but the choice of method should follow the measurement question.
CFA becomes important when a scale claims a specific dimensional structure. Measurement invariance matters when comparing groups. IRT is particularly useful when researchers care about item discrimination or measurement precision across severity levels. Network analysis becomes relevant when the aim is to investigate conditional symptom relationships rather than only total scores or latent dimensions.
The broader conclusion: measurement is more than reliability
Our analysis demonstrated why evaluating the PHQ-9 and DASS-21 required more than reporting Cronbach’s alpha. Alpha showed good internal consistency. CFA broadly supported the DASS-21 three-factor structure but simultaneously revealed striking overlap between its latent dimensions. Measurement invariance supported comparisons between medical and non-medical students. IRT showed that PHQ-9 symptoms contributed unequal amounts of information and that the scale was most precise across a particular part of the depression continuum. Network analysis then provided another perspective, identifying self-worth, concentration and related symptoms as structurally prominent within a combined PHQ-9 and DASS-21 network. :contentReference[oaicite:15]{index=15}
Taken together, these analyses did not produce a simple verdict that either questionnaire is universally valid or invalid. They produced something more useful: a detailed description of where the instruments performed well, where their constructs overlapped, which items carried more information, and which findings require cautious interpretation and future replication.
That is the value of moving beyond Cronbach’s alpha. Good psychometric research asks not merely whether a questionnaire is internally consistent, but what exactly it measures, how its items behave, whether comparisons are defensible, where its precision is strongest and what its symptom-level structure can—and cannot—tell us.
Frequently asked questions
Is Cronbach’s alpha enough to validate the PHQ-9 or DASS-21?
No. Cronbach’s alpha provides evidence about internal consistency but does not establish factor structure, measurement invariance, item discrimination, measurement precision or symptom-network structure. Validation should be matched to the claims researchers intend to make from the scale.
What did CFA show about the DASS-21?
The proposed three-factor depression, anxiety and stress model showed adequate-to-good fit according to the reported indices. However, depression, anxiety and stress factors were extremely highly correlated, particularly depression with stress and anxiety with stress. :contentReference[oaicite:16]{index=16}
Which PHQ-9 symptoms were most informative in the IRT analysis?
Self-worth, feeling down, concentration and appetite showed among the highest discrimination and information in this sample, while the loss-of-interest item had the lowest discrimination estimate. These findings are sample-specific and should not be interpreted as a ranking of clinical importance. :contentReference[oaicite:17]{index=17}
What did network analysis add?
Network analysis examined conditional relationships among all 30 PHQ-9 and DASS-21 items. It highlighted self-worth, feeling down, concentration and related symptoms in prominent central or bridging positions, providing an item-level perspective that total scores and CFA alone do not provide. :contentReference[oaicite:18]{index=18}
Can central symptoms be considered treatment targets?
Not from this study alone. Cross-sectional centrality identifies structural prominence in the estimated network, not causal influence. Longitudinal and intervention studies would be needed before concluding that targeting a particular symptom produces downstream clinical improvement.
Medical disclaimer
This article discusses psychometric research and mental-health screening instruments for educational purposes. The PHQ-9 and DASS-21 are questionnaires and do not by themselves establish an individual psychiatric diagnosis. Personal symptoms, self-harm thoughts or concerns about mental health should be assessed by an appropriately qualified healthcare professional. Urgent or immediate safety concerns require prompt professional assistance.
Key takeaways
- Cronbach’s alpha showed good internal consistency for the PHQ-9 and all three DASS-21 subscales, but reliability alone could not establish structural validity.
- CFA supported the DASS-21 three-factor model while revealing extremely high correlations among depression, anxiety and stress latent factors.
- Measurement invariance testing supported meaningful DASS-21 comparisons between medical and non-medical students in this sample.
- IRT showed that PHQ-9 items differed substantially in discrimination and information, with self-worth and concentration among the strongest items.
- Network analysis highlighted self-worth, concentration and related depressive symptoms as structurally prominent, but cross-sectional centrality should not be interpreted causally.
Frequently asked questions
Is Cronbach’s alpha enough to validate the PHQ-9 or DASS-21?
What did CFA show about the DASS-21 in the Pakistani student sample?
Which PHQ-9 items were most informative in the IRT analysis?
What does network analysis reveal that a total score cannot?
Does a central symptom in a network automatically become a treatment target?
References
- Asghar T, Hassan A, Sahar I, Tahir M, Shahid B, Komal K. Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students: a cross-sectional study in Pakistan. BMC Psychology. 2026. https://doi.org/10.1186/s40359-026-05332-5 :contentReference[oaicite:19]{index=19}