Skip to content
TAtaimoorasghar.com

August 29, 2026 · 15 min read

From PHQ-9 Scores to Symptom Networks: A Deeper Look at Student Mental Health

A 602-student Pakistani study shows how PHQ-9 scores, item response theory and symptom networks reveal detail hidden by total depression scores.

By Taimoor Asghar

A PHQ-9 total score can tell us how much depressive symptom burden a student reports, but it cannot show which symptoms are carrying the most information, which symptoms sit at the center of a broader symptom pattern, or how depression overlaps with anxiety and stress. A recent study of 602 university students in Lahore, Pakistan, combined conventional scale scores with confirmatory factor analysis, item response theory, and symptom network analysis. The result is a more detailed picture: the overall score still matters, but self-worth, concentration, low mood, and related symptoms may provide additional information about how distress is organized within this student population.

Why move beyond the PHQ-9 total score?

The Patient Health Questionnaire-9, or PHQ-9, is one of the most widely used brief measures of depressive symptoms. Its nine items correspond to core depressive symptoms and are scored from 0 to 3, producing a total score that summarizes symptom severity. That simplicity is a major strength. A total score is easy to calculate, compare across groups, and use in screening or research.

Yet every summary score compresses information. Two students can receive the same PHQ-9 total while endorsing very different symptoms. One might report poor sleep, fatigue, and concentration problems; another might report low mood, reduced interest, and feelings of worthlessness. Their totals may be identical, but their lived experiences and the pattern of symptoms behind those totals are not.

This distinction matters in student mental health research because university life involves multiple overlapping pressures: academic workload, sleep disruption, financial concerns, social transitions, uncertainty about the future, and pre-existing mental-health vulnerabilities. A single score can quantify burden, but symptom-level analyses may reveal which parts of that burden are especially informative in a given population.

What the Pakistani student study examined

The source study, published in BMC Psychology in August 2026, included 602 undergraduate students from universities in Lahore: 424 medical students and 178 non-medical students. Participants completed the PHQ-9 and the 21-item Depression Anxiety Stress Scales, commonly known as the DASS-21. The investigators did not stop at group averages. They used multivariable regression, confirmatory factor analysis, item response theory, symptom network analysis, and measurement-invariance testing to examine how the instruments behaved and how individual symptoms contributed to the overall picture.

The full source article can be read in the BMC Psychology study on PHQ-9 and DASS-21 symptom profiles among Pakistani university students.

This layered approach is useful because each method asks a different question. Regression asks which observed characteristics are associated with higher scores. Confirmatory factor analysis asks whether the questionnaire structure fits the data. Item response theory asks how strongly individual items discriminate across levels of the underlying trait. Network analysis asks how symptoms are connected to one another after accounting for the rest of the network.

What the conventional scores showed

At the score level, non-medical students had a higher median PHQ-9 score than medical students: 10 compared with 9. The adjusted analysis also found higher PHQ-9, DASS Depression, and DASS Anxiety scores among non-medical students, while the adjusted difference in DASS Stress did not reach conventional statistical significance.

Those group differences should be interpreted carefully. The study also reported low model R-squared values, ranging from 0.029 to 0.041. In practical terms, the variables included in the regression models explained only a small proportion of the variation in symptom scores. Academic discipline therefore cannot be treated as a simple explanation for student mental health. The finding is better understood as evidence that, in this sample, non-medical students reported somewhat higher depressive and anxiety symptoms after adjustment, while much of the individual variation remained unexplained by the measured predictors.

Female gender was associated with higher scores across the PHQ-9 and DASS-21 domains. Each additional hour of sleep was associated with a lower PHQ-9 score, and prior depression treatment showed the largest positive association with PHQ-9 scores among the reported predictors. Because the design was cross-sectional, these associations do not establish direction of causation. For example, more sleep could relate to lower depressive symptoms, depressive symptoms could disrupt sleep, or both could be influenced by other factors.

Reliability is useful, but it is not the whole psychometric story

Researchers often begin evaluating a questionnaire by reporting internal consistency. In this study, Cronbach’s alpha ranged from 0.82 to 0.88 across the scales, values generally considered adequate for group-level research. But a high alpha does not prove that a scale is unidimensional, does not show that every item is equally informative, and does not establish that theoretically distinct subscales are empirically well separated.

That is why the study added confirmatory factor analysis and item response theory. These methods can reveal features that a single reliability coefficient cannot.

DASS-21 factor structure: acceptable fit with substantial overlap

The DASS-21 is designed to assess three related constructs: depression, anxiety, and stress. In the Pakistani student sample, the confirmatory factor analysis showed adequate model fit, with a comparative fit index of 0.954, Tucker-Lewis index of 0.948, and root mean square error of approximation of 0.052.

At the same time, the latent correlations between some DASS-21 factors were extremely high. The reported depression-stress correlation was 0.939, while anxiety-stress was 0.949. Those values suggest that although a three-factor model could fit the observed responses reasonably well, the constructs were very strongly intertwined in this sample.

This does not mean that depression, anxiety, and stress are literally the same phenomenon. It does mean that researchers should be cautious about assuming that three labeled subscale scores represent sharply separated psychological dimensions in every population. The distinction may be theoretically meaningful while still being difficult to separate empirically when symptoms co-occur strongly.

What item response theory adds to a PHQ-9 score

Classical scoring gives every PHQ-9 item a straightforward contribution to the total: higher responses add more points. Item response theory, or IRT, asks a more nuanced question. How well does each item distinguish between people at different levels of the underlying depressive-symptom trait?

In the study’s graded-response IRT analysis, items related to self-worth, concentration, feeling down, and appetite showed the highest discrimination. The interest or anhedonia item showed the lowest discrimination, with a reported discrimination parameter of 0.468.

“Discrimination” has a technical meaning here. It does not mean that one symptom is clinically important and another is unimportant. A highly discriminating item is one whose response probabilities change more sharply as the measured latent trait changes. An item with lower discrimination can still be clinically meaningful; it simply contributes differently to measurement in that dataset.

This is a central reason to look beyond summed scores. If some symptoms carry more psychometric information than others, a total score treats unequal measurement contributions as though they were interchangeable points. IRT helps researchers see where a scale is especially sensitive and where it may provide less differentiation.

From symptom lists to symptom networks

Network analysis takes another conceptual step. Traditional psychometric models often assume that an underlying condition or latent trait gives rise to observed symptoms. A network perspective instead examines symptoms as nodes connected by statistical relationships. For depression, the nodes might include low mood, sleep disturbance, fatigue, concentration difficulty, appetite change, and negative self-evaluation.

The network approach does not automatically prove that one symptom causes another. In cross-sectional data, an edge usually represents a conditional association rather than a demonstrated causal pathway. Still, network models can help identify symptoms that occupy structurally important positions within the observed pattern.

In the Pakistani student study, self-worth, concentration, and downheartedness emerged as the most central nodes. This broadly complements the IRT results, where self-worth and concentration also showed strong discrimination. When two different analytical frameworks both highlight similar symptoms, the convergence is scientifically interesting, although it should not be overstated as proof that these symptoms are universal treatment targets.

What does “central symptom” actually mean?

A central symptom is one that has relatively strong or numerous connections with other symptoms according to a particular network metric. If concentration difficulty is central, for example, it may be statistically linked with several other symptoms after the model accounts for the remaining nodes.

It is tempting to translate centrality directly into a clinical instruction: “Treat the most central symptom first.” That conclusion is not justified by a single cross-sectional network. Centrality can depend on the sample, the questionnaire, the estimation method, the chosen centrality statistic, and the stability of the network. A symptom can also appear central because it is measured in a way that overlaps with multiple neighboring experiences.

The study explicitly included stability testing, which is important because centrality rankings can be fragile. Reported stability coefficients ranged from 0.31 to 0.44. These values support a cautious reading: the network provides useful exploratory information, but exact rankings should not be treated as immutable biological or clinical facts.

Why self-worth and concentration deserve attention

The prominence of self-worth and concentration is noteworthy for student populations. Academic environments repeatedly demand sustained attention, memory, planning, comparison with peers, and evaluation through examinations or grades. Difficulty concentrating can therefore have consequences that extend beyond the symptom itself. Poor concentration may interfere with studying, missed academic goals may worsen distress, and the resulting feedback loop could amplify other difficulties.

Negative self-worth may also be especially consequential in environments where achievement is highly visible. Feelings of failure or low personal value can coexist with low mood, withdrawal, and reduced motivation. However, the study does not establish that academic pressure caused these symptoms, nor does it show that changing self-worth or concentration alone would reduce the entire symptom network.

A better interpretation is that these symptoms were psychometrically and structurally prominent in this specific dataset. For campus screening and follow-up conversations, that finding encourages clinicians and researchers not to focus only on the total score. Asking what is driving the score may reveal more actionable context.

What the network does not tell us

Symptom networks can look visually intuitive, which can make them easy to overinterpret. Several limitations are worth keeping in view.

  • Cross-sectional networks do not establish temporal order. They cannot show whether low self-worth came before concentration problems, whether the reverse occurred, or whether both changed together because of another factor.
  • Centrality is sample-dependent. A network estimated in Lahore university students may not reproduce exactly in working adults, adolescents, clinical patients, or students in other countries.
  • Questionnaire items define the available nodes. A network cannot reveal symptoms or contextual factors that were never measured.
  • Group networks describe aggregate patterns. The structure estimated across hundreds of students should not be assumed to represent the symptom dynamics of a particular individual.
  • Statistical importance is not identical to clinical urgency. A less central symptom can still require immediate attention, particularly when it involves safety, severe impairment, or suicidal thoughts.

Why combining CFA, IRT, and networks is stronger than relying on one method

The most useful feature of the study is not any single statistic. It is the combination of methods. Confirmatory factor analysis evaluates the broad measurement structure. Reliability estimates describe internal consistency. IRT examines item-level measurement behavior. Network analysis examines symptom-to-symptom relationships. Regression places scale scores in relation to participant characteristics.

Together, these methods answer different layers of the same question: what do the questionnaires measure, how well do they measure it, which items contribute the most information, how do symptoms cluster or connect, and how do total scores differ across groups?

This multidimensional approach is especially valuable for instruments such as the PHQ-9 and DASS-21 because both are frequently reduced to totals. Totals are practical, but they are not the only information contained in the responses. A student with a PHQ-9 score of 10 is not simply a “10.” The pattern that produces that score may matter for research interpretation and for a clinician’s conversation with that student.

Measurement invariance strengthens group comparisons

Whenever researchers compare questionnaire scores between groups, they should ask whether the instrument functions similarly across those groups. If an item has a different meaning or measurement relationship in medical and non-medical students, an observed score difference could partly reflect measurement artifacts rather than a true difference in symptom burden.

The study reported support for measurement invariance across academic discipline. That finding strengthens the interpretation of the medical versus non-medical comparisons because it suggests that the scales were measuring comparable constructs across those groups. It does not eliminate every source of bias, but it addresses an important psychometric prerequisite that is often ignored in simple group comparisons.

What this means for university mental-health screening

The findings support a balanced view of screening. Brief questionnaires remain useful because they are scalable, standardized, and easy to administer. A PHQ-9 or DASS-21 score can help identify elevated symptom burden and can support population-level monitoring. But screening should not end with a number.

For student-support services, symptom-level review can add context. A student whose score is driven largely by sleep disruption and fatigue may need a different conversation from someone whose score reflects persistent low mood, impaired concentration, and severe negative self-evaluation. The questionnaire is a starting point for assessment, not a substitute for clinical judgment.

The study’s results also argue against focusing campus mental-health attention only on medical students. In this sample, non-medical students reported higher adjusted depression and anxiety scores. The broader implication is that universities should avoid assuming risk solely from discipline labels. Mental-health systems should be accessible across faculties and should recognize that vulnerability can arise from many overlapping personal and environmental factors.

How researchers should interpret these findings

Researchers planning similar studies can draw several methodological lessons. First, report more than Cronbach’s alpha when making claims about validity or structure. Second, test whether group comparisons are supported by measurement invariance. Third, consider item-level methods when a total score may conceal heterogeneous symptom patterns. Fourth, use network centrality cautiously and report stability rather than presenting visually prominent nodes as definitive targets.

Replication is also essential. The most central or discriminating symptoms in one sample may shift in another. Future work could test whether self-worth and concentration remain prominent in other Pakistani universities, in longitudinal cohorts, or in clinically assessed students. Longitudinal or intensive repeated-measures designs would be particularly useful because they can investigate whether changes in one symptom precede changes in another over time.

A deeper reading of student mental health

The key lesson is not that PHQ-9 totals should be abandoned. They remain compact and interpretable summaries. The lesson is that totals answer only one level of the problem.

In this 602-student study, conventional analyses identified differences in average symptom burden, while psychometric and network analyses revealed additional structure. Self-worth and concentration repeatedly emerged as informative features, DASS-21 domains showed very high latent overlap, and the network results suggested that some symptoms occupied more central positions than others. At the same time, low regression R-squared values and moderate network-stability estimates warn against overly simple conclusions.

Student mental health is not captured by one score, one faculty label, or one central node. A stronger approach combines severity scores with symptom patterns, measurement quality, context, and clinical assessment. For research, this produces a more precise understanding of what questionnaires are actually telling us. For practice, it reinforces a basic principle: the number can flag distress, but the pattern behind the number is where much of the meaningful detail begins.

Frequently asked questions

What is the PHQ-9 used for?

The PHQ-9 is a nine-item self-report questionnaire used to assess the presence and severity of depressive symptoms. It is commonly used for screening and symptom monitoring, but a score by itself does not establish a psychiatric diagnosis.

What is symptom network analysis?

Symptom network analysis models symptoms as interconnected nodes rather than reducing all responses to a single total. It can identify which symptoms have stronger conditional relationships with others, but cross-sectional networks should not be interpreted as proof of causal pathways.

Which PHQ-9 symptoms were most informative in this study?

In the item response analysis, self-worth, concentration, feeling down, and appetite showed the highest discrimination. In the symptom network, self-worth and concentration were also among the most central nodes, alongside downheartedness.

Did medical students have worse mental-health scores than non-medical students?

No. In this Lahore sample, non-medical students had a higher median PHQ-9 score and higher adjusted PHQ-9, DASS Depression, and DASS Anxiety scores. However, the statistical models explained only a small proportion of overall variation, so academic discipline should not be treated as the main determinant of an individual student’s mental health.

Can a symptom network tell clinicians what to treat first?

Not by itself. Central symptoms can generate useful hypotheses, but cross-sectional network centrality does not prove that targeting a particular symptom will cause improvement in the rest of the network. Clinical decisions require individual assessment, severity, safety, functional impact, patient preferences, and other relevant information.

Medical disclaimer: This article is for educational and research communication purposes only. The PHQ-9 and DASS-21 are screening and symptom-measurement tools and are not substitutes for an individualized clinical evaluation. Anyone experiencing severe distress, suicidal thoughts, marked functional decline, or other urgent mental-health concerns should seek prompt assessment from an appropriate healthcare professional or emergency service.

Key takeaways

  • A PHQ-9 total score summarizes depressive symptom burden but can hide substantial differences in the symptom patterns that produce the same score.
  • In the 602-student Lahore study, self-worth and concentration stood out in both item response and symptom network analyses.
  • The DASS-21 showed adequate overall factor-model fit, but very high correlations among some latent domains suggested substantial overlap between depression, anxiety and stress.
  • Non-medical students reported higher adjusted depression and anxiety scores than medical students in this sample, but academic discipline explained only a small share of overall score variation.
  • Symptom network centrality should be interpreted cautiously because cross-sectional associations do not establish causality and exact rankings may be sample-dependent.

Frequently asked questions

What is the PHQ-9 used for?
The PHQ-9 is a nine-item self-report questionnaire used to assess depressive symptom presence and severity. It is widely used for screening and symptom monitoring, but a score alone does not establish a psychiatric diagnosis.
What is symptom network analysis?
Symptom network analysis models symptoms as interconnected nodes and estimates relationships among them. It can reveal structurally prominent symptoms, but cross-sectional networks do not prove causal pathways.
Which PHQ-9 symptoms were most informative in this study?
The study reported high item discrimination for self-worth, concentration, feeling down and appetite. Self-worth and concentration were also among the most central symptoms in the network analysis.
Did medical students have worse mental-health scores than non-medical students?
No. In this Lahore sample, non-medical students had a higher median PHQ-9 score and higher adjusted PHQ-9, DASS Depression and DASS Anxiety scores, although the models explained only a small proportion of total variation.
Can a symptom network tell clinicians what to treat first?
Not on its own. Network centrality can generate hypotheses, but clinical decisions require individual assessment, safety considerations, functional impact, patient preferences and other relevant evidence.

References

  1. Asghar T, Hassan A, Sahar I, et al. Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students: a cross-sectional study in Pakistan. BMC Psychology. 2026. https://doi.org/10.1186/s40359-026-05332-5
  2. Kroenke K, Spitzer RL, Williams JBW. The PHQ-9: validity of a brief depression severity measure. Journal of General Internal Medicine. 2001;16(9):606-613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x
  3. Borsboom D, Cramer AOJ. Network Analysis: An Integrative Approach to the Structure of Psychopathology. Annual Review of Clinical Psychology. 2013;9:91-121. https://doi.org/10.1146/annurev-clinpsy-050212-185608
  4. Lovibond PF, Lovibond SH. The structure of negative emotional states: comparison of the Depression Anxiety Stress Scales (DASS) with the Beck Depression and Anxiety Inventories. Behaviour Research and Therapy. 1995;33(3):335-343. https://doi.org/10.1016/0005-7967(94)00075-U