Skip to content
TAtaimoorasghar.com

August 16, 2026 · 14 min read

Which Depression Symptoms Matter Most? What Item Response and Network Analysis Revealed

A Pakistani university study used item response and network analysis to reveal which PHQ-9 depression symptoms carried the most information.

Depression questionnaires are usually summarized as total scores, but individual symptoms do not necessarily contribute equally to what those scores tell us. In a recent study of 602 university students in Lahore, Pakistan, item response theory and symptom network analysis both highlighted self-worth, concentration, and depressed mood as particularly informative features of the PHQ-9 symptom profile. Appetite also showed strong discrimination in the item response model, while loss of interest showed comparatively low discrimination. These findings do not mean that some symptoms can simply be ignored. Instead, they show how looking beneath the total score can reveal important differences in how individual symptoms behave within a specific population.

By Taimoor Asghar

Why Look Beyond a Total Depression Score?

The Patient Health Questionnaire-9, or PHQ-9, contains nine items corresponding to major depressive symptoms. A respondent rates how frequently each symptom has occurred during the previous two weeks, and those responses are combined into a total score.

Total scores are useful. They provide a compact measure of overall symptom burden and can support screening, monitoring, epidemiological research, and structured clinical assessment. Yet summing nine responses into one number inevitably removes some information.

Two students could have the same PHQ-9 total while reporting very different experiences. One might primarily report depressed mood, poor concentration, and feelings of worthlessness. Another could arrive at a similar score through sleep disturbance, fatigue, appetite changes, and reduced interest. Numerically, the totals may be identical. Psychologically and clinically, however, the symptom patterns are not necessarily interchangeable.

This is one reason researchers increasingly complement conventional reliability and factor-analysis methods with techniques that examine individual symptoms. Item response theory asks how strongly each item relates to the underlying construct being measured. Network analysis approaches the problem differently by examining how symptoms are statistically connected to one another.

Our recently published study of PHQ-9 and DASS-21 symptom profiles among Pakistani university students applied both approaches alongside conventional psychometric analyses. The combination offered a more detailed picture of depression measurement than a total score alone could provide.

What Was Studied?

The cross-sectional study included 602 undergraduate students from universities in Lahore, Pakistan. Of these, 424 were medical students and 178 were non-medical students. Participants completed the PHQ-9 as well as the 21-item Depression Anxiety Stress Scales, or DASS-21.

The analysis was broader than simply comparing average questionnaire scores. It included descriptive statistics, group comparisons, multivariable regression, reliability assessment, confirmatory factor analysis, measurement invariance testing, item response theory using a graded response model, and symptom network analysis with stability assessment.

That broader psychometric framework matters because different methods answer different questions. Reliability can indicate whether items behave consistently as a scale. Factor analysis examines whether observed responses conform to an expected latent structure. Item response theory examines individual item performance across an underlying symptom continuum. Network analysis examines relationships among symptoms themselves.

What Does Item Response Theory Tell Us?

Item response theory, commonly abbreviated IRT, shifts attention from the questionnaire as a whole to the performance of individual items. Rather than assuming that every symptom contributes identically to measurement, IRT estimates properties of each item.

The study used a graded response model, an IRT model designed for ordered response categories such as the PHQ-9’s frequency options.

Understanding item discrimination

One especially useful IRT parameter is discrimination. In simplified terms, discrimination describes how sharply an item differentiates between people at different positions on the underlying construct being measured.

A highly discriminating item tends to change its response probabilities relatively strongly as the underlying depression level changes. Such an item can therefore provide substantial information about differences along that latent continuum in the population being studied.

A lower discrimination parameter does not automatically mean that an item is clinically unimportant or that it should be removed. Clinical importance, diagnostic relevance, content validity, and statistical discrimination are distinct concepts. An item can describe a meaningful feature of depression even if its responses do not strongly distinguish levels of the latent construct in one particular dataset.

Which PHQ-9 Symptoms Were Most Discriminating?

The item response analysis identified several symptoms with comparatively high discrimination. These included:

  • negative feelings about oneself or impaired self-worth;
  • difficulty concentrating;
  • feeling down, depressed, or hopeless;
  • changes in appetite.

Together, these results suggest that cognitive, affective, and selected somatic symptoms were particularly informative about differences in the underlying depression construct within this Pakistani university sample.

The finding regarding self-worth is notable because it goes beyond simply asking whether someone feels sad. Feelings of failure, inadequacy, or having let oneself or one’s family down can represent an important cognitive dimension of depressive experience. In a university environment, such thoughts may intersect with academic expectations, social comparison, family expectations, uncertainty about future careers, and perceived personal performance. The present cross-sectional data cannot establish why the item performed strongly, but its psychometric prominence is worth examining in future work.

Concentration was also highly discriminating. Concentration difficulties can be especially consequential for students because sustained attention is central to lectures, reading, examination preparation, clinical learning, assignments, and decision-making. Again, the statistical result should not be interpreted as evidence that concentration problems uniquely cause depression. It indicates that responses to this item differentiated levels of the measured depression construct relatively well in this sample.

Depressed mood itself was another highly discriminating symptom, which is intuitively consistent with its central role in the clinical description of depression.

Appetite change also showed high discrimination in the IRT analysis. This is a useful reminder that symptom information was not confined to cognitive items. Depression is heterogeneous, and changes involving sleep, energy, appetite, psychomotor functioning, cognition, mood, and self-evaluation can all form part of a person’s presentation.

What About Loss of Interest?

One of the more interesting results concerned the PHQ-9 item assessing diminished interest or pleasure, often described as anhedonia. In this dataset, the interest item had the lowest reported discrimination parameter, with an estimated value of 0.468.

That finding requires careful interpretation.

It does not mean that anhedonia is clinically irrelevant. It does not demonstrate that the PHQ-9 should be rewritten based on this study. And it certainly does not mean that clinicians should disregard loss of interest when evaluating a patient.

Instead, the result means that this particular item was less effective than several other PHQ-9 items at differentiating positions on the modeled latent depression continuum in this sample.

Item performance can depend on population characteristics, language and interpretation, symptom prevalence, response distributions, cultural context, study design, and statistical model assumptions. Replication in independent Pakistani student samples and other populations would therefore be necessary before making broader measurement conclusions.

What Does Symptom Network Analysis Add?

Item response theory still conceptualizes responses largely in relation to an underlying latent construct. Network analysis offers a different perspective.

In a symptom network, individual symptoms are represented as nodes, while statistical relationships between them are represented as connections or edges. Instead of treating symptoms only as interchangeable indicators of one hidden disorder, network methods allow researchers to investigate the structure of relationships among the symptoms themselves.

This can be valuable because depressive symptoms commonly occur together in patterns. Sleep difficulties may relate to fatigue. Fatigue may coexist with concentration problems. Negative self-evaluation may occur alongside depressed mood. Network analysis attempts to characterize such conditional relationships statistically.

However, a cross-sectional symptom network should not be mistaken for a causal map. If two symptoms are connected, the analysis does not prove that one produces the other. Direction of causation cannot generally be established from a single cross-sectional assessment.

Which Symptoms Were Most Central in the Network?

The network analysis identified self-worth, concentration, and downheartedness as the most central nodes.

This is particularly interesting because two of those features—self-worth and concentration—also emerged among the strongest discriminating items in the IRT analysis. Depressed mood or downheartedness was likewise prominent across the analyses.

The convergence does not make the two techniques equivalent. IRT discrimination and network centrality measure different properties. A symptom can discriminate well without being highly central in a network, and a central network node is not automatically the best screening question. Nevertheless, when different analytical approaches independently draw attention to similar symptoms, the pattern becomes scientifically interesting.

Why self-worth may deserve greater research attention

Self-worth appears particularly noteworthy because it emerged as both highly discriminating and highly central. The PHQ-9 item encompasses feeling bad about oneself, perceiving oneself as a failure, or feeling that one has disappointed oneself or one’s family.

For student mental-health research, this dimension deserves attention beyond its contribution to the PHQ-9 total score. Academic performance, perceived expectations, financial concerns, competition, family pressures, and uncertainty about employment could plausibly interact with self-evaluation, although the present study was not designed to establish those pathways.

Future longitudinal research could examine whether changes in negative self-worth precede, follow, or develop alongside other depressive symptoms. Such designs would provide stronger evidence about temporal relationships than a cross-sectional network can provide.

Why concentration matters in university populations

Concentration was another symptom that appeared prominently in both analytical frameworks. Its practical significance may be particularly visible in student populations because cognitive demands are embedded in daily academic life.

Difficulty concentrating can interfere with reading, examination performance, retaining information, completing assignments, and participating effectively in classes. At the same time, concentration difficulties are nonspecific. They can occur with inadequate sleep, anxiety, stress, physical illness, medication effects, substance use, attention disorders, and other conditions. A concentration complaint should therefore never be treated as proof of depression in isolation.

Does a Central Symptom Make the Best Treatment Target?

Not necessarily. This is one of the most important limitations to understand when interpreting network studies.

A symptom that occupies a statistically central position in a cross-sectional network is not automatically the symptom that should be targeted first in treatment. Centrality is a property of the estimated network and can depend on the sample, variables, model, regularization choices, and measurement quality.

Intervening on a central node might theoretically influence other symptoms, but demonstrating that requires longitudinal or experimental evidence. Cross-sectional centrality alone cannot establish that treating one symptom will cause improvements elsewhere in the network.

Clinical care must also consider factors that questionnaire models do not fully capture, including functional impairment, suicide risk, medical conditions, substance use, personal circumstances, prior episodes, treatment history, patient preferences, and co-occurring psychiatric symptoms.

How Stable Were the Network Findings?

The study did not report centrality estimates without evaluating their stability. Reported stability coefficients ranged from 0.31 to 0.44.

These values call for appropriate caution. They suggest that some network conclusions were more robust than others and reinforce the need to view exact rankings as exploratory rather than permanent hierarchies of depressive symptoms.

This distinction is essential when communicating network research. A visually prominent node in a network diagram can appear definitive, but estimates always contain uncertainty. Replication in larger independent samples can help determine whether the prominence of self-worth, concentration, and depressed mood persists across universities, regions, age groups, and cultural contexts.

What the Findings Mean for Depression Screening

The study concluded that cognitive and self-worth symptoms emerged as particularly informative for screening purposes. That conclusion should be understood as an argument for paying attention to symptom-level information, not for replacing validated multi-item questionnaires with a handful of questions.

A total PHQ-9 score remains useful because depression is multidimensional and heterogeneous. Examining individual responses alongside the total score can provide additional context.

For researchers, symptom-level analysis may help answer questions such as:

  • Which items provide the most measurement information in a particular population?
  • Do the same items perform similarly across medical and non-medical students?
  • Which symptoms occupy prominent positions within the observed symptom network?
  • Does item performance replicate across languages, countries, institutions, and demographic groups?
  • Do symptom networks change over time or after an intervention?

These questions extend mental-health measurement beyond simply reporting mean questionnaire scores.

Why the Findings Should Not Be Overgeneralized

The study has several boundaries that matter when interpreting the item-level results.

First, it was cross-sectional. The data represent symptom relationships measured at one period rather than changes tracked through time. As a result, causal or temporal conclusions cannot be drawn from the symptom network.

Second, the sample consisted of 602 undergraduate students in Lahore. Psychometric characteristics are properties of measurements within particular populations and contexts; they should not automatically be assumed to apply to all Pakistani adults, patients with diagnosed major depressive disorder, adolescents, older adults, or populations in other countries.

Third, questionnaire responses are self-reported. They measure reported symptom frequency rather than establishing a clinical diagnosis.

Fourth, although certain symptoms emerged as more discriminating or central, statistical prominence is not equivalent to clinical severity, urgency, diagnostic necessity, or treatment priority.

Finally, the publisher currently identifies the available article as an early-access, unedited version that may undergo further editorial processing. Readers using exact numerical values should consult the current published record.

What This Adds to Student Mental-Health Research

Much student mental-health research reduces depression to prevalence thresholds, mean scores, or comparisons between groups. Those summaries have value, but they provide only one layer of information.

Item-level approaches can reveal a richer structure. The present findings suggest that negative self-worth, concentration difficulties, and depressed mood deserve particular attention when researchers study depressive symptom patterns among university students in Pakistan. Appetite also carried relatively strong discriminatory information in the IRT model, while loss of interest discriminated less strongly in this dataset.

The most compelling aspect is not the creation of a definitive ranking from first to ninth. Rather, it is the demonstration that symptoms contributing to the same PHQ-9 total can behave differently statistically.

This has implications for future psychometric studies. Researchers can evaluate whether the same symptom hierarchy is reproduced in new samples, whether item functioning changes by demographic or academic group, whether centrality patterns persist longitudinally, and whether symptom-level changes predict meaningful outcomes beyond changes in the total score.

The Bigger Lesson: Depression Is More Than a Number

Questionnaire totals are intentionally reductive: they transform complex experiences into interpretable numerical summaries. That is one of their strengths. But the convenience of a total score should not lead us to assume that all nine symptoms behave identically.

In this Pakistani university sample, item response theory indicated that self-worth, concentration, depressed mood, and appetite were among the most discriminating PHQ-9 items, whereas the interest item had the lowest discrimination. Network analysis independently placed self-worth, concentration, and downheartedness among the most central symptoms.

The overlap makes these cognitive and affective symptoms especially interesting targets for further research. It does not establish a universal hierarchy of depression symptoms, nor does it justify diagnosing or treating depression according to individual questionnaire items alone.

The practical message is more measured: researchers and clinicians can learn something from the total score, but they may learn more when they also examine the pattern underneath it.

Medical disclaimer: This article is for educational and research discussion only. The PHQ-9 is a screening and symptom-assessment instrument and does not by itself establish a diagnosis of depression. Individual symptoms, questionnaire scores, and statistical findings should be interpreted in appropriate clinical context by qualified healthcare professionals. Anyone experiencing significant psychological distress, thoughts of self-harm, or concerns about their mental health should seek appropriate professional assessment or urgent assistance when necessary.

Key takeaways

  • Individual PHQ-9 symptoms did not contribute equally to measurement in this sample of 602 Pakistani university students.
  • Item response theory identified self-worth, concentration, depressed mood, and appetite among the PHQ-9 items with the highest discrimination.
  • The loss-of-interest item had the lowest reported discrimination parameter, but this does not make anhedonia clinically unimportant.
  • Network analysis identified self-worth, concentration, and downheartedness as the most central symptoms, creating notable overlap with the item response findings.
  • Cross-sectional network centrality does not demonstrate causation or identify a proven treatment target.
  • Symptom-level analysis can complement total depression scores, but the findings require replication before being generalized to other populations.

Frequently asked questions

Which PHQ-9 symptoms were most informative in the study?
Item response theory indicated that self-worth, concentration, depressed mood, and appetite showed the highest discrimination among the PHQ-9 items in this Pakistani university sample.
Which depression symptoms were most central in the network analysis?
Self-worth, concentration, and downheartedness emerged as the most central nodes in the study’s symptom network, although the authors also reported stability coefficients of 0.31 to 0.44 and the rankings should therefore be interpreted cautiously.
Does high network centrality mean a symptom causes other depression symptoms?
No. A cross-sectional network identifies statistical relationships among symptoms but cannot establish causal direction. Longitudinal or experimental studies are needed to determine whether changing one symptom causes changes in others.
Does the low discrimination of the PHQ-9 interest item mean anhedonia is unimportant?
No. The interest item had the lowest discrimination parameter in this particular sample, but statistical discrimination is not the same as clinical importance. Anhedonia remains an important depressive symptom, and item performance can vary across populations and settings.
Can individual PHQ-9 symptoms be used instead of the total score?
Individual responses can provide useful additional information, but this study does not support replacing the validated PHQ-9 total score or a comprehensive clinical assessment with only a few selected symptoms.
Can the PHQ-9 diagnose depression?
The PHQ-9 is widely used to assess and screen for depressive symptoms, but a questionnaire score alone does not establish an individual clinical diagnosis. Diagnosis requires appropriate clinical evaluation and context.

References

  1. Asghar T, Hassan A, Sahar I, et al. Psychometric properties and symptom profiles of the PHQ-9 and DASS-21 among medical and non-medical university students: a cross-sectional study in Pakistan. BMC Psychology. 2026. https://doi.org/10.1186/s40359-026-05332-5