Abstract
The General Health Questionnaire-28 is a well-known symptom-based rating scale of mental health. Several studies have investigated its latent structure using confirmatory factor analysis. This study questions this approach on several substantive points, most notably the inability for symptoms to interact using confirmatory factor analysis, and argues for the use of network analysis instead. Network results demonstrate the method’s utility to improve our understanding of the rating scales’ symptom structure. Insights include a much richer understanding of comorbidity on the General Health Questionnaire-28 and the identification of particularly salient symptoms affecting the network. It yields substantive information of interest to researchers and practitioners alike.
Keywords
Introduction
Determining the substantive structure of psychological conditions, both theoretically and empirically, is a long-standing issue in psychology and psychiatry (Borsboom et al., 2019; Kendler, 2017). And it appears to be far from being resolved. This becomes evident when one considers the absence of meaningful progress in our understanding of mental disorders, notwithstanding several decades worth of research on the biological bases of mental conditions, whether cellular, metabolic, or genetic (Adam, 2013). This does not mean that there has not been noteworthy progress on biological–mental associations, only that, despite these developments, this approach has been unsuccessful in providing uniquely biological accounts for mental conditions (Borsboom et al., 2019). Many remain convinced that the answer to mental illness lies in reductive explanatory frameworks. The view that mental disorders are nothing but brain disorders (Insel and Cuthbert, 2015) is firmly rooted with the tacit assumption that it is a matter of time until all will be revealed. As such, the view that psychopathological conditions have core biological causes (Borsboom et al., 2019) remains dominant in much of psychiatry and clinical psychology. This view is also represented in our psychometric conceptualizations of psychopathology.
Well-known screening measures of psychopathology such as the General Health Questionnaire-28 (GHQ-28) indeed reflect the view of latent constructs effecting symptoms (Schmittmann et al., 2013). They characteristically comprise so-called reflective or common-cause models, meaning that they contain a number of symptoms assumed to represent some underlying condition(s). Accordingly, an underlying latent variable (i.e. depression) accounts for the patterns of covariance among the symptoms with the assumption that they have nothing in common when the latent variable is removed (i.e. local independence). Or in the parlance of medical science, the presumed disease (i.e. depression) fully explains the presence of the symptoms, which all ought to disappear once the disease is cured (Borsboom et al., 2019). Usually, scores on such rating scales are then summed to represent the severity of the underlying condition.
A substantial literature, however, reveals the shortcomings of the sum-score approach to psychopathology. According to Fried and Nesse (2015), sum scores obscure important information that is better explained and understood at symptom level. Using depression as an example, they argue that while no biomarkers for depression as a syndrome have been found, much of this type of research appears quite promising at the symptom level. Another problem is that antidepressants have differential effects on symptoms, in the sense that they improve different symptoms to different degrees. At the same time, antidepressants can produce side effects which are considered symptoms of depression itself. For instance, common side effects of antidepressants include problems with sleep, concentration, weight change, and appetite change; however, these symptoms are also part of the typical diagnostic criteria for depression.
A further foundational and substantive problem of sum scores is what they are assumed to represent. For example, Fried (2017) has shown that, among seven of the most frequently used measures of depression, there is a serious absence of content overlap among them. He identified 52 different symptoms, of which 40 percent appeared in only one scale, and only 12 percent appeared in all seven scales. In addition, from a common-cause (hereafter used interchangeably with latent variable) model perspective, one would expect depression scales to be unidimensional; however, empirical analysis shows that they tend to be multidimensional (Bagby et al., 2004; Gullion and Rush, 1998; Shafer, 2006).
The fact that psychopathological constructs are “measured” with a vast array of different symptoms and that there appears to be no definitive set of symptoms for each mental condition highlight what Kendler (2017) refers to as the constitutive and indexical views of psychiatric diagnosis. Accordingly, in a “constitutive relationship, criteria definitively define the disorder. Having a disorder is nothing more than meeting the criteria. In an indexical relationship, the criteria are fallible indices of a disorder understood as a hypothetical, tentative diagnostic construct” (Kendler, 2017: 2054). The Diagnostic and Statistical Manual of Mental Disorders (DSM) more or less approximates the constitutive position. The indexical perspective in contrast view symptoms (DSM and non-DSM) as fallible representations with which disorders could be indexed.
Arguably, the most crucial issue with common-cause models, sum scores, and reductionist approaches in general is that symptoms cannot interact directly—and perhaps causally—with each other (Borsboom et al., 2019) since it is assumed that they are independent conditional on the latent variable (Epskamp et al., 2018). This implies that symptoms arise and vary in intensity only as a function of the underlying disorder. In such models, for instance, symptoms such as lack of sleep, fatigue, and trouble concentrating are considered independent indicators of an underlying disorder (i.e. depression) that do not directly influence each other, although it is very reasonable to assume a pattern of influence so that a lack of sleep → fatigue → trouble concentrating (Schmittmann et al., 2013).
Network analysis, in contrast, offers a way to model this critical aspect of clinical symptomology. The method has in recent years gained substantial interest among researchers (Borsboom, 2017; Borsboom and Cramer, 2013; Borsboom et al., 2019; Cramer et al., 2010; Fried and Cramer, 2017; Fried et al., 2017). In network models, patterns of covariation are examined with the assumption that symptoms can mutually influence each other. In addition, network models enable explicit consideration of comorbidity where symptoms can function as causal bridges between mental disorders, for example, between major depression and generalized anxiety disorder (Cramer et al., 2010; Fried and Cramer, 2017).
Against this backdrop, this study conducts a network analysis on the 28-item version of the General Health Questionnaire (GHQ-28; Goldberg and Hillier, 1979) using data from a confirmatory factor analysis previously published in this journal (De Kock et al., 2014). The data were made available by the authors. While confirmatory factor analysis is a latent variable technique and represents a common-cause model, network analysis allows a different perspective. Whereas models based on the former “explain symptom covariation by a latent factor that is viewed as the common cause of all symptoms, network models suggest that syndromes are constituted by the connections among symptoms” (Fried and Nesse, 2015: 79). Thus, confirmatory factor analysis is focused on symptoms and their common relationship with the underlying disorder. It also assumes no interrelation among the symptoms. Network models, on the other hand, view the disorder as a product, or manifestation, of symptom interrelations. As such, it should provide a richer understanding of the symptoms and syndromes represented on the GHQ-28.
It is worth noting that the GHQ-28 was developed to be a general measure rather than a diagnostic one. The original idea was that each score could be interpreted as reflecting the likelihood that a person would get a psychiatric diagnosis upon formal psychiatric assessment (Goldberg and Hillier, 1979). Nonetheless, the conditions represented on the GHQ-28 were, and still are being, conceptualized and examined using latent variable models (De Kock et al., 2014; Goldberg and Hillier, 1979), which carries with it the shortcomings noted above.
This article aims to showcase the utility of network analysis to better understand the symptomology of this measure. Using the same data as those used for the confirmatory factor analysis, network analysis will by implication also illustrate the limitations of a latent variable perspective on the symptoms and syndromes represented on the GHQ-28. In fact, it will be argued that a latent variable application to the GHQ-28 is theoretically inappropriate. Although it is acknowledged that latent variable modeling is convention to test postulated measurement models for purposes of examining construct validity, the argument here is that such a model is theoretically questionable and that it presents an impoverished view of the GHQ-28’s symptom structure.
Method
Participants and procedure
As reported originally by De Kock et al. (2014), participants were 523 employees of the South African National Defense force. All were of indigenous African descent, with the majority being male (83.6%). The sample was quite diverse regarding educational level although most participants had no post-schooling qualification (61.8%). Approximately 40 percent of the participants were between 18 and 25 years of age. Study participation was voluntary and informed consent was obtained verbally after the aim of the research was explained to the participants. More information regarding the data collection procedure can be found in De Kock et al. (2014) and are not repeated here.
Measures
The present data are based on the GHQ-28 (Goldberg and Hillier, 1979). It measures four aspects of mental health, namely, somatic symptoms, anxiety and insomnia (hereafter anxiety), severe depression, and social dysfunction. The participants were asked to report their general mental health over the last few weeks on a four-point Likert-type scale ranging from 1 = not at all to 4 = much more than usual. The item content is provided in Table 1.
GHQ-28 descriptions of the nodes reflected in the network.
GHQ-28: General Health Questionnaire-28; SM: somatic; A: anxiety; SD: social dysfunction; D: depression.
Data analytic procedure
An EBICglasso network was estimated. This is a regularized partial correlation network suited for ordinal data, given its use of polychoric correlations as input (Epskamp and Fried, 2018). In this network structure, nodes represent variables (i.e. symptoms) and edges (lines) represent the fact that two nodes are related after conditioning on all other variables in the data. Edges have weights, which are partial correlation coefficients in the case of EBICglasso models. Like regular correlations, partial correlations range between −1 and 1, but unlike regular correlations they reflect the association between two nodes after controlling for the influence of all other variables in the network (Costantini et al., 2015; Epskamp and Fried, 2018). No edges are drawn between nodes when the partial correlations between them are exactly zero. EBICglasso estimation further makes use of regularization to prevent the identification of spurious edges. Such edges are non-zero partial correlations that appear as weak edges in the network; however, they are likely to be false positives in the sample data due to sampling variation (Costantini et al., 2015). To prevent overinterpretation of spurious effects, least absolute shrinkage and selection operator (lasso; Tibshirani, 1996) regularization is used. In essence, it is a general shrinkage parameter that has the effect of forcing small partial correlations (which are likely to be spurious) to zero.
Centrality indices were inspected next for each of the GHQ-28 symptoms in the estimated network (Opsahl et al., 2010). Three measures, namely, strength, closeness, and betweenness centrality, were evaluated. According to Costantini et al. (2015), strength centrality is the sum of all absolute edge weights for a node with all other nodes. The inverse of the sum of distances of a node to all other nodes provides the closeness centrality index. And the betweenness centrality indicates how frequently a node is in the shortest path between all other nodes. Stated simply, strength quantifies how well a node is directly connected to other nodes, closeness quantifies how well a node is indirectly connected to other nodes, and betweenness quantifies how important a node is in the average path between two other nodes (Epskamp et al., 2018: 3).
Primary weight was placed on strength centrality as it has been found to be the most stable index in psychopathological networks (Epskamp et al., 2018). Nodes were placed using the Fruchterman–Reingold algorithm (Fruchterman and Reingold, 1991), so that nodes with higher centrality values are more likely to end up in the center of the plot, and nodes with lower centrality will be placed toward the periphery. However, caution is recommended regarding substantive conclusions based on exact node placement, as the algorithm can be somewhat erratic, when, for example, there are extremely small differences in edge weights affecting node placement (Epskamp and Fried, 2018).
To examine the accuracy and stability of the network, a bootstrap analysis was performed as recommended by Epskamp et al. (2018). This robustness analysis provides (a) information regarding the stability of the order of centrality estimates and (b) 95 percent confidence intervals (CIs) for all estimated connections (Kendler et al., 2018). Node order stability was examined using a case-dropping bootstrap. This procedure systematically drops proportions of cases and examines how stable the correlation remains with the original data. Specifically, it examines the proportion of data that can be dropped and still retain with 95 percent probability, a correlation of 0.7 or higher between the subsetted and original centrality indices (Epskamp and Fried, 2018). Turning to edge weight CIs, it is important to note that they require careful interpretation when considering the regularization component of EBICglasso estimation. Given that the lasso has already done model selection, the edges included in the network for the present sample are non-zero, even if their CIs overlap with zero. As such, CIs are not interpreted in the conventional way as a test of whether the parameter is significantly different from zero (Epskamp et al., 2018).
Results
The estimated network presented in Figure 1 shows the items (symptoms) of the GHQ-28 as numbered nodes. A couple of features emerge from a visual inspection. The first is that the symptoms do appear to cluster together in clinically meaningful substructures. Anxiety and somatic symptoms cluster in the bottom and depression symptoms appear throughout the middle, with social dysfunction symptoms on top. A second and important feature is that there are also many intermingled symptoms, which suggest that they are substantially interrelated (Kendler et al., 2018). Notably, the anxiety and somatic symptoms did not separate into two distinct clusters as one might have expected, but were in fact quite interrelated. In addition, one depression symptom (D5), which entails feeling incapacitated due to excessive levels of nervousness, were amid this anxiety–somatic cluster, rather than clustering with the other depression symptoms.

Network containing the 28 symptoms of the GHQ. Blue lines represent positive associations, red lines the negative ones, and the thickness and brightness of an edge indicate the association strength.
A summary of interconnectedness, indicated by the centrality indices, are presented in Figure 2 for each of the 28 symptoms of the GHQ-28. The results appear to be in line with the visual summary interpretation. There was some correspondence across all three centrality indices for nodes SM6 and D6. The former reflects a sense of excessive pressure in one’s head (being squeezed) and the latter to wishing one was dead. Further, for closeness and betweenness centrality, feelings of hopelessness (D7) was most central, while suicide ideation (D2) and feeling high-strung (A7), were most closeness central only.

Betweenness, closeness, and node strength centrality estimates for the 28 symptoms of the GHQ (see Table 1 for symptom descriptions of the nodes).
The 95 percent CIs for each edge estimated using 2500 bootstraps are displayed in Figure 3. The results of a case-drop bootstrap analysis are presented in Figure 4. The results of these robustness analyses are contextualized in section “Discussion.”

Network stability of GHQ-28 symptoms. The graph indicates the edge weights (solid red line) and the 95 percent confidence intervals around these edge weights (gray). The bootstrapped mean CIs are indicated with the black line.

The correlation of the centrality of nodes in the original network with the centrality of networks sampled while dropping participants. When the correlation after dropping a substantial amount of participants is high, it means that the centrality estimates in the original network can be considered stable.
Discussion
The primary aim of this article was to examine the symptoms and symptom clusters of the GHQ-28 using network analysis. While the GHQ-28 is a well-known and frequently used measure, it is based on a sum-score, common-cause perspective of the relevant psychopathological conditions. The shortcomings of such a perspective were highlighted earlier, and it was argued that network analysis could enable a more informed and nuanced perspective on the symptom structure of the GHQ-28.
The symptoms of the GHQ-28 were mostly positively connected to each other. The strongest edges in the network indicate that the symptoms clustered largely in their expected syndromes (i.e. depression, social dysfunction). However, there was some important intermingling of symptoms, especially among the anxiety–somatic symptoms which did not cluster separately, with the strongest edges among nodes A3 (feeling under strain) and SM3 (feeling exhausted and without energy), A4 (feeling edgy and bad-tempered) and SM4 (feeling ill), as well as A5 (scared and panicky for no good reason) and SM4 (feeling ill). The high degree of intermingling present here essentially reflects a single anxiety–somatic cluster. This would suggest a high degree of mutual interaction among these symptoms. While undirected networks do not allow directional causal inference (Epskamp et al., 2018), in this case it is probably reasonable to assume causality flowing from anxiety to somatic symptoms. However, the inverse is of course possible, for example, when somatic symptoms are experienced due to a serious disease, which in turn would cause anxiety regarding one’s future health. A more common explanation, however, is that the anxiety induced in everyday life would give rise to somatic symptoms.
Moreover, depression node D5 (feeling incapacitated due to excessive nervousness) was located amid this anxiety–somatic cluster, and not within the depression cluster. Given the content overlap, this is not surprising. It does, however, raise questions regarding the common-cause model conceptualization of depression in the GHQ-28, along with its concomitant sum-score utility. This node (D5) also had strong edge weight connections with nodes A6 (feeling everything is on top of me) and SM6 (feeling as if head is being squeezed).
The many blue edges between the clusters further suggest a high level of mutual influence among the clusters themselves, meaning, broadly, that if an individual suffers from one symptom cluster, there is good chance of suffering from another, or several other clusters, or, at minimum, from some of the symptoms of one or more clusters, but not necessarily from all the symptoms in each cluster. There were, however, also a few negative edges (red lines) in the network, for example, between nodes SD5 (feeling useful) and SM4 (feeling ill), and SD6 (feeling capable of making decisions) and SM4 (feeling ill).
There was a fair amount of consistency on the centrality indices. The nodes with the highest concurrent levels of strength, closeness, and betweenness centrality were SM6 (feel as if head is being squeezed and about to explode) and D6 (wishing one was dead). These were the two most important nodes in the network with regard to their general connectedness and ability to influence other nodes. The five nodes with the highest betweenness centrality were SM6 (feel as if head is being squeezed), D7 (thinking of taking own life), SM4 (feeling as if ill), SD4 (satisfied with the ability to perform tasks), and D1 (thinking you are worthless). The four nodes with the highest closeness centrality were D7 (thinking of taking own life), A7 (feeling nervous), D2 (feel life is hopeless), and SM4 (feel as if ill). The highest strength centrality was observed for nodes SM6 (feel as if head is being squeezed and about to explode), SD4 (satisfied with the ability to perform tasks), and SD6 (feeling capable of making decisions).
The above results need, however, to be weighed against the robustness analyses. As described earlier, the edge weight CIs presented in Figure 3 require careful interpretation given the use of regularization with EBICglasso estimation. As such, the bootstrapped null CIs should not be interpreted as if those edges are not different from zero (Epskamp et al., 2018). Nonetheless, for the present results, the CIs for most edges are quite wide, suggesting a fair amount of variability in edge weight estimation. This implies caution regarding the substantive conclusions for the purpose of generalizing to the broader population.
With regard to centrality stability, it can be seen from Figure 4 that betweenness and closeness drop steeply as smaller subsets are sampled. The correlation stability (CS) coefficient indicates that betweenness (CS(cor = 0.7) = 0.13) and closeness (CS(cor = 0.7) = 0.21) are not sufficiently stable. Node strength appears to perform better (CS(cor = 0.7) = 0.44), but did not quite reach the preferred cut-off of 0.5 (Epskamp et al., 2018). As a result, only node order strength should be considered interpretable, whereas the node order of the betweenness and closeness indices are likely not sufficiently stable for meaningful inference.
It should, on the whole, be evident from the present results that a network analysis of the GHQ-28 provides much information regarding symptom and cluster interaction, which cannot be gleaned from a latent variable perspective. Its assumptions are too constrained and as such cannot reflect the dynamic interactions present on the GHQ. Thus, it underscores the impoverished view provided by a common-cause approach to the current symptom structure. At the same time, it renders questionable the usefulness of the proposed latent variables (syndromes) and their concomitant sum scores, as evidenced by the fact that some symptoms appeared in unexpected clusters rather than within their theoretical “home” clusters.
Moreover, the high degree of anxiety–somatic symptom interaction further highlights the inappropriateness of a latent variable approach to the GHQ. For example, the strong relation between the anxiety and somatic subscales has also emerged in previous confirmatory factor analytic studies (De Kock et al., 2014; Werneke et al., 2000). On these grounds, De Kock et al. (2014) argued that the two subscales might represent a broader dimension of psychological distress. They proceeded to test such a model which yielded good fit. While this is a very reasonable step to take in latent variable modeling, it is theoretically problematic in this case. Now, a new construct is needed to explain both the anxiety and somatic symptoms. This is clearly not a unidimensional construct, and it removes the possibility that anxiety can be the cause of the somatic symptoms (theoretically, a very likely explanation, albeit not the only one).
It is further questionable whether the somatic and social dysfunction subscales of the GHQ-28 can be justified at all within a latent variable framework. This requires thinking about the ontological status of these constructs. At the risk of oversimplification, is social dysfunction indicated by its symptoms, or is it a coherent latent construct that can be conceived of independent of its symptoms? I would argue that this is a case of interpretational confounding and that these constructs might be better considered from a formative rather than a reflective perspective (Bagozzi, 2007; Howell et al., 2007).
Misspecification, whether it is a latent model where a formative model is more appropriate, or a latent model where a network model better represents the psychological phenomena in question, has clear implications for how it is understood clinically. This article sought to demonstrate the problem in the context of the GHQ-28. Moreover, such problems do not only have theoretical significance but also produce important measurement issues (Rhemtulla et al., 2018).
Thus, while latent variable models can be successfully fit to the GHQ, this research argues that it does not allow for adequate representation of its symptom structure. Network analysis offers new insights regarding comorbid structure on the GHQ-28 and allows for a range of new theoretical possibilities to consider. It also enables better understanding regarding the progression among symptoms and clusters. Consider, for instance, well-connected nodes such as SM6 and its strong relations to symptoms in all other clusters. It shows how some symptoms are more important than others regarding the progression of mental conditions, which can be used by practitioners to inform treatment options.
Nevertheless, the present findings regarding the network of the GHQ-28 represent a first approximation of its structure. Several points should be borne in mind when drawing conclusions from the present data. First, the sample comprised adult Africans only, all active in a military setting. While these data are of great value, given the dearth of research on exclusively African populations, and the fact that most psychological knowledge is produced in the so-called western, educated, industrialized, rich, and democratic samples (WEIRD; Henrich et al., 2010), it too represents a small slice of humanity, albeit an important one. Second, the robustness checks also suggest caution when interpreting edges given the variability observed from the bootstrapped CI results, so too the node order on the centrality indices, most notably for betweenness and closeness centrality. Third, research on different population groups will be required to show if the network structure observed in this study reflects stable interactions reflective of the GHQ-28 broadly, or a pattern unique to the present sample.
It is recommended that future studies be conducted with larger samples to further examine the robustness of the present results. In addition, it is important that future studies be done in a variety of different populations. Replications across different countries and cultural/ethnic and gender groups will reveal whether or not the present results are reflective of the GHQ-28 symptomology generally, or reveal the need for nuanced considerations of comorbidity and treatment implications for different populations.
Conclusion
This study argued that a latent variable approach to the GHQ-28 is useful, but has several limitations. Some of these include the assumption or belief that all psychological disorders have root biological causes consistent with the disease models in medical science, with related concerns regarding the appropriateness of latent variable conceptualizations and concomitant sum scores. Against this backdrop, network analysis was introduced and its utility was demonstrated on the data previously used for a latent variable application to the symptom structure of the GHQ-28. It was shown that the network perspective provides new insight with much potential implications for its users, bearing in mind the noted caveats.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
