Abstract
Keywords
The Center for Epidemiological Studies Depression Scale (CES-D) is a 20-item self-rating questionnaire developed by Radloff (1977) that is used widely in research and clinical settings for screening depressive symptoms. It has been used with various age groups, including older adults (e.g., Ros et al., 2011), adult community samples (e.g., Sheehan, Fifield, Reisine, & Tennen, 1995), and adolescents (e.g., Cheng, Yen, Ko, & Yen, 2012), and in different cultures and cross-cultural studies (e.g., Mackinnon, McCallum, Andrews, & Anderson, 1998). Although the CES-D was initially developed for epidemiological studies (Radloff, 1977) it has also been used in clinical populations (Weissman, Sholomskas, Pottenger, Prusoff, & Locke, 1977) and primary care settings (Thomas & Brantley, 2004).
The CES-D has empirically supported scales for Depressed Affect (DA), (lack of) Positive Affect (PA), Somatic Symptoms and Retarded Activity (SS), and Interpersonal Difficulties (ID). However, the recommended (Radloff, 1977) and common practice for screening depressive symptoms has been to use the total score based on all items, rather than the scores for the different scales (Edwards, Cheavens, Heiy, & Cukrowicz, 2010). If the use of the total score is to have credibility, it would be necessary to demonstrate, at the very least, that most of the covariance of the ratings on the CES-D items can be explained by a single general factor, and for the general factor to have good reliability and external validity. To date these have not been clearly demonstrated. For a multidimensional scale such as the CES-D, the support for a general factor can be evaluated using a bifactor model (an orthogonal factor model, with one general depression factor on which all the CES-D items load, and specific factors for DA, PA, SS, and ID items). For CES-D ratings provided by a group of older adults, the current study examined support for a bifactor model, and also the internal consistency reliability and external validity of the factors in this model.
In the original validation study, based on three independent groups, Radloff (1977) found support for a four-factor oblique model, with factors for DA, PA, SS, and ID (see Figure 1). These factors were derived from principal components analysis. This procedure involves the analysis of the variables in a data set (such as the ratings of the CES-D items) simultaneously to uncover how the items can be grouped together. While exploratory factor analysis (EFA) does the same, it differs from principal components analysis. Principal components analysis establishes the groups or more specifically the components via the linear relationships of the variables, whereas EFA establishes the groups or more specifically the factors in terms of the common variances shared between the variables. Variances that are not shared, called error variances, are not considered when producing the factors. Thus, EFA solutions are based on the common factor model. Confirmatory factor analysis (CFA) is also based on the common factor model. However, unlike EFA which is exploratory, the goal of CFA is to validate a proposed (a priori) factor model. CFA studies of the CES-D have confirmed the original oblique four-factor model (Knight, Williams, McGee, & Olaman, 1997; Sheehan et al., 1995), although a few studies have supported a three-factor model, with factors for combined DA and SS, PA, and ID (e.g., Dick, Beals, Keane, & Manson, 1994). There has been little support for a one-factor model. The four-factor model has also been supported in older adults (Chapleski, Lamphere, Kaczynski, Lichtenberg, & Dwyer, 1997; Foley, Reed, Mutran, & DeVellis, 2002; Hertzog, Van Alstine, Usala, Hultsch, & Dixon, 1990; Ros et al., 2011), the target group in the current study. CFA studies have also shown support for a second-order factor model, with the four-factor structure as primary factors and a single second-order factor of general depression (Hertzog et al., 1990; Knight et al., 1997; Sheehan et al., 1995; see Figure 1).

CES-D models examined in the study.
For the second-order factor model, McCallum, MacKinnon, Simons, and Simons (1995) used Schmid–Leiman parameterization to compute the loadings of the items on the general factor and the primary (more accurately specific) factors for a group of older adults. The findings indicated that loadings on the DA and SS specific factors were generally low (mostly less than .20), loadings for the PA and ID factors were appreciable (around .50), and loadings on the general factor were high (all ≥0.578). The general factor accounted for 76.43% of the total common variance, whereas the specific factors for DA, PA, SS, and ID accounted for 4.98%, 5.67%, 5.29%, and 7.64%, respectively. These findings can be interpreted in terms of a dominant single general factor for the CES-D that accounts for most of the covariance in the CES-D item scores. However, as this study used only 16 of the 20 CES-D items and only two of the original four response categories, this study did not provide a test of the bifactor model for the standard 20-item CES-D.
Apart from using the Schmid–Leiman parameterization on a second-order factor model, another approach that can be used to evaluate the presence of a dominant general factor is the bifactor model (Reise, Moore, & Haviland, 2010). From a confirmatory perspective, a bifactor model is a latent structure with a general factor, on which all items load, and two or more domain specific factors. The general and specific factors are uncorrelated (orthogonal model). The general factor represents the covariance among all the scale items, whereas the specific factors represent domain specific item response covariance that is not accounted for by the general factor. There are also no cross-loadings of items across the specific factors. Figure 1 shows a conceptual path diagram of the bifactor model as applied to the CES-D. As shown, it is a confirmatory factor model with five orthogonal factors: a general factor on which all items load, and specific factors for the DA, PA, SS, and ID items. The general factor explains the covariance across all the DA, PA, SS, and ID items, whereas the specific factors explain the unique variance of the items within the specific factors.
The bifactor model is different from a conventional second-order factor model. The higher order factor in a second-order factor model represents a superordinate factor explaining the covariance of the primary factors, and not the items, as is the case of the general factor in a bifactor model (Reise et al., 2010). Another important difference is that unlike the second-order factor model, the bifactor model does not impose the proportionality constraint (i.e., for a given set of subtests, the ratios of variance attributable to the respective first-order factor to variance attributable to the general factor are constrained to be the same). Because of this constraint, the relationships of the general and the specific factors in this model with external variables can only be examined if the regression path of the general factor or one of the specific factors to the external variable is fixed at zero (Reise et al., 2010). Thus, the external validity of all factors cannot be examined simultaneously. Consequently, the bifactor model can be considered to be better than the Schmid–Leiman parameterization on the second-order factor for evaluating the factor loadings on the general and specific factor and the external validity of these factors (Reise et al., 2010). If it can be demonstrated that the CES-D general factor accounts for most of the variance for the CES-D ratings, and that it has good internal reliability and external validity, then it could provide good support for using the total CES-D score for research and clinical practice.
For a group of older adults, Yang and Jones (2008) examined the applicability of the bifactor model for a nine-item version of the CES-D, with specific factors for “dysphoria,” “psychosomatic,” and “lack of positive affect.” They found support for the bifactor model, with the loadings for all items being higher on the general factor than the specific factors. In a subsequent study involving the same group, Yang, Tommet, and Jones (2009) examined the applicability of the bifactor model for the 20-item version of the CES-D. Although there were specific factors for DA, PA, SS, and ID, based on EFA solution, one of the original DA items was allocated to the PA specific factor, and another one to the SS specific factor. This study found support for their bifactor model, with the factor loadings for most of the items being higher on the general factor than the specific factors. However, as the items allocated to the specific factors in this study were slightly different from how they would be allocated in terms of the original factor structure (Radloff, 1977), and as the study did not examine the internal reliability and external validity of the general and specific factors, the findings in the study can be seen as limited. There is clearly a need for more studies in this area.
In the current study, we examined and compared the fit of several CES-D models, using all 20 item ratings provided by older adults from the general community. Initially, CFA was used to examine the fit for the bifactor model. In this model there were five orthogonal factors: a general factor on which all items loaded, and specific factors for DA, PA, SS, and ID. The items for the specific factors were identical to those in the original four-factor model proposed by Radloff (1977). We compared the fit of this model with a one-factor model (all 20 items loading on a single factor), the four-factor model proposed by Radloff (1977), and the second-order factor model, with four lower order factors loading on a single higher order factor. The lower order factors were DA, PA, SS, and ID and the items for these factors were again identical to those in the original four-factor model proposed by Radloff (1977). The four models tested are depicted in Figure 1. We then examined the salience of the factor loadings, the amount of common variance explained, and the internal consistency reliability values of the general and specific factors in the bifactor model. Additionally, we examined the external validity of the factors of the bifactor model in terms of their associations with independent measures of depression and suicidal ideation. In this respect, there are data showing that CES-D total scores correlate well with clinical ratings of depression (Roberts & Vernon, 1983), and with suicidal ideation (McLaren, Gomez, Bailey, & Van der Horst, 2007). Based on past findings reported by Yang et al. (2009) and McCallum et al. (1995), we expected support for the bifactor model, with higher loadings, variance explained, and reliability value for the general factor than the specific factors. We also expected the general factor to have stronger associations with depression and suicidal ideation than the specific factors, with the association being higher for depression than for suicidal ideation (since depression is more aligned to the general factor than is suicidal ideation).
Method
Participants
A convenience sample of 1,178 older adults aged 65 years and older participated in the study. Apart from age, there was no other exclusionary criterion for participation. The sample consisted of 573 men aged from 65 to 96 years (M = 74.98, standard deviation [SD] = 7.48) and 605 women aged from 65 to 98 years (M = 75.01, SD = 7.65). The men and women did not differ significantly on age, t(1176) = −0.06, p > .05. The majority of men were married (n = 313, 54.6%), with the other men being widowered (n = 149, 26.0%), divorced (n = 56, 9.8%), or single (n = 55, 9.6%). The majority of men had completed secondary school or equivalent (n = 401, 70.0%), with the remaining having completed primary school (n = 87, 15.2%), or a university degree (n = 85, 14.8%). Almost half of the women were married (n = 284, 46.9%), with the remaining women being widowed (n = 180, 29.8%), divorced (n = 95, 15.7%), or single (n = 7.6%). Most of the women had completed secondary school or equivalent (n = 453, 74.9%), with the others having completed primary school (n = 73, 12.1%), or a university degree (n = 79, 13.0%). There was a significant association between gender and marital status, χ2(3, N = 1168) = 14.35, p = .002, with more men being married (n = 313; Haberman’s standardized adjusted residuals statistic or HAR = 2.6, p < .01), and more women being divorced (n = 95; HAR = 3.0, p < .01) than that expected by chance (n for married = 597, and n for divorced = 151). There was no relationship between gender and highest level of education, χ2(2) = 3.74, p > .05. Men (M = 11.10, SD = 10.31) and women (M = 10.54, SD = 9.44) scored similarly on the CES-D, t(1166) = 0.98, p > .05. Based on a total scale score of 16 or more that is generally used to identify older adults with a higher probability of clinical depression (Lewinsohn, Seeley, Roberts, & Allen, 1997), there was no relationship between gender and percentage scoring at a clinical level, χ2(2) = 3.74, p > .05. For the current sample, 24.3% of men and 22.7% of women had a CES-D total score of 16 or greater. Apart from this information, we had no other information on the mental health of participants.
Materials
The Center for Epidemiologic Studies Depression Scale
As noted in the introduction, the CES-D (Radloff, 1977) is a 20-item self-report measure which assesses the number of depressive symptoms that a person has experienced (e.g., I felt that I could not shake off the blues even with help from my family). This measure comprises scales for DA (seven items; Items 3, 6, 9, 10, 14, 17, and 18), PA (four items; Items, 4, 8, 12, and 16), SS (seven items; Items 1, 2, 5, 7, 11, 13, and 20), and ID (two items; Items 15 and 19). Participants rated each item on a 4-point scale that represented the number of days the participant experienced the emotion or thought in the statement (0 = less than 1 day, 1 = 1-2 days, 2 = 2-4 days, 3 = 5-7 days) over the preceding 7 days. The four PA items are reverse scored. Higher scores indicate more severe depressive symptoms. In terms of internal consistency, past studies have shown that the alpha coefficient of the full CES-D scale is high, with values usually above .85 (e.g., Knight et al., 1997; Radloff, 1977; Ros et al., 2011). Knight et al. reported values of .87, .78, .75, and .44 for DA, PA, SS, and ID, respectively. Cronbach’s alpha internal consistency values for DA, PA, SS, and ID were .87, .80, .80, and .59, respectively, in the current sample.
Suicide Subscale in the General Health Questionnaire
The Suicide Subscale of the General Health Questionnaire (GHQ; Goldberg & Hillier, 1979) has seven items. Although all seven items together are intended to measure suicidal ideation, three of the items (Items 1, 2, and 4) tap depression symptoms, and the other four items (Items 3, 5, 6, and 7) tap suicidal ideation (Watson, Goldney, Fisher, & Merritt, 2001). The three depression items and the four suicidal ideation items were used in the current study as indicators to model latent factors for depression and suicidal ideation, respectively. To avoid confusion with the factors in the CES-D, these factors are referred to henceforth as GHQ–Depression and GHQ–Suicidal Ideation. Consistent with previous research (Goldney, Winefield, Tiggemann, Winefield, & Smith, 1989; Hamilton & Schweitzer, 2000), the responses for items were scored from 0 to 3, with higher scores indicative of depression or suicidal ideation. Cronbach’s alpha internal consistency values for GHQ–Depression and GHQ–Suicidal Ideation were .66 and .76, respectively, for the current sample.
Procedure
Ethics approval was obtained from Federation University’s Human Research Ethics Committee. The sample was recruited for a variety of smaller studies, each assessing depressive experiences in older adults as one of the aims. All participants were from the general community, and they were recruited through several sources from the State of Victoria, Australia. Participants were recruited in person. At various venues (e.g., assisted living facilities, social and sporting clubs), one of the researchers gave a brief presentation about the study. They explained the procedure, and interested individuals were given an envelope with the questionnaires and a participant information letter. The questionnaires were completed by participants at the time and returned to the researcher. A total of 1,400 questionnaires were distributed, with 1,178 usable questionnaires being returned, indicating a response rate of 84%.
Analytical Procedure
Initially, the normality of the CES-D items was examined using PRELIS 2.51 (Joreskög & Sorböm, 1996). Table 1 shows the mean and standard deviation (SD) scores, and the skewness and kurtosis values of the items. All items showed significant skewness and kurtosis (p < .001). There was also significant multivariate skewness (110.665, z = 117.348, p < .001) and kurtosis (796.316, z = 51.020, p < .001). Given this, the study used maximum likelihood with robust estimation to ascertain statistical fit. This procedure, referred to as MLMVχ2 in Mplus, tests the closeness of fit between the unrestricted sample covariance matrix and the restricted (model) covariance matrix, correcting for the lack of normality.
Mean and Standard Deviation, and Skewness, and Kurtosis of the CES-D Items.
Note. CES-D = Center for Epidemiologic Studies Depression Scale. All skewness and kurtosis were significant (p < .001).
All the CFA models in the study were conducted using Mplus (Version 7) software (Muthén & Muthén, 2013). As all types of χ2 values, including MLMVχ2, are inflated by large sample sizes, the fit of the models was also examined using two commonly used practical fit indices: the root mean square error of approximation (RMSEA) and the comparative fit index (CFI). The guidelines suggested by Hu and Bentler (1998) are that RMSEA values of .06 or less be taken as good fit, values >.060 to .08 be considered moderate fit, values >.08 to .10 be considered marginal fit, and values >.10 be considered poor fit. For the CFI, values of .95 or higher are taken as indicating good model–data fit, values of .90 and <.95 are taken as acceptable fit, and values <.90 as poor fit. It is to be noted, however, that the appropriateness of these “benchmarks” has yet to be established for bifactor analyses (West, Taylor, & Wu, 2012). To determine statistical differences between models, the difference in MLMVχ2 values was used. As this difference is not distributed as a chi-square, it is necessary to adjust for this difference. This was done using the scaling correction formula for MLMV on the Mplus website.
The internal consistency reliability of the general and specific factors was computed using McDonald’s omega hierarchical (ωH). Zinbarg, Revelle, Yovel, and Li (2005) have recommended this method for a scale that has a hierarchical factor structure, such as a bifactor structure. For a bifactor model, the ωH of the general factor is model-based reliability computed in terms of the proportion of variance for the general factor to the total observed variance for the scale. For a specific factor, the ωH is model-based reliability computed in terms of the proportion of variance for the relevant specific factor to the total observed variance for the subscale (Brunner, Nagy, & Wilhelm, 2012). According to Zinbarg et al. (2005), if the value of ωH for the general factor is high, then there is justification for summing the responses of the items in the scale. The converse of this argument is that if the value of ωH for a specific factor is high, then there is justification in using its scale score, independent of the support for a general factor.
In order to examine the association of GHQ–Depression and GHQ–Suicidal Ideation with the factors in the bifactor model, the bifactor model was extended to include the latent factors for GHQ–Depression and GHQ–Suicidal Ideation, and these factors were correlated with the general and specific factors. The GHQ–Depression and GHQ–Suicidal Ideation factors were correlated. According to Hemphill (2003), coefficients of <.2 are small effects, .2 to .3 are medium effects, and >.30 are large effects. These guidelines were used in the study to interpret the effect sizes of the correlations in these models.
Results
Fit of the Bifactor Model and Comparisons With Other Models
Table 2 shows the results of all the CFA models tested. 1 For the bifactor model, the factor loadings of the two items on the ID scale were set equal, and Items 17 and 10 of the DA scale were also constrained to be equal. Without these modifications the bifactor model showed nonidentification and inadmissible solution. Based on guidelines proposed by Hu and Bentler (1998), the goodness-of-fit values (RMSEA and CFI) for the one-factor model indicated poor fit. For the four-factor model, the second-order factor model, and the bifactor model, there was good fit in terms of their RMSEA values, and acceptable fit in terms of their CFI values. For the four-factor model, the correlations involving PA with DA, SS, and ID were .36, .31, and .32, respectively, whereas the correlations among the other factors ranged from .90 to .70. For the second-order factor model, the loadings for DA, PA, SS, and ID on the general factor were .99, .36, .90, and .77, respectively. Table 2 also shows that there was no significant difference in fit between the four-factor model and the second-order factor model, and both these models showed better fit than the one-factor model. The bifactor model had better fit than all the other models tested. Overall, therefore, the findings indicated that of the four models tested, the bifactor model was the optimum structural model to represent ratings on the CES-D.
Fit of the Factor Models of the CES-D.
Note. CES-D = Center for Epidemiologic Studies Depression Scale; CFI = comparative fit index; CI = confidence interval; RMSEA = root mean square error of approximation.
p < .001.
Salience of Factor Loadings of the General and Specific Factors of the Bifactor Model
Table 3 presents the factor loadings and the decomposed variance estimates for the bifactor model. In an absolute sense, with the exception of the four PA items, for all other items the loadings on the general factor were relatively higher than the loadings on the specific factors. As shown in the table, for the general factor, three of the four items belonging to the PA factor had nonsalient loadings, and all other items had salient loadings (based on Thurstone’s recommendation that values ≥ .30 are salient). For the specific factors, salient loadings were found for two (out of seven) DA items, three (out of four) PA items, one (out of seven) SS item, and both the ID items.
Completely Standardized Factor Loadings and Sources of Variance of the CES-D in the Bifactor Model.
Note. CES-D = Center for Epidemiologic Studies Depression Scale; G = General; DA = Depressed Affect; PA = Positive Affect; SS = Somatic Symptoms; ID = Interpersonal Distress. For each item the residual is 1 − (variance for the item on the general factor + variance for the item on the specific factor). As an example, the residual for Item 3 is 1 − (.572 + .002) = .426.
Variance Explained by the General and Specific Factors of the Bifactor Model
Table 3 also shows that the specific factors for DA, PA, SS, and ID accounted for 4.27%, 18.89%, 4.72%, and 3.53%, respectively, of the common variance. Thus, together, the specific factors accounted for 31.41% of the common variance, and the general factor accounted for the remaining 68.57% of the common variance. The PA specific factor accounted for 60.14% of the total common variance accounted for by all the specific factors.
Internal Consistency Reliability Values of the General and Specific Factors
The ωH value for the general factor was .85. The ωH values for DA, PA, SS, and ID were .049, .708, .075, and .225, respectively. Thus, the value for the general factor was high, the value for PA was moderately high, and the values for DA, SS, and ID were low, based on values between .70 and .79 being moderate, and values of .80 or more as being high (Clark & Watson, 1995).
Associations of Depression and Suicidal Ideation With the Factors in the Bifactor Models
Table 4 shows the correlations of GHQ–Depression and GHQ–Suicidal Ideation latent factors with the factors in the bifactor model. As shown, the general factor and the PA specific factor were significantly and positively associated with GHQ–Depression and GHQ–Suicidal Ideation, with large and medium effect sizes, respectively. None of the other correlations were significant. Using Steiger’s (1980) z-to-r transformation formula for comparing correlations measured on the same participants, the correlation for the general factor with GHQ–Depression was significantly higher than the correlation for the general factor with GHQ–Suicidal Ideation, z = 3.74, p < .001. The correlation for the PA specific factor with GHQ–Depression did not differ significantly from that for the PA specific factor with GHQ–Suicidal Ideation, z = 1.14, ns.
Correlations of the Factors in the Bifactor Model With the GHQ–Depression and GHQ–Suicide Ideation.
Note. CES-D = Center for Epidemiologic Studies Depression Scale; GHQ = General Health Questionnaire.
p < .001.
Additional Analyses Based on Previous Results
For the PA items we found mostly nonsalient loadings on the general factor and salient loadings on its specific factor, relatively high variance accounted by the PA items to the PA specific factor, moderately high ωH value for the PA specific factor, and significant associations of the PA specific factor with GHQ–Depression and GHQ–Suicidal Ideation. Therefore, we examined the support for a modified bifactor model in which the PA items were not part of the general factor, and independent of the DA, SS, and ID specific factors. The goodness-of-fit values for the modified bifactor model were MLMVχ2(df = 156) = 454.32, p < .001; RMSEA = .040 (90% CI = .036-.045); CFI = .92. Although this model showed reduced statistical fit compared with the standard (original) bifactor model, ΔMLMVχ2(Δdf = 4) = 41.60, p < .001, like the standard bifactor model, its RMSEA value indicated good fit, and its CFI value indicated acceptable fit, thereby indicating somewhat comparable qualitative fit across them. Like the general factor in the standard bifactor model, the ωH value for the general factor in the modified bifactor model was high, at .89. Also, the general factor in the modified bifactor model was significantly and positively associated with GHQ–Depression (r = .40) and GHQ–Suicidal Ideation (r = .30), with large effect sizes.
Discussion
Initial analysis indicated good support for the original CES-D bifactor model, with this model showing better support than the four-factor model, the higher order factor model, and the one-factor model. The four-factor model and the second-order factor model showed mixed fit, with the RMSEA showing good fit and the CFI showing adequate fit. The one-factor model showed poor fit. Thus, there was some support for the four-factor model and the second-order factor model, but not the one-factor model.
The fit findings for the four-factor model, the second-order factor model, and the one-factor model are consistent with existing data on the CES-D (Knight et al., 1997; Sheehan et al., 1995), including data involving older adults (Chapleski et al., 1997; Foley et al., 2002; Hertzog et al., 1990; Ros et al., 2011). The support of the bifactor model in the current study is also consistent with that reported by McCallum et al. (1995), which used Schmid–Leiman parameterization on a second-order factor model to examine a bifactor conceptualization of the CES-D. However, unlike the current study that used all 20 items in the CES-D and the original four response categories, the study by McCallum et al. used only 16 items and only two response categories. The support here for the bifactor model can also be interpreted as consistent with the study by Yang et al. (2009). Indeed for a slightly modified bifactor model (in which one of the original DA items was allocated to the PA specific factor, and another one to the SS specific factor), they reported comparable fit (RMESA of .025, compared with .037 in the current study). However, as the current study provided an evaluation of the bifactor model of the standard CES-D items, response categories, and items-to-factors structure, the findings for the bifactor model in the current study extend existing data in this area.
The findings for the original bifactor model in the study showed that with the exception of three PA items, all other items had salient loadings on the general factor. With the exception of the four PA items, for all other items the loadings on the general factor were relatively higher than the loadings on the specific factor. For the specific factors, salient loadings were found for mainly PA and ID items. Overall, the general factor accounted for slightly more than twice the common variance of all the specific factors together. The specific factors for DA, PA, SS, and ID accounted for 4.27%, 18.89%, 4.72%, and 3.53%, respectively, of the common variance. Somewhat comparable with these findings, the study by McCallum et al. (1995) found that the general factor accounted for 76.43% of the total common variance, while the specific factors for DA, PA, SS, and ID accounted for 4.98%, 5.67%, 5.29%, and 7.64%, respectively. Taken together, these findings indicate that the general factor is relatively more dominant than the specific factors. While the general factor explains most of the covariance in the scores of the CES-D items, the findings in the current study showed that this is especially so for the DA, SS, and ID items.
The findings of the current study also showed that for the original bifactor model, the internal consistency reliability value (based on ωH) for the general factor was .848 and the internal consistency reliability values were .049, .708, .075, and .225 for DA, PA, SS, and ID, respectively. The value for the general factor was high, the value for PA was moderately high, and the values for DA, SS, and ID were low, based on values between .70 and .79 being moderate, and values of .80 or more as being high (Clark & Watson, 1995). Thus, there was weak support for the internal consistency reliabilities for DA, SS, and ID. In contrast, there was acceptable support for the internal consistency reliabilities for the general factor and the specific factor for PA.
The current study also examined the external validity of the factors of the original bifactor model in terms of their associations with GHQ–Depression and GHQ–Suicidal Ideation. The findings showed that the general factor and PA were significantly and positively associated with GHQ–Depression and GHQ–Suicidal Ideation, with large and medium effect sizes, respectively. None of the other correlations were significant. Also, the correlation for the general factor with GHQ–Depression was statistically significantly higher than the correlation for the general factor with GHQ–Suicidal Ideation. Taken together, the correlation findings can be interpreted as supportive of the external validity of the general factor and bifactor model. Also, the finding that the PA specific factor was associated with GHQ–Depression and GHQ–Suicidal Ideation can be taken as supporting the external validity of this factor.
The lack of support for the internal consistency reliabilities of the DA-, SS-, and ID-specific factors indicates that there is no justification for using their scale scores, as the general factor captures most of the variance in these scales. In contrast, the support for the internal consistency reliabilities for the general factor and the PA specific factor raises the possibility that the total and PA scale scores could be used. However, as the differential meaning of a general factor that includes the PA items from a PA factor with the same set of PA items as in the general factor score is unclear, this may not be prudent. Our findings showed that although the fit of a modified bifactor model in which the PA items were not part of the general factor and were independent of the DA, SS, and ID specific factors, was less than the standard bifactor model, the difference in fit between the modified and the standard model was low. Also, the ωH value for the general factor in the modified bifactor model was high, and it was significantly and positively associated with GHQ–Depression and GHQ–Suicidal Ideation, with large effect sizes. Thus, the general factor in the modified bifactor model had comparable support for its internal consistency reliability and external validity. Given this, it can be argued that despite the fact that the modified general factor had slightly lower fit than the standard general bifactor model, from a practical viewpoint, it would be simpler to use a total score without including the scores of the PA items (i.e., sum of DA, SS, and ID scores), and the PA scale score. We recommend this method of scoring the CES-D. It is to be noted in this respect, that our scoring preference is consistent with that proposed by other researchers (Edwards et al., 2010; Schroevers, Sanderman, van Sonderen, & Ranchor, 2000).
In summary, the findings in this study supported the original bifactor model for ratings of the CES-D. For this model, the general factor was relatively more dominant than the specific factors. The general factor explained most of the covariance in the scores of the CES-D items for DA, SS, and ID items. Most of the covariance in the scores of the PA scale was explained by the PA specific factor rather than the general factor. The findings also supported a modified bifactor model in which the PA items were not part of the general factor, and independent of the DA, SS, and ID specific factors Taken together, these findings were interpreted as indicating that the CES-D is best scored in terms of a total score using only the DA, SS, and ID items, and also a score for PA using the four PA items. Thus, our scoring recommendation is inconsistent with the use of the total score based on all the DA, SS, ID, and PA items, as recommended (Radloff, 1977) and used generally for screening depressive symptoms (Edwards et al., 2010). The use of a total score based on only the DA, SS, and ID items is, however, consistent with that suggested by some researchers (Edwards et al., 2010; Schroevers et al., 2000).
In concluding, the findings and interpretations made in this study have limitations. First, although we had an 84% return rate, we had no information on the 16% who did not respond. Thus, we do not know how the missing data from these individuals affect our findings. Second, because we examined older adults, the findings here could be unique to this age group, and not applicable to other age groups. Related to this limitation is that we used a convenience sample. Third, even for older adults, as the sample involved a normal community group of older adults, the findings may not be applicable to other groups of older adults, such as those with mental health or cognitive problems. Indeed past studies have shown that the CES-D is valid and reliable for identifying clinical levels of depression in clinical groups who are at high risk for developing major depression (Hann, Winter, & Jacobsen, 1999; Pandya, Metz, & Patten, 2005). Fourth, it is possible that factors such as socioeconomic, ethnicity, and marital status could influence CES-D ratings (Yang et al., 2009). The failure to control for these effects in the study could have confounded the results. Fifth, the findings reported here are based on a single study. As a consequence, there is a need for cross-validation of the findings before the findings can be generalized. Sixth, for the bifactor model, the factor loadings of both items on the ID scale, and Items 17 and 10 of the DA scale, were constrained to be equal. Without these modifications, the bifactor model showed nonidentification and inadmissible solution. It is uncertain if the findings were confounded by these modifications. Although we have highlighted a number of limitations, we believe that the findings in the current study add to the literature on the bifactor model of the CES-D, and the adequacy of using the total CES-D score for research and clinical practice, at least with older adults. It would be useful if more studies were conducted in this area, taking into consideration the limitations highlighted here.
Footnotes
Acknowledgements
The authors wish to thank Angelina Crea, Mitchell Hobbs, Simon Morris, Karyn Newnham, Jayne Turner, and Kellie Wilson for collecting the data.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
