Abstract
This study presents new analyses of NEO Personality Inventory–Revised (NEO-PI-R) responses collected from a large British sample in a high-stakes setting. The authors show the appropriateness of the five-factor model underpinning these responses in a variety of new ways. Using the recently developed exploratory structural equation modeling (ESEM) technique, the authors show that model fits improve markedly over conventional confirmatory factor analyses (CFA) of the same data set, but that (a) factor interpretations do not change under ESEM analyses, (b) ESEM factor scores, just like CFA factors scores, correlate at near unity with sums of observed scores, (c) NEO-PI-R facets under ESEM analyses are invariant across gender, and (d) ESEM highlights the inappropriateness of alpha and beta as a higher order representation of NEO-PI-R facets, whereas a CFA approach might lead researchers to believe in the appropriateness of these higher order factors. These results, coupled with the existing validity evidence for the NEO-PI-R, suggest that the five-factor structure is the most parsimonious structure for summarizing NEO-PI-R responses from high-stakes settings in the United Kingdom.
More than half a century since its discovery by early personality researchers, the five-factor model of personality (FFM) is now cemented as the dominant framework for describing consistent differences and similarities between the ways people think, feel, and behave (Chamorro-Premuzic, 2007; Goldberg, 1990, 1993; McCrae & Costa, 1987). The literature underpinning the FFM is overwhelming and the breadth of consensus on the value of the FFM is considerable (Chamorro-Premuzic & Furnham, 2010). This latter point is clearly illustrated by opening paragraphs in numerous journal articles, which contain strikingly similar descriptions of how the FFM model is observed across cultures and languages as well as demographic variables such as gender and age (e.g., Saucier & Goldberg, 1998; Saucier & Ostendorf, 1999). In fact, apart from a few notable exceptions (e.g., Block, 1995, 2001), disputes over the FFM’s legitimacy today are centered on issues of refinement rather than questioning the fundamental utility of the model. Some of the aspects that are still being debated relate to the existence or otherwise of a higher order structure of the FFM (e.g., Digman, 1997; Rushton & Irwing, 2009), the status of personality dimensions that are not well represented by the FFM such as honesty-humility (Ashton & Lee, 2008) and ambition (Hogan & Chamorro-Premuzic, in press), the facet-level structure of the FFM (Perugini & Gallucci, 1997), and whether or not a neuropsychological basis exists for what is essentially a taxonomy of phenotypes (De Young, 2010).
The NEO Personality Inventory–Revised (NEO-PI-R), arguably the leading psychometric measure of the FFM, has played a towering role in bringing researchers to this point in our understanding of personality. Although the research base for the NEO-PI-R is vast, and findings have been consistently replicated in numerous and diverse contexts, the NEO-PI-R model has an Achilles’ heel: namely, that confirmatory factor analyses (CFAs) do not yield evidence in support of the FFM when judged by traditionally accepted psychometric standards (Church & Burke, 1994; McCrae, Zonderman, Costa, Bond, & Paunonen, 1996). In fact, the fit of CFA models to NEO-PI-R data has routinely been so poor that McCrae et al. (1996) suggested abandoning CFA. The inability of researchers to fit an appropriate CFA model means that there still remain researchers who argue for both fewer (Rushton & Irwing, 2009) and more (Ashton & Lee, 2008) factors of personality. This is in part because of the recognition that inappropriate use of CFA may result in overestimation or underestimation of the number of personality factors extracted.
Exploratory structural equation modeling (ESEM), a recent development in psychological measurement, offers the potential to reconcile these seemingly opposing vantage points. This is because ESEM has a number of advantages over traditional CFA approaches that mean it could be more appropriate for modeling personality data. These advantages include relaxation of the assumption that items have factorial complexity of one (i.e., no cross-loadings of items or facets), the availability of standard errors for parameter estimates in an exploratory setting, and an assessment of fit using goodness-of-fit indices available in traditional structural equation modeling frameworks (Asparouhov & Muthén, 2009; Marsh et al., 2010). The flexibility of ESEM, as contrasted with CFA, is illustrated in Figure 1 for a hypothetical two-factor model. The CFA model in panel 1A assumes zero loadings on the nontarget factor. On the other hand, the ESEM model in panel 1B allows nonzero loadings on factors other than the primary targeted factor, as illustrated by the dotted arrows. It is thought that this flexibility is more realistic, and these relaxed assumptions are expected to lead to improved fit.

Confirmatory factor analysis (CFA) versus exploratory structural equation modeling (ESEM) representations of a two-factor model
The need for a methodological innovation such as ESEM to reconcile the conflicting evidence from CFA analyses of NEO-PI-R data sets and other methods of validation is well illustrated by research carried out on the short form of the NEO-PI-R. The shorter NEO-FFI (NEO Five-Factor Inventory; 60 vs. 240 items) attracts considerable research interest, no doubt because of its quicker administration time (less than 10 minutes vs. more than 35 minutes for the longer NEO-PI-R). Egan, Deary, and Austin (2000), however, in addition to providing British norms based on 1,025 adults, suggested that the instrument “requires modification and improvement before it can be regarded as measuring the five independent personality traits” (p. 907). In a recent article, Marsh et al. (2010) used the ESEM technique, coupled with theoretically appropriate correlated residuals, to examine fit for the NEO-FFI. Results showed that using ESEM on the NEO-FFI provided a much better model fit than that CFA has exhibited up until this point.
Despite the fact that ESEM has been shown to result in improved model fit for the FFM over traditional CFA approaches, several questions of fundamental importance to assessment specialists using the NEO-PI-R remain unanswered. First, no studies exist that apply ESEM to longer forms of the NEO-PI-R. As a consequence, there is no evidence as to whether the improved fit observed for the NEO-FFI also exists for the NEO-PI-R. There is also no evidence on whether the FFM primary factor interpretations change as a result of the application of ESEM. Critically, a further issue that remains to be investigated is whether ESEM produces person scores that are well approximated by the simple sums of candidates’ observed scores. Because these simple sums are the most routinely interpreted scores in applied settings, the answer to this question has important implications for the accuracy of norm data that is so integral to attributing meaning to personality profiles. Although Marsh, Liem, Martin, Morin, and Nagengast (2011) examined gender invariance for the NEO-FFI, no study has yet examined whether factor scores from ESEM models, with their improved fits, yield FFM scores on the long form NEO-PI-R that are invariant across gender. Finally, we expect researchers and practitioners alike would like to know if an alternative higher order structure exists for more parsimoniously describing personality based on FFM dimensions derived using the more flexible ESEM approach. Several higher order structures of the FFM have been proposed. In a seminal article, Digman (1997) proposed that covariation between the higher order factors of Extraversion and Openness could be explained by a factor he labeled alpha, while covariation among the remaining facets of Neuroticism, Agreeableness, and Conscientiousness could be explained by a factor he called beta. DeYoung (2010) labeled similar higher order structures plasticity and stability based on psychometric as well as neuropsychological evidence. More recently, proponents of a general factor of personality have presented psychometric evidence for what must be the most basic representation of personality so far (Rushton & Irwing, 2009). Given the increased flexibility of ESEM over traditional CFA methods, it would be of interest to practitioners to know whether there is a more parsimonious and well-fitting representation of the FFM that they might use in their applied work based on the factor scores emerging from ESEM. In this article, we answer these questions using a large sample of white-collar British workers who completed the NEO-PI-R in a high-stakes selection setting.
Method
Participants
In all, 13,234 British adults were tested over a 10-year period as part of an assessment center run by chartered organizational psychologists. Each participant completed a number of self-report and ability tests as well as other exercises and an interview. Tests were completed for the purpose of personnel selection or career progression (internal promotion), representing a high-stakes testing scenario as results were used to inform these consequential decisions. Participants’ age ranged from 18 to 67 years and most participants were employed as middle or senior managers in a range of British-based companies. Of the 13,234 adults who participated in the assessment center, 4,937 participants completed the NEO-PI-R and were included in the current study, out of whom 25% were female. The average age of the sample was 44 years and the standard deviation was 14 years.
Measure
The NEO-PI-R (Costa & McCrae, 1992) questionnaire is a 240-item questionnaire designed to measure the FFM traits as well as six primary facets for every trait. Items are responded on a 5-point Likert-type scale that ranged from strongly disagree to strongly agree. The test is untimed but takes approximately 35 minutes to complete. Although a wealth of research exists providing evidence for the validity and the reliability of this instrument, most data are derived from student and low-stakes settings (Chamorro-Premuzic & Furnham, 2010).
Procedure
Participants were tested in an assessment center setting for selection and promotion purposes. The questionnaire was untimed and most participants took between 30 and 45 minutes to complete it. They were asked to respond honestly and were promised, and received, full feedback on their scores at a later point.
Analyses
Descriptive analyses
SPSS 17.0 was used for data cleaning. The calculation of the correlations and the simple descriptive analyses are presented in tables later in the text.
Structural equation models
All structural equation modeling was carried out using the Mplus computer program (Version 6.1; Muthén & Muthén, 2006) with the MLR estimator to handle issues related to nonnormality of the data. We first fitted both CFA and ESEM models to males and females separately, prior to examining measurement equivalence, which must be demonstrated before concluding that survey items measure constructs similarly across populations. Geomin rotation, an oblique rotation method allowing the factors to be intercorrelated, was performed in the case of ESEM models. The consensus today among methodologists is that oblique rotation should always be preferred. This is because we most constructs in social sciences turn out to be intercorrelated, and an oblique rotation will uncover an orthogonal (i.e., uncorrelated) solution if one is appropriate anyway (MacCallum, 1998). CFA does not use rotations because the patterns of fixed and free loadings are specified a priori.
Measurement equivalence
Measurement equivalence prevails if two individuals with equal standing on the construct assessed, but sampled from different populations (i.e., genders), have equal expected observed scores on the measurement instrument (Drasgow, 1984). Today, numerous publications exist that outline the same steps for showing measurement invariance in multiple group models (e.g., Horn & McArdle, 1992; Millsap, 1995; Vandenberg & Lance, 2000). First, an unconstrained baseline model is estimated where the only parameters equated across genders are those required for identification purposes. If a satisfactory fit for the baseline model is observed, configural, or weak invariance, is said to hold. In other words, the same number of factors exists in the data from both groups and items have the same pattern of zero and nonzero loadings in both groups. Next, factor loadings are constrained to be equal across both groups. If constraints on the factor loadings do not reduce fit appreciably, metric, or strong invariance is said to hold. Next, the intercepts for the items are held equal across groups. If the item intercepts constraints do not appreciably reduce fit, strict, or scalar invariance is said to hold. Although further constraints related to the factor variances, factor covariances, and residuals are possible (cf. Marsh et al., 2011), fulfillment of these additional constraints are not requirements necessary for comparisons of mean differences on the construct under study.
Model selection
We report multiple indices in addition to the model chi-square, because its sensitivity to sample size can lead to rejection of theoretically appropriate models (e.g., Byrne, 1998). We selected the most appropriate model out of these sequences of models based on an overall assessment of the following indices: root mean square error of approximation (RMSEA; Steiger, 1990), Bentler’s (1990) comparative fit index (CFI), and Bentler and Bonnet’s (1980) nonnormed fit index (NNFI). The RMSEA is a measure of badness of fit per degree of freedom, and conventionally values less than .05 are considered indicative of good fit (Brown & Cudeck, 1993). The CFI and Tucker–Lewis index (TLI) range between 0 and 1 and values more than .9 are considered acceptable fit, whereas values more than .95 indicate excellent fit (Hu & Bentler, 1995). Importantly, Marsh, Hau, and Wen (2004) have suggested that these standards are unlikely to be achieved with CFA when models as complex as the FFM are analyzed.
Results
Confirmatory Factor Analysis Results
Table 1 presents the results of CFA analyses of facet-level data using Mplus. These results show that the separate models for males and females fitted the data poorly when compared with the conventional standards outlined above. This pattern of poor fit was also observed for the baseline model for configural invariance. In particular, in addition to chi-square statistics, which are highly significant in all cases, RMSEA is more than .10, which indicates poor model fit. Moreover, CFI and TLI are clearly not anywhere near .90 and .95. Because of this, we did not proceed further to investigate metric and scalar invariance, as CFA models are inappropriate for these data.
Fit Statistics for Traditional Confirmatory Factor Analysis Models
Note. df = degrees of freedom; RMSEA = root mean square error of approximation; CFI = comparative fit index; TLI = Tucker–Lewis index, SRMR = standardized root mean square residual.
Exploratory Structural Equation Modeling Results
Results from ESEM analyses indicated an improvement in fit for the male-only and female-only models to what can be argued to be acceptable levels by conventionally accepted standards. The improvement in fit between the CFA and ESEM models is approximately comparable in magnitude with the changes that Marsh et al. (2010) observed by moving from CFA to ESEM on the short form of the NEO-PI-R, the NEO-FFI. The better model fit observed here is most likely because of the relaxation of CFA conditions where each facet is only allowed to load on its target factor and has zero loadings on every other factor. Marsh et al. (2010) referred to this CFA model as the independent clusters CFA model. Table 2 shows that for the separate male and female ESEM models, while chi-square remains highly significant; RMSEA and CFI now indicate adequate fit, whereas TLI is very close to the acceptable level.
Fit Statistics for Exploratory Structural Equation Models
Note. df = degrees of freedom; RMSEA = root mean square of approximation; CFI = comparative fit index; TLI =Tucker–Lewis index, SRMR =standardized root mean square residual.
Measurement Equivalence Results
The model fit for the baseline model testing configural invariance with ESEM, presented in Table 2, indicates fit statistics similar to the single-group ESEM results. Because this is a substantial improvement over CFA results, and fit statistics approach accepted standards, we proceeded to examine metric and scalar invariance. The results in Table 2 show that the imposition of metric and scalar invariance constraints does not reduce fit appreciably. These findings show that the ESEM solution is a more appropriate solution than the CFA solution for these data, and importantly, that there is no differential facet functioning in this data set across gender. The correlations between the factors for the scalar invariance ESEM solution were small to moderate. The largest correlation, at −.45, was between Factor I and Factor V. The smallest correlation, at −.06, was between Factor IV and Factor V.
Interpretation of Exploratory Structural Equation Modeling Factors
One of the primary advantages of the ESEM approach is that it relaxes the assumption that items or facets have zero loadings on all factors other than the target factor. This model restriction is unrealistic for personality data, which are known to be factorially complex (Marsh et al., 2007). Along with the increased flexibility, however, comes the need to interpret the factors as one would in an exploratory factor analyses. In other words, it is quite possible under ESEM that the pattern of factor loadings will not support the a priori FFM structure and patterns of factor loadings need to be examined. Interpretation of the factors that emerge from ESEM ultimately requires judicious interpretation of the loading pattern and significance of the loadings for each of the facets. It seems reasonable, however, to have as our requirement that for a factor to be considered an a priori component of the FFM, all, or at least the majority of the facets that measure the factor ought to have their highest loadings on it, and that all these loadings should be significant.
By these criteria, we see that Factor I is Neuroticism. For both males and females, all the loadings of the N facets are significant on Factor I, and moreover, all facets except N5 Impulsiveness have their highest loadings on Factor I (Table 3). For males, N5 has a higher loading on Factor II and Factor V than it does on Factor I, and only a marginally smaller loading on Factor IV. For females, N5 has a greater loading on Factor V and substantial loadings on Factor II and Factor IV. Constraining these loadings to zero under the independent clusters model that underpins CFA will certainly detract from model fit. In addition to N5 not having its highest loading on Factor I, which is ostensibly Neuroticism, the facet N6 vulnerability has nontrivial loadings on Factor V for men and women. Using the same criteria, Factor II is Extraversion. All Extraversion facets except E3 assertiveness have their strongest loading on Factor II, and all loadings for Extraversion facets on Factor II are significant for men and women. E3 not only has a greater loading in both the male and female samples on Factor IV than it does on Factor II but it also has a smaller and still considerable loading on Factor V.
Loading Parameter Estimates and Significance From Standardized ESEM Solution
Note. Est. = estimated geomin rotated factor loading; p = two-tailed p value. N = Neuroticism; N1 = Anxiety; N2 = Angry Hostility; N3 = Depression; N4 = Self-Consciousness; N5 = Impulsiveness; N6 = Vulnerability; E = Extraversion; E1 = Warmth; E2 = Gregariousness; E3 = Assertiveness; E4 = Activity; E5 = Excitement Seeking; E6 = Positive Emotion; O = Openness to Experience; O1 = Fantasy; O2 = Aesthetics; O3 = Feelings; O4 = Actions; O5 = Ideas; O6 = Values; A = Agreeableness; A1 = Trust; A2 = Straightforwardness; A3 = Altruism; A4 = Compliance; A5 = Modesty; A6 = Tender Mindedness; C = Conscientiousness; C1 = Competence; C2 = Order; C3 = Dutifulness; C4 = Achievement Striving; C5 = Self-Discipline; C6 = Deliberation.
The preponderance of evidence points compellingly toward Factor III representing the FFM Openness dimension. For males and females, all Openness facets have their strongest loading on Factor III except O3 feelings, which loads higher on Factor II (that we have labeled Extraversion) in both men and women. Moreover, the need for flexibility in modeling nontarget loadings is again illustrated by sizable secondary loadings for O6 values. Factor IV in the ESEM solution is Agreeableness, because the highest loading for each agreeableness facet is on Factor IV, and all these loadings are significant. Here again, sizable secondary loadings show the inappropriateness of the independent clusters model for these data. Factor V emerges strongly as Conscientiousness with all conscientiousness facets having their highest loadings on the final factor, Factor V, all of which are significant. In sum, the geomin rotated ESEM solutions for men and women reveal two important points. First, ESEM reveals clear support for the FFM. Second, substantial cross-loading for certain facets, most notably N5, E3, and O3, shows the inappropriateness of the independent clusters model assumptions of traditional CFA approaches if the goal is to obtain strong fit from structural equation modeling.
Is There a Higher Order Structure Based on the NEO-PI-R’s ESEM Solution?
Although second-order ESEM factor models are not currently possible in Mplus, it is possible to save the factor scores from ESEM and subject these to further analysis using ESEM. We used this procedure to investigate one- and two-factor representations of the factor scores that emerged from the best fitting CFA and ESEM models described earlier. First, the single-factor model based on CFA scores did not converge, suggesting that a single-factor model is inappropriate for these data. The single-factor model for the ESEM scores did converge, but fit was so bad as to preclude further interpretation (RMSEA = .32, CFI = .71, TLI = .43). The two-factor model based on CFA factor scores did converge. The factors clearly resembled alpha or plasticity with high Extraversion and Openness loadings (N = −.01, E = .77, O = .87, A = .21), and beta or stability with high-reverse Neuroticism, Agreeableness, and Conscientiousness loadings (N = −.83, E = .45, O = −.01, A = .18, and C = .80). Although we might conclude the existence of higher order factors on the basis and fit (RMSEA = .13, CFI = .99, TLI = .91), something is clearly wrong at a chi-square of 88.86 and a single degree of freedom. Moreover, these values inspected casually would disguise the fact that the values are based on an even worse fitting first-order model (i.e., the multiple-group CFA baseline model presented in Table 1). Despite evidence of alpha and beta based on loading patterns, the model fit and knowledge that the factor solution is based on a poorly fitting first-order solution demands caution before interpreting the supposed alpha and beta factors as substantive factors. Importantly, the two-factor ESEM solution that analyzed the first-order ESEM factors would not converge because of nonpositive definite covariance matrix, suggesting that the higher order solution was not appropriate. Thus, the application of ESEM highlights the inappropriateness of a higher order solution for the Big Five based on these data. The rationale we offer for this result is an intriguing explanation suggested by Ashton, Lee, Goldberg, and De Vries (2009). The essence of the argument of Ashton et al. is that suppressed secondary loadings on nontarget constructs can lead to correlations between factors. Although these correlations among factors can be modeled by higher order factors, the higher order factors accounting for these correlations will be spurious, because the correlations on which they are based are artifactual. On the basis of these results, and if well-fitting models are of concern to practitioners, the most general level of the FFM that yields adequate and defensible fit, so far, is at the level of the five factors.
Correlations Between ESEM Factors, CFA Factors, and Observed Variable Totals
The way that norms for the NEO-PI-R are typically used, as is the case with other personality tests, involves creating observed item totals for candidates and calculating the proportion of a representative norm sample that scores the same as or lower than the candidate. It is today well known that correlations between latent variable scores and observed variable totals are often very high (Fan, 1998). Correlations between latent variable scores and external criteria and observed variable scores and external criteria are also known to be similar (Ferrando & Chico, 2007), justifying use of observed variable scores in the scoring and feedback process. However, assessment practitioners who use the NEO-PI-R will want to know whether the observed variable scores they routinely use are still appropriate, now that a more flexible latent variable modeling approach has yielded better model fit to the NEO-PI-R.
To investigate this issue, we saved the factor scores from both the poor fitting CFA models along with the better fitting ESEM models, and correlated each set of scores both with each other and with their observed variable equivalents. Results, presented in Table 4, indicated that the correlations between both forms of latent scores (i.e., ESEM factor scores or CFA factor scores) and observed variable scores were near unity in almost all instances. This suggests that practitioners may continue to use observed scores in their work because the improved model fit for the ESEM scores does not substantially affect the scores that emerge from these analyses. The exception to this pattern is for the Agreeableness factor, where the correlations are still in excess of .90 for the CFA-observed correlations, but the correlations between the ESEM and CFA and ESEM and observed variable scores drop to .80 and .84, respectively. However, we expect that even these correlations of .80 and above will give many practitioners the confidence to continue using observed variable scores for even the Agreeableness dimension of the NEO-PI-R.
Correlations Among ESEM, CFA, and Observed Variable Scores
Note. CFA–OBS = Correlation between confirmatory factor analysis factor scores and corresponding observed variable sums; ESEM–OBS = correlation between exploratory structural equation modeling factor scores and corresponding observed variable sums; CFA–ESEM = correlations between confirmatory factor analysis factor scores and exploratory structural equation modeling factor scores.
Demographic Descriptive for the Current British Sample
Applied measurement specialists reading this article might be interested in descriptive statistics for this sample to enable them to calculate norms for interpretation in their own work. They might also be interested to examine gender differences or age-related patterns of association with personality. In response, in Table 5, we present means for the overall sample as well as for males and females separately. Effect sizes, calculated as Cohen’s d, are presented for gender. These show that overall the effect sizes are small to moderate, with the largest effect size at the factor level occurring for the Openness domain (d = .52 favoring females) whereas the largest facet-level effect size occurred on O3 feelings (d = .55 favoring females). The magnitude of the differences observed, where just a handful of facets show moderate-sized differences of a half standard deviation, mirrors the findings of Costa, Terracciano, and McCrae (2001). Costa et al. reported that the largest gender difference for the NEO-PI-R was .44 for the N6 vulnerability facet of Neuroticism. Other similarities also exist between our study and theirs, for example, females are uniformly higher on all Neuroticism facets in both studies. The United Kingdom, however, was not represented in Costa et al.’s study. We also recommend caution in making comparisons as it is today clear that the purpose of the personality testing has an impact on the mean levels observed for personality measures (e.g., De Fruyt et al., 2006). That is to say, score inflation in selection settings is likely to render comparisons with data sets from lower stakes personality testing situations inappropriate.
Sample Descriptive Statistics for NEO-PI-R Facets
Note. NEO-PI-R = NEO Personality Inventory–Revised; LB = lower bound; UB = upper bound; N = Neuroticism; N1 = Anxiety; N2 = Angry Hostility; N3 = Depression; N4 = Self-Consciousness; N5 = Impulsiveness; N6 = Vulnerability; E = Extraversion; E1 = Warmth; E2 = Gregariousness; E3 = Assertiveness; E4 = Activity; E5 = Excitement Seeking; E6 = Positive Emotion; O = Openness to Experience; O1 = Fantasy; O2 = Aesthetics; O3 = Feelings; O4 = Actions; O5 = Ideas; O6 = Values; A = Agreeableness; A1 = Trust; A2 = Straightforwardness; A3 = Altruism; A4 = Compliance; A5 = Modesty; A6 = Tender Mindedness; C = Conscientiousness; C1 = Competence; C2 = Order; C3 = Dutifulness; C4 = Achievement Striving; C5 = Self-Discipline; C6 = Deliberation.
Table 5 contains the correlation between the NEO-PI-R domain scores, along with its facets, and age. These results show that, by and large, there are only weak associations between personality and age. Where significant correlations between age and personality are observed, for example with agreeableness, the results show that the correlations are in fact quite small. This means that there is only a very minor association between personality and age. Moreover, significant associations when the correlations are so small suggest that they are the result of high power from the large sample size, rather than showing a substantively meaningful relationship between age and personality.
Discussion
Although the NEO-PI-R is one of the most well researched instruments available for the assessment of broad dimensions of personality, and large norm databases exist, researchers have not been able to show satisfactory model fit for it using modern psychometric methods such as CFA. The current results highlight that CFA, the most common method to analyze the FFM factors, makes unrealistic assumptions with regard to the factorial complexity of each of the NEO-PI-R facets. These facets clearly have nonzero loadings on numerous factors, and this violates what Marsh et al. (2007) referred to as the independent clusters model. Importantly, we showed that the factors retain their a priori interpretations when modeled at facet level using ESEM. The factor scores estimated based on the well-fitting ESEM model were then shown to be correlated almost perfectly with both the scores estimated based on the CFA model and their observed variable total counterparts. The biggest discrepancy occurred for agreeableness, although correlations were still large (.80 with CFA scores and .84 with ESEM scores). The ESEM model also showed measurement equivalence for the number of factors (configural invariance), facet loadings (metric invariance), and facet intercepts (scalar invariance). This suggests that the relation between facet-level scores and the latent personality dimension is the same for both males and females. Finally, analyses of the ESEM factors revealed a structure that resembled Digman’s (1997) alpha and beta for ESEM analyses of factor scores derived from CFA, but this model fitted poorly, and was based on factor scores from a first-order model that had even worse fit. Under ESEM analyses of first-order factor scores derived from a well-fitting first-order model, a two-factor model was shown to be inappropriate because of inadmissible solutions. The most general level of interpretation appropriate for responses to the NEO-PI-R, therefore, is still the FFM if model fit is of concern.
Limitations of this study ought to be mentioned to facilitate interpretation of results and guide future research. First, these data are collected in a setting where individuals were being considered for selection and promotion. Results observed here might not generalize to other situations in which personality might be measured (e.g., research purposes, or solely for development). Care should be taken then when examining the consistency between these findings and other studies using data from other contexts. Our study used ESEM to model facet-level data, and it would be useful to apply the same analyses at the item level for the NEO-PI-R. Finally, although ESEM has not revealed any strong implications for changing applied practice, we have only so far examined measurement models (i.e., modeled responses to the questionnaire itself). It might be that when ESEM is used in a broader setting that included criterion variables we will identify criteria for which the better fit yields predictive improvements. Future research should investigate this issue.
Because of the unique nature and characteristics of this sample we also presented descriptive statistics and examined gender differences and age associations for each of the NEO-PI-R factors and facets. These show that gender differences are small, and age-related associations are weak. Indeed, it is because past research has shown only small demographic associations with personality (Sackett, Schmitt, Ellingson, & Kabin, 2001) and still predicts performance that, tests such as the NEO-PI-R are used so widely in occupational settings. Because the gender differences were small and associations with age were weak, the descriptive statistics presented here might be used by industrial–organizational psychologists wanting to create norms for high-stakes selection in the United Kingdom without too much concern over impact against women and older workers at mid-to advanced career stages.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
