Abstract
In the field of international educational surveys, equivalence of achievement scale scores across countries has received substantial attention in the academic literature; however, only a relatively recent emphasis on scale score equivalence in nonachievement education surveys has emerged. Given the current state of research in multiple-group models, findings regarding these recent measurement invariance investigations were supported with research that was limited in scope to few groups and relatively small sample sizes. To that end, this study uses data from one large-scale survey as a basis for examining the extent to which typical fit measures used in multiple-group confirmatory factor analysis are suitable for detecting measurement invariance in a large-scale survey context. Using measures validated in a smaller scale context and an empirically grounded simulation study, our findings indicate that many typical measures and associated criteria are either unsuitable in a large group and varied sample-size context or should be adjusted, particularly when the number of groups is large. We provide specific recommendations and discuss further areas for research.
In the field of education, international large-scale assessments, such as the Trends in International Mathematics Study (TIMSS) and the Programme for International Student Assessment (PISA), are tasked with measuring what students know and can do internationally. Further, non-performance based education studies, such as the Teaching and Learning International Survey (TALIS), seek to measure study participants on latent variables that deal with attitudes, perceptions, and experiences. In both types of study, performance on tasks or scores on other latent variables are typically summarized in terms of measurement model-based scale scores (Olson, Martin, & Mullis, 2008; Organisation for Economic Co-operation and Development [OECD], 2010a).
Internationally, surveys are not limited to the educational sphere. Indeed, internationally comparative studies abound. For example, the World Health Organization administers a World Health Survey. Similarly, UNICEF is responsible for three cycles of the Multiple Indicator Cluster Survey, which monitors the situation of women and children globally. Regardless of the survey topic or specific data source, an important criterion for comparing scale scores in an international context is that the latent variable is understood and measured equivalently across all countries. This property is typically (although not only) referred to as measurement invariance (Meredith, 1993), lack of bias (Lord, 1980) or absence of differential item functioning (Hambleton & Rogers, 1989; Mellenbergh, 1982; Swaminathan & Rogers, 1990). Although equivalence of scale scores in large-scale achievement tests has received substantial attention in the academic literature (e.g., Allalouf, Hambleton, & Sireci, 1999; Ercikan, 2002, 2003; Grisay & Monseur, 2007; Hambleton, Merenda, & Spielberger, 2005), only a relatively recent emphasis on scale score equivalence in international nonachievement education surveys has emerged (OECD, 2010a).
To that end, the 2008 cycle of TALIS used multiple-group confirmatory factor analysis (MG-CFA; Jöreskog, 1971) to provide evidence of score comparability on a number of scales designed to measure and compare teachers internationally in areas such as beliefs and practices (OECD, 2010a). TALIS researchers supported their findings regarding measurement invariance with research that was limited in scope to few groups and relatively small sample sizes (Chen, 2007; French & Finch, 2006). It is notable that relying on these limited studies is not a deficiency of TALIS, but rather an indication of a general lack of scholarly evidence regarding the feasibility and tolerance of multiple-group models in international large-scale assessment or other large-scale survey-type contexts. This paucity is reflected in relatively recent calls for more research into the performance of fit indices when the number of groups is greater than two (Meade, Johnson, & Braddy, 2008) and when the sample sizes are large (French & Finch, 2006). In response, the proposed study examines the extent to which typical model evaluation measures in MG-CFA are suitable for detecting measurement invariance in large-scale educational studies such as TIMSS, PISA, and TALIS, where the number of groups is large and the sample size within groups is varied—from relatively small to relatively large. Although, as noted, there exist many other international surveys, we use TALIS as a prominent example and one from which findings should generalize to other international surveys and multiple-group comparisons with large numbers of groups and varied sample sizes.
Background
As international educational surveys have grown in number and in numbers of participants, the psychometric and cross-cultural scholarly communities have noted numerous methodological challenges and opportunities that arise in this particular context. To point to just a few, issues on instrument adaptation (Hambleton, 2002), scale score comparability (Oliveri & von Davier, 2011), and even defining “culture” (LeTendre, 2002) are manifold when dozens of countries—with different languages and cultures—participate. From a psychometric perspective, even on a short scale the number of possible pairwise, cross-country differences on any parameter can be large. For example, a four-item scale where only cross-country factor loadings are under consideration, the possible pairwise differences are
Given the wealth of studies that examine the performance of MG-CFA in the two-group case (e.g., Chen, 2007; Cheung & Rensvold, 2002; French & Finch, 2006; Meade et al., 2008; Meade & Lautenschlager, 2004), we do not delve deeply into the details surrounding MG-CFA. Instead, we briefly discuss the method in general, summarize the findings from current research in the field, and describe the approach used in TALIS, which we consider in the current article. Limiting our empirical focus provides us with a platform from which to more broadly consider the method of investigating cultural equivalence in a large-group, varied sample size context. Furthermore, the scope and scale of TALIS is general enough for the findings to appeal to a wide audience. Finally, the method adopted by TALIS researchers is one generally recommended in the literature and used more broadly by researchers who are interested in providing some evidence of the scale score validity of a measure in multiple populations.
Cultural/Measurement Equivalence/Invariance
With the growth of cross-cultural studies such as TIMSS and TALIS, an interest in analyzing these types of data has also burgeoned in recent years. This assertion is supported by a June 2013 search for scholarly (peer reviewed) journal articles on the Academic Search Premier database using the Boolean search string “(cultural or measurement) and (invariance or equivalence).” This search resulted in 40 articles published from 1980 to 1989, 210 articles published from 1990 to 1999, and an incredible 2,545 articles published from 2000 to 2009. This search includes articles undertaking within-country comparisons (e.g., evaluating measurement equivalence by gender or Mexican compared with European Americans). On the other hand, this search might well omit inquiries into broader model invariance, including general latent variable model invariance. The example nevertheless underscores the scale and growth of research concerned with this method.
Generally, investigations of measurement invariance focus on the degree to which comparisons on the latent variable of interest (e.g., teacher beliefs) can be validly compared across heterogeneous populations. Although approaches vary (Schmitt & Kuljanin, 2008; Vandenberg & Lance, 2000) and can depend on the research question (Bollen, 1989), the influential work by Horn, McArdle, and colleagues (Horn & McArdle, 1992; Horn, McArdle, & Mason, 1983) and Meredith (1993) guides much of the current applied measurement invariance work. We first begin with the general factor model given by
Albeit with some variations,
2
the general approach followed in practice is that if the null hypothesis,
Given the well-known sensitivity of the chi-square test of model fit to sample size (Bagozzi, 1977; Bentler & Bonett, 1980), dozens of fit-evaluation alternatives have been offered. In practice, a small handful of these are recommended for use in empirical research using multiple-group models. In particular, Cheung and Rensvold (2002) recommended a change in the comparative fit index (CFI; Bentler, 1990) equal to or greater than −.010 as evidence of non-invariance. This finding was supported in further research by French and Finch (2006) in the multivariate normal context. Furthermore, Chen (2007) suggested that changes in CFI (ΔCFI) equal to or greater than −.010 supplemented by a change in the root mean square error of approximation (RMSEA; Steiger & Lind, 1980) less than or equal to .015 were indicative of noninvariance when sample sizes were equal across groups and larger than 300 in each group. Chen also recommended ΔCFI not less than −.005 and and ΔRMSEA at least as small as .010 when sample sizes were unequal and each sample size was smaller than 300. Finally, Chen recommended changes in the standardized root mean square residual (SRMR; Bentler, 1995) no larger than .005 to test for scalar invariance in the case of small, uneven samples to ΔSRMR of less than or equal to .030 when testing for metric invariance in the case of large, even sample sizes. Operationally, the OECD has adopted the following criteria: ΔCFI at least as large as −.010 and ΔRMSEA less than or equal to .010 as sufficient evidence of noninvariance of either slopes or intercepts (OECD, 2010a). We maintain consistency with the OECD in the current article and adopt the same criteria for evaluating levels of invariance.
For a number of practical reasons, much methodological research in the area of measurement invariance has been limited to the two-population case, where there is a reference population and a focal population (Holland & Thayer, 1988). In spite of this limitation in the methodological literature, applied researchers can and do compare larger numbers of populations (De Beuckelaer, Lievens, & Swinnen, 2007; Torsheim et al., 2012) but they typically use criteria validated for two-group investigations (Chen, 2007; Cheung & Rensvold, 2002; French & Finch, 2006). As such, little is known about the performance of these fit statistics and indices for determining measurement invariance when the number of groups is greater than two and the sample sizes are varied. The current article begins to investigate this issue. In particular, we pose the following research question: Are currently accepted measures for evaluating measurement invariance suitable when the number of groups and within-group sample sizes are more typical of those found in cross-national surveys?
To provide further empirical context for the current article, the number of groups participating in several of the most current international education surveys is included in Table 1. We also note that TALIS 2008 is currently the only study with reported empirical results of measurement invariance analyses. As such, we use these results as the foundation for our investigation. To that end, TALIS analysts evaluated several scales for measurement invariance and although the hypothesis of both configural and metric invariance were supported (OECD, 2010a), scalar invariance was generally not supported. To get a sense of the degree of difference in TALIS parameter estimates across countries, we fit configural invariance models to two arbitrarily chosen teacher scales (Classroom Disciplinary Climate, with four items and Structuring Teacher Practices, with five items). In both cases, we adhered to current operational procedures and assumed that the observed variables were normally distributed. Descriptive statistics for the intercept and slope estimates on both scales across the 23 countries can be found in Table 2, where we can see large differences in both slopes and intercepts. Differences are particularly notable on the Structuring Teacher Practices scale, where the range of the slopes across countries and items is about 2 and the intercept differences range from 3.30 to 3.94, depending on the item. Given that these items have five response options (from 1 to 5), cross-country differences of this magnitude in intercepts and slopes are quite significant.
Number of System-Level Participants in Several International Education Studies.
Note. TIMSS = Trends in International Mathematics Study; PISA = Programme for International Student Assessment; TALIS = Teaching and Learning International Survey.
Descriptive Statistics From Normal MG-CFA Models Fit to Two Selected TALIS Scales Across 23 Countries.
Note. MG-CFA = Multiple-group confirmatory factor analysis; TALIS = Teaching and Learning International Survey.
Method
Study Design
We used a simulation study in order to address our research question as it related to the performance of MG-CFA and associated fit criteria when the number of groups was relatively large and sample sizes within each group were varied. To more closely mimic actual data and operational procedures, data were simulated as ordered categorical but analyzed with a normal model. The process used to select population item- and person-parameter values is discussed subsequently. To obtain reasonable values for our population parameters, we selected the same two scales described above (Structuring Teacher Practices and Disciplinary Climate), to which we fit unidimensional ordered categorical MG-CFA models to the TALIS 2008 data. Based on the empirical results, we then calculated the mean and variance of each of the parameter values across countries and scales to arrive at a reasonable basis for our generating distributions. For example, we estimated the mean and variance for all slopes across the 23 participating countries and selected scales to be 1.255 and 0.506, respectively. See Table 3 for population means and variances along with generating distributions used in the simulation study. We then randomly drew from each generating distribution to create population values for each of the relevant item parameters. Generating population values were largely assumed to be normally distributed with the exception of the residual variance parameters, which we assumed to be distributed as
Parameter-Generating Distributions Based on Empirical Analysis of TALIS Scales.
Note. TALIS = Teaching and Learning International Survey.
Latent variable means were sampled once for each group from N(0.223, 0.436), which helped to ensure meaningful between-group differences. We sampled latent variable variances for each group from a uniform distribution (U[0.5, 2.5]) because our originally chosen distribution,
We used the generating distributions from which to create a number of conditions, where we manipulated several factors, including the length of the scale, the number of noninvariant items, the nature of noninvariance, and the number of groups. For each simulated condition, the population model was assumed to be unidimensional, which is consistent with the approach used by the OECD for developing noncognitive scales (OECD, 2010a). Furthermore, to create reasonable conditions for examining the performance of typical multiple-group models and associated evaluation criteria, the models were considered to be correctly specified for all groups under consideration. In other words, the hypothesis of configural invariance should hold across all groups. A description and rationale for including each of these study factors values is provided below.
Number of groups
The number of groups examined in the study was set at 10 or 20. Recall that one of the purposes of the study was to examine measurement invariance criteria in a relatively large group setting. Based on this motivation, these group sizes more closely approximate the operational contexts of large-scale surveys, where group sizes typically range from around a dozen, particularly in the case of field trials, to 60 or more countries for the main survey in larger studies. For our particular analysis, group sizes of 10 and 20 are similar to the TALIS context, which is the only study that currently has published results associated with establishing measurement invariance for noncognitive scales (OECD, 2010a).
Scale length
The number of items per scale considered in the study was set at five or six. This scale length is typical of scales in studies such as TALIS, where most noncognitive scales (e.g., teacher beliefs about teaching profession) are composed of just a few items (OECD, 2010a).
Noninvariant items
The number of items affected by noninvariance was zero, two, or three, regardless of the length of the scale. In five-item scales with noninvariant items this implied that either 40% or 60% of the items were affected by cross-group parameter noninvariance. In the six-item scales, this meant that 33% or 50% of the items were noninvariant.
Nature of noninvariance
To examine the impact of different sources of parameter noninvariance on fit evaluation measures, three types of noninvariance were simulated for the affected items: slope, threshold, or both slope and threshold. For the relevant conditions we used an approximate one-third split across groups for the purposes of simulating noninvariance. The noninvariant parameters were assigned such that in the 10-group condition, 3 randomly selected groups served as the reference group, 4 randomly chosen groups were assigned higher values of noninvariant parameters, whereas the remaining 3 groups were assigned lower values of noninvariant parameters. Similarly, in the 20-group condition, using random assignment again, 7 groups served as the reference group, whereas 7 groups were assigned higher valued noninvariant parameters, followed by the remaining 6 groups that were assigned lower population values on the parameter of interest. For all conditions, we fixed the number of thresholds per item to 4 (implying five response options), which is generally representative of noncognitive scales in TALIS and other international studies. For both referent and noninvariant items, all parameters were sampled once.
Sample sizes for the current study varied from 600 to 6,000 per group in both the 10- and 20-group conditions. In particular, sample sizes started with 600 and increased by 450 and 250 in the 10- and 20-group conditions, respectively. Selected sample sizes were randomly assigned to each group to avoid confounding sample sizes with either invariance or values of the latent variable means or variances. In addition, assigned sample sizes were a fixed characteristic of a group across conditions within a set group size. For example, the sample size assigned to Group1 was 1,170 across all 10-group conditions and 1,500 across all 20-group conditions. In other words, whereas the sample size was randomly assigned to each group, it was fixed across conditions for that particular group. These sample sizes and group numbers were considerably different from previous simulation research that examined measurement invariance issues (Cheung & Rensvold, 2002; French & Finch, 2006; Lubke & Muthén, 2004; Meade & Lautenschlager, 2004); however, they better represented empirical data structures in the international studies of interest including TALIS, PISA, and TIMSS, where group sample sizes range from the hundreds to thousands.
Analysis
Data generation and analysis
Data were simulated in Mplus 6.12 (Muthén & Muthén, 1998-2010) according to an ordered categorical model with parameters as previously described (see Table 3 for distribution means and variances for population parameters). To set the scale and origin of the latent variable for multiple-groups data generation we used the methods described in Millsap and Yun-Tein (2004). In line with operational procedures in TALIS, the data were subsequently analyzed under an assumption of normality in the observed variables. And the origin and scale of the latent variable were set by fixing the first group’s latent variable mean to 0 and latent variable variance to 1 while the first factor loading in each group is fixed at 1. Although this is a meaningful misspecification and can result in interpretation problems (Lubke & Muthén, 2004), it is the current method of choice, if investigations of measurement invariance are done at all in international educational assessments. Our study design yielded a total of 28 conditions (2 group sizes × 2 scale lengths × 2 noninvariant item sets × 3 sources of noninvariance + 4 fully invariant conditions) and each condition was replicated 500 times.
Evaluation criteria
In order to examine the performance of MG-CFA fit measures in a large-scale assessment context, we proceeded according to the operational methods used by the OECD (2010a). In particular, MG-CFA models were fit to all groups’ data simultaneously, by condition, where first a test of configural invariance was followed by increasingly restrictive tests of metric and scalar invariance. We assess the fit of the configural model to each condition’s data via the chi-square test of model fit, CFI, Tucker–Lewis Index (TLI), RMSEA, and SRMR. Accordingly, the chi-square test should be statistically not significant, CFI and TLI should be no smaller than .950, RMSEA should be no larger than .050, and SRMR should be no larger than .080. Similarly, for each subsequently restrictive set of models (metric followed by scalar invariance), we began the model evaluation by first examining the overall fit of the model to the data via the five fit measures just listed. We term the fit measures that result from the test of configural invariance and overall model fit measures for metric and scalar invariance the overall fit measures. Next, we tested for metric invariance followed by scalar invariance. To examine the plausibility of metric and scalar invariance, we used the chi-square difference test, ΔCFI, and ΔRMSEA. Again, similar to Chen (2007) and in line with operational OECD procedures, we considered changes in CFI not larger in magnitude than −.010 and changes in RMSEA less than or equal to .010 as reasonable evidence of invariance. We refer to the results of this group of tests as relative fit results to differentiate between these results (where a more restrictive set of assumptions are compared with a less restrictive set of assumptions) and the overall fit results just noted. For all results, we report the average fit statistics and indices across the 500 replications for each condition and invariance test.
Results
In the following section, we present the results for the five-item scale (overall fit at each assumed level of invariance followed by relative fit for metric and scalar invariance). We then present the results in a similar fashion for the six-item scale. Results for the fit measures are reported as averages across the 500 replications within each respective condition along with standard deviations to indicate the degree to which results varied across replications.
Five-Item Conditions
Overall fit
As shown in Table 4, Panel A, the average chi-square values under configural invariance were very large relative to the respective degrees of freedom across each condition, suggesting that the chi-square test detected misfit at configural invariance in all conditions. Average chi-square values for overall fit under an assumption of metric invariance were predictable in that the results supported further deterioration in model fit for all conditions, including when data were generated as fully invariant. We also observed that the chi-square test statistics were consistently higher in the 20-group conditions than in equivalent 10-group conditions. As discussed in the “Method” section, the data were simulated as ordered categorical but the models assumed normality of observed variables, which suggested here that the chi-square had very high power to detect this type of model misspecification.
Average Fit Values for Configural Invariance (Overall Fit) in Five-Item (Panel A) and Six-Item (Panel B) Conditions for 10 and 20 Groups.
Note. CFI = comparative fit index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; TLI = Tucker–Lewis index.
Although the chi-square test results point to misspecification in all conditions that assumed configural invariance, the fit indices, namely the RMSEA, CFI, TLI, and SRMR, yielded better results in all invariance conditions and group sizes. In particular, the CFI and TLI were all larger than .980 and the SRMR never exceeded .020 in all conditions, suggesting acceptable model fit. The RMSEA, on the other hand, yielded values in the range of .059 to .073, which is slightly higher than the typical .050 cutoff.
As the level of assumed invariance was increased from configural to metric and from metric to scalar invariance, as expected, the fit indices yielded poorer model fit (see Table 5). For example, when the slopes were assumed invariant, fit indices yielded lower CFI and TLI values (ranging between .840 and .900, in Conditions 3 through 6 and 11 through 14) and higher SRMR values and RMSEA values greater than .20 when metric invariance was examined, whereas in fully noninvariant and conditions with threshold noninvariance (Conditions 1, 2, and 7 through 10), fit indices remained at acceptable levels, although the RMSEA values suggested only marginal fit. Similarly, expected patterns were found when scalar-invariant models were fit to the data in all conditions. In particular, indices suggested acceptable fit only in fully noninvariant conditions, particularly when the number of groups was 10, with marginal RMSEA values. In addition, the number of groups had differential impacts across the fit indices, although in general, 10-group conditions had more favorable outcomes than their 20-group counterparts. For example, when data were generated with two noninvariant slopes and thresholds, all fit indices suggested better model fit in the 10-group condition compared with the 20-group condition. In contrast, when two thresholds were modeled as noninvariant, all fit indices except for the SRMR slightly favored the 20-group over the 10-group condition in terms of better fit.
Overall Average Fit Values for Metric (Panel A) and Scalar (Panel B) Invariance in Five-Item Conditions.
Note. CFI = comparative fit index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; TLI = Tucker–Lewis index. Under Level heading, 0, 2, and 3 correspond to number of invariant items; N = none, T = thresholds, S = slope, S&T = slopes and thresholds; Nj = number of groups.
Relative fit
As shown in Table 6, the chi-square difference test was statistically significant, suggesting rejection of both metric and scalar invariance, for all conditions, including both fully-invariant conditions. In addition, and as expected, an increase in the number of groups (from 10 to 20) resulted in consistently larger chi-square difference test statistics for both metric and scalar invariance models across all studied conditions.
Relative Fit Tests Results for Five-Item Conditions.
Note. CFI = comparative fit index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual.
With respect to the relative fit indices, several interesting findings emerged. Overall, fit indices performed well at examining hypotheses of metric invariance. For example, ΔCFI supported metric-invariant hypotheses for 10-group data that were compatible with the models (Conditions 1, 7, and 9), whereas all conditions with noninvariant slopes yielded values much greater in magnitude than the −.010 cutoff. In the equivalent 20-group conditions, ΔCFI values were close to or somewhat greater than the −.010 cutoff value (−.010, −.014, and −.014 for Conditions 2, 8, and 10, respectively). This finding provides some evidence that for larger groups, a more liberal criterion for establishing measurement invariance might be warranted. Findings associated with the ΔRMSEA indicated good performance for the 10-group fully invariant condition (Condition 1) and two-item scalar noninvariant condition (Condition 7). In both situations, the ΔRMSEA values were below the generally accepted .010 criteria. In the 20-group conditions, where data were either fully invariant (Condition 2) or scalar noninvariant (Conditions 8 and 10), the ΔRMSEA indices ranged from .022 to .026, also suggesting that a more liberal criterion is warranted in situations with larger number of groups. Furthermore, in the 10-group condition where three items were scalar noninvariant, the average ΔRMSEA value was .022.
In examining relative model fit for scalar invariance (i.e., further imposing equality constraints on the model parameters), fit indices were largely effective at correctly identifying fully invariant data and also at discriminating between metric- and scalar-invariant data. Specifically, the average ΔRMSEA and ΔCFI values were mostly within typically accepted criteria for both 10- and 20-group, fully invariant data (Conditions 1 and 2), suggesting that the hypothesis of scalar invariance would be supported for both these conditions, assuming that a metric invariance hypothesis was supported in a previous step. Additionally, data generated to have only noninvariant thresholds had associated relative fit indices that were outside generally acceptable cutoff values for 10- and 20-group cases (Conditions 7, 8, 9, and 10). Both these would provide some evidence that typically accepted cutoffs for relative fit indices were reasonably well suited for retaining a hypothesis of fully invariant data and for determining if the metric-invariant data were additionally scalar invariant.
As Table 6 shows, in several instances, additional constraints yielded average relative fit index values that would support a scalar invariance hypothesis for data that were not commensurate with this assumption. In general, however, this finding is not consequential as long as evidence against metric invariance was observed in a previous step. For example, in Condition 3, where two slopes were noninvariant, average ΔCFI reported was −.006. Similarly, in Condition 6, where three slopes are noninvariant, the average ΔRMSEA reported was .001. But in both conditions, the fit indices associated with a test for metric invariance were well outside acceptable criteria, rendering a further test of scalar invariance unnecessary.
Six-Item Conditions
Overall fit
Under the assumption of configural invariance (Table 4, Panel B), average chi-square values were very high and relative to the degrees of freedom, average values were highly statistically significant for each condition. This suggests that, according to the chi-square test of model fit, none of these data meet minimum criteria for configural invariance. According to Table 7, Panel A, chi-square values for overall fit under an assumption of metric invariance are predictable, in that imposing restrictions on the data resulted in a decrement in fit under all conditions, even for data that are fully invariant. Similarly, when a model that assumed scalar invariance was fit to data for all conditions (Table 7, Panel B), model fit declined, according to the chi-square test. Furthermore, the chi-square test exhibited typically expected performance in that for data generated under equivalent conditions (e.g., full invariance), the test statistics were consistently higher when the number of groups was greater (i.e., 10 compared with 20). Similar to the five-item conditions, these findings suggest that the chi-square has very high power to detect this sort of model misspecification. And, given that in each condition the simulated data are commensurate with an assumption of configural invariance, 3 the chi-square test will generally not be well suited for evaluating the hypothesis of configural invariance in this sort of situation.
Overall Average Fit Values for Metric (Panel A) and Scalar (Panel B) Invariance in Six-Item Conditions.
Note. CFI = comparative fit index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; TLI = Tucker–Lewis index. Under Level heading, 0, 2, and 3 correspond to number of invariant items; N = none, T = thresholds, S = slope, S&T = slopes and thresholds; Nj = number of groups.
Despite the finding that the chi-square test of model fit rejects all models under the assumption of configural invariance, the fit indices, including the RMSEA, CFI, TLI, and SRMR performed much better in all conditions and group sizes. In particular, the CFI and TLI were all larger than .98 and the SRMR was not larger than .02, suggesting that these indices support the hypothesis of configural invariance. One exception to this trend was the RMSEA, none of which averaged below the typically accepted cutoff of .05. Instead, under all conditions, this index varied between .058 and .079, indicating less-than-optimal model fit. And, as expected, further restrictions on the models resulted in deteriorations in these indices. Furthermore, violations of hypothesized levels of invariance were easily detected by all fit indices. As an example, data that were simulated such that three slopes and three sets of thresholds varied across the groups had associated overall fit indices that fell well outside any conventionally accepted criteria in both the 10- and 20-group conditions when a model that assumed metric invariance was fit to the data. In contrast, data that were commensurate with a more stringent level of invariance than the level under consideration only experienced slight deteriorations in model fit (e.g., a metric-invariant model fit to fully invariant data). One notable finding here is that when only the number of groups differed but all other parameters were equal across a pair of conditions, the SRMR tended to be larger for the 20-group condition and this was particularly pronounced for models that assumed scalar invariance. This finding is somewhat surprising given that this index is not explicitly a function of sample size (Bentler, 1995).
Relative fit
Based on the results presented in Table 8, the chi-square difference test is strongly statistically significant, suggesting rejection of both metric and scalar invariance hypotheses, for all conditions, including for fully invariant data. And, as expected, the test statistic values were larger when only group sizes differed across two comparable conditions (e.g., when two slopes are metric noninvariant). Similar to the findings of overall fit, the chi-square difference test is overly powerful under the conditions examined and is likely not a useful measure for examining measurement invariance under these conditions.
Relative Fit Tests Results for Six-Item Conditions.
Note. CFI = comparative fit index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual.
A number of interesting findings emerged with respect to the relative fit indices, which we discuss currently. Based on the six-item conditions examined, the ΔRMSEA generally performed very well at examining hypotheses of metric invariance in the 10-group condition. In particular, fully invariant (Condition 1) and scalar noninvariant data (Conditions 7 and 9) had associated ΔRMSEA well below the .010 criteria. And all conditions with noninvariant slopes had fit indices larger than .10 (more than 10 times typical cutoff values). In contrast, the ΔRMSEA values for 20-group conditions where the data were commensurate with a hypothesis of equal slopes (Conditions 2, 8, and 10) were all larger than .025, suggesting that a more liberal criterion might be warranted with larger numbers of groups. Similarly, the ΔCFI supported metric-invariant hypotheses of 10-group data that were commensurate with the models (Conditions 1, 7, and 9), whereas all conditions with noninvariant slopes had relative fit index values much greater in magnitude than the −.010 cutoff. And similar to the ΔRMSEA measure, the ΔCFI values for commensurate 20-group data (Conditions 2, 8, and 10) were somewhat over the −.010 cutoff. By condition, average values were −.013, −.016, and −.017, respectively. This again indicates that a more liberal criterion is warranted for larger numbers of groups.
Added equality constraints on the model parameters to examine the hypothesis of scalar invariance resulted in fit indices that were largely effective at identifying fully invariant data and at discriminating between metric- and scalar-invariant data. Specifically, the average ΔRMSEA and values were well within typically accepted criteria for both 10- and 20-group, fully invariant data, suggesting that the hypothesis of scalar invariance would be supported for both these conditions, assuming that a metric invariance hypothesis was supported in a previous step. Furthermore, data generated to have only noninvariant thresholds had associated relative fit indices that were well outside generally accepted cutoff values for both group sizes (Conditions 7, 8, 9, and 10). As with the five-item conditions, these results offer some evidence that generally accepted cutoffs for both relative fit indices are reasonably well suited for retaining fully invariant data and for determining whether metric-invariant data are additionally scalar invariant. And although in some conditions, additional constraints resulted in average relative fit index values that would generally provide evidence in favor of the scalar invariance hypothesis, this is of little consequence as long as evidence against metric invariance were heeded in a previous analytic step. Examples of this include the average ΔCFI value for Condition 3, where two slopes are noninvariant (−.005) and the average ΔRMSEA value where three slopes are noninvariant (.006). We note one seemingly odd finding: that the ΔRMSEA is negative (albeit small) for some conditions. Given that the RMSEA is a function of the ratio of the chi-square statistic to the degrees of freedom for the model under consideration (Bentler, 1995; Hu & Bentler, 1999), this finding can be explained by the fact that this ratio was, on average, smaller for the scalar-invariant models due to a relatively large increase in the degrees of freedom obtained by imposing equality constraints on the intercepts (1975.61/181 = 5.38 compared with 1583.61/136 = 12.12, respectively). And given that the chi-square difference test and ΔCFI are in expected directions, this finding should be interpreted as within the acceptable cutoff and points to conditions under which this measure can be slightly negative.
Discussion
In making comparisons cross-nationally, there is ambiguity regarding whether differences in scale score means can be attributed to authentic differences between countries or to cross-country measurement differences, because of cultural response biases, translation errors, or cultural differences in understanding the underlying construct. Thus, without evidence to support measurement equivalence, any claims or conclusions regarding comparative differences are necessarily weak (Horn, 1991; Vandenberg & Lance, 2000). One means for investigating the comparability of scale scores is through investigation of measurement invariance; however, little is known about the performance of this method and typically used fit measures when the number of groups is large and the sample sizes are large and varied. As such, the current article provides some evidence regarding the performance of MG-CFA in a large-scale assessment context.
To provide some insights into the degree to which typically accepted MG-CFA model-to-data fit measures are applicable in contexts where the number of groups is large and sample sizes are varied, we conducted a simulation study that mimics operational procedures currently used by the OECD, an international organization involved in educational surveys and evaluation. To that end, we generated multiple-group measurement models that drew on an empirical analysis of two TALIS scales to arrive at plausible item- and person-parameters. We examined several factors in this study including two different scale lengths (five and six items), different sources of noninvariance (in slopes, thresholds, and both slopes and thresholds), and two different numbers of groups (10 and 20). To further align our study with authentic conditions, we simulated unidimensional ordered categorical data; however, in line with operational OECD procedures, we applied normal models to these data. We then evaluated several widely accepted measures of model fit in the multiple-group context. We subsequently summarize several interesting findings along with implications for applied researchers and future areas of research.
Before proceeding to a more detailed discussion we draw attention to one important point: We recognize that the disjuncture between the generating models (ordered categorical) and the analytic models (normal) are inconsistent and that this choice is theoretically not best practice. However, given the current state of the practice, this is a purposeful choice and reflects currently used methods by agencies that conduct investigations of measurement invariance, if they are conducted at all.
Under the studied conditions, our results provide some evidence that the chi-square test is likely not useful as a test of overall model fit, as this test strongly suggested model misfit across all studied conditions. Furthermore, the chi-square difference test was also not a suitable measure in this context for the same reasons. These findings are not surprising, given that the studied conditions included both large number of groups, large sample sizes, and we began with a theoretically misspecified model (one that assumes normality when the data are ordered categorical). Despite this result, we do note that the chi-square test statistics (overall and relative) behaved as expected across all conditions in that these values were generally larger for larger numbers of groups and as more constraints were placed on these models. We also acknowledge that our results support earlier findings (Cheung & Rensvold, 2002; Meade et al., 2008), which found that alternative fit indices are often preferable over chi-square tests of model fit in the multiple-group context.
Given the results of the chi-square test statistics, the following discussion is focused on the utility of the overall and relative fit indices in the studied conditions. Of the fit indices considered, our findings suggested that currently accepted cutoffs for the CFI, TLI, and SRMR are generally suitable. In particular, for overall fit evaluations of configural, metric, and scalar invariance, commensurate models fit to data were overwhelmingly supported by CFI greater than .95 and TLI greater than .95. In fact, under every condition studied, neither of these statistics was less than .97 when the model was consistent with the data, providing evidence that a slightly more stringent cutoff would be reasonable in this context. In most cases, an SRMR of less than .08 was also suitable for determining overall fit at different levels of invariance; however, a few exceptions existed. For example, in the 20-group, six-item, fully invariant data condition, a metric-invariant model resulted in an SRMR of .104. Likewise, a few other conditions produced similar results. On first inspection, this might suggest that a more liberal SRMR criterion should be considered (e.g., .10 or .11); however, in several instances, hypothesized models that did not correspond to the data resulted in SRMR values of around .10. As such, we recommend that the SRMR is not used in isolation, if it is used at all. Rather, it should be used in conjunction with the CFI and TLI. And where inconsistencies in these measures arise, the analyst would be better served by relying on the CFI and TLI. In contrast, our findings lead us to recommend a slightly more liberal criterion for the RMSEA, particularly when groups are relatively large. In particular, we found several instances where a model-to-data match did not meet the minimum RMSEA criterion and this finding was more pronounced in the 20-group conditions; however, the RMSEA was highly sensitive to model-to-data mismatches. This leads us to suggest an RMSEA cutoff of around .10 when there are at least 10 groups.
In terms of the relative fit indices (ΔCFI and ΔRMSEA), we note several recommendations based on our findings. Regardless of scale length, we observed a tendency for the ΔRMSEA associated with a hypothesis of metric invariance to increase as the number of groups increased. To that end, in all 20-group conditions and in one 10-group condition ΔRMSEA was at least twice the magnitude of the traditionally accepted value of .010. Furthermore, when a hypothesis of metric invariance was under consideration, the ΔRMSEA was effective at detecting misspecification, with associated values well outside accepted cutoffs. As such, we recommend that a more liberal ΔRMSEA value be used for evaluating metric invariance when large numbers of groups are under consideration. In particular, .030 appears to be a sensible cutoff in these contexts. This finding did not translate to investigations of scalar invariance, where the traditional cutoff of .010 performed well at identifying scalar invariance.
In terms of results for the ΔCFI, we found that the typical value of −.010 suffered some deficiencies in identifying models that were metric invariant, with what also appears to be a slight dependency on the number of groups. Furthermore, when the models were not consistent with a hypothesis of metric invariance, associated ΔCFI values were well outside generally accepted criteria. As such, we suggest that a slightly more liberal criterion of around −.020 be adopted, especially for larger group sizes. Similar to the results for the ΔRMSEA, this finding did not extend to tests for scalar invariance and the status quo is a reasonable and safe choice under the studied context.
In spite of findings that suggest some revised guidelines when considering measurement invariance in larger numbers of groups, any simulation study is subject to limitations. In particular, our analysis only looked at unidimensional scales, in line with OECD operational procedures. Additional research may be needed to examine the performance of these statistics and indices for contexts where scales are typically multidimensional. Furthermore, we only considered two possible group sizes (10 and 20) and given that prominent international educational studies are increasingly developing an interest in examining measurement invariance, it would also be worthwhile to examine this study for even larger numbers of groups. Given the typical number of items comprising noncognitive scales in many international surveys (OECD, 2010a, 2010b), we examined scales of only two lengths: five and six items. In other contexts, longer scales might be more prevalent and further research with longer scales is then likely necessary.
In spite of these limitations, we provided some evidence that particular well-established criteria for determining model fit in the multiple-group case may not be the most suitable under the conditions studied (relatively large numbers of groups, large sample sizes, ordinal data modeled as normal). To that end, slightly more liberal RMSEA measures are recommended as a means for establishing configural and metric invariance but traditionally accepted changes in RMSEA are recommended for determining scalar invariance. In contrast, more stringent CFI and TLI values should be expected to perform equally well at determining overall fit for configural-, metric-, and scalar-invariant models. And similar to the ΔRMSEA, we suggest a more relaxed cutoff for the ΔCFI measures when testing for metric invariance but traditional cutoffs are suitable for determining scalar invariance. We also recommend that the overall SRMR be used in conjunction with other measures or not at all for determining the overall plausibility of configural, metric, and scalar invariance. Finally, we note that our article was concerned only with the detection, not the cause, of noninvariance. As such, we advocate for an approach whereby once measurement nonvariance is detected at some level of the hierarchy, follow-up analyses are necessary to attempt to locate the source of noninvariance in the scale (e.g., in consultation with culture studies experts and/or linguists to examine potential sources of variability).
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors(s) declared receipt of the following financial support for the research, authorship, and/or publication of this article: This work was partially funded by a contract from the Organisation for Economic Co-operation & Development.
