Abstract
Psychometricians have argued that measurement invariance (MI) testing is needed to know if the same psychological constructs are measured in different groups. Data from five experiments allowed that position to be tested. In the first, participants answered questionnaires on belief in free will and either the meaning of life or the meaning of a nonsense concept called “gavagai.” Since the meaning of life and the meaning of gavagai conceptually differ, MI should have been violated when groups were treated like their measurements were identical. MI was severely violated, indicating the questionnaires were interpreted differently. In the second and third experiments, participants were randomized to watch treatment videos explaining figural matrices rules or task-irrelevant control videos. Participants then took intelligence and figural matrices tests. The intervention worked and the experimental group had an additional influence on figural matrix performance in the form of knowing matrix rules, so their performance on the matrices tests violated MI and was anomalously high for their intelligence levels. In both experiments, MI was severely violated. In the fourth and fifth experiments, individuals were exposed to growth mindset interventions that a twin study revealed changed the amount of genetic variance in the target mindset measure without affecting other variables. When comparing treatment and control groups, MI was attainable before but not after treatment. Moreover, the control group showed longitudinal invariance, but the same was untrue for the treatment group. MI testing is likely able to show if the same things are measured in different groups.
Introduction
[I]f the error variances are different then there are either different variables operating on the measure across groups or the same set of variables operate differently across groups. There simply is no alternative explanation. (DeShon, 2004, p. 144)
The principal aim of measurement invariance (MI) testing is assessing whether a given psychological instrument measures the same thing in different groups. This is typically assessed through a process in which models are iteratively fitted until factor configurations, loadings, and intercepts, are equated between groups. 1 The interpretation of equal factor configurations (configural invariance) is that groups reference the same structures of constructs when dealing with the items on a given psychological measure. Equal loadings (metric invariance) are interpreted to mean that groups think of psychological measures the same way when they provide their answers. Equal intercepts (scalar invariance) imply the levels of psychological measures can be considered to represent the levels of a measurement’s target construct to the same degree.
These three steps have been defended as all that is needed to achieve MI (Little et al., 2007). A major component that is frequently disregarded is the residual variances—the equating of which is known as “strict (factorial) invariance.” Strict invariance is generally agreed to be the last stage of MI testing and subsequent constraints fall under what is known as structural invariance testing. Though these concepts are related, they are not the same (Vandenberg & Lance, 2000). If the conditions of MI testing are fulfilled up to scalar invariance, we can compare the means of latent variables and make the statement that performance differences between groups are driven systematically by differences in the levels of those latent variables. If we achieve strict invariance, we can say that a psychological instrument measures the same things between populations because all influences, including ones that are not explicitly modeled, are equated during that step. 2
Steps for Multi-Group Confirmatory Factor Analytic Measurement Invariance Testing.
Note. This table is adapted from Beaujean (Beaujean, 2014, tbl. 4.1). For additional information, see Beaujean (2014, chp. 4).
It is important to recognize that MI, taken as the fulfillment of the three steps prior to testing strict invariance
3
does not allow us to state that we are measuring the same things in different groups. Scalar invariance alone is still consistent with a scenario in which unmeasured variables drive differences in scoring (Lubke et al., 2003), and it can mask differences in specific indicator means, affecting our ability to trust our manifest and—consequently—latent scaling (Lubke & Dolan, 2003). Additionally, MI short of strict invariance is consistent with a scenario in which tests (and factors) are unequally reliable in different groups. Early practitioners of MI testing and its methodological ancestors were aware of this fact and what it implied for measurement (Cronbach, 1947; Meredith, 1993); more recently, this fact has also been explicitly articulated: “Following Cronbach’s (1947) statement on error, if the error variances are different then there are either different variables operating on [a] measure across groups or the same set of variables operate differently across groups. There simply is no alternative explanation…. Given this knowledge, it is completely unclear why it would ever be acceptable to conclude that measures are invariant when the error variances are not homogeneous across groups…. “Error variance is not only a random process, but also the effect of unmodeled sources of systematic variance that affect measured responses. The effect of those unmodeled sources of systematic variance on the evaluation of MI was demonstrated for the single indicator case. In those data, the measurement process is systematically different across groups and yet this difference is only manifested in the existence of error variance heterogeneity. The expected value for the response across groups (e.g., intercept and slope) was invariant…. [M]ulti-indicator models do not reduce the need to demonstrate error variance homogeneity when assessing the functioning of a measure across groups.” (DeShon, 2004, pp.‘s 144, 147–148).
While it is incorrect to say that invariance of the loadings and intercepts of a measurement model indicates that the same thing(s) is (are) necessarily being measured in different groups, their invariance does indicate that some of the same thing(s) is (are) being measured. Such a finding does not make strong guarantees, however. On its own, the mere achievement of this level of MI ignores potential differences in causes specific to one or another group. To the extent reliability differs, so too must the causes of responses and what is measured in different groups, regardless of the random or systematic nature of the residual variances. Consider Cronbach’s (1947) statement that “All methods of studying reliability make a somewhat fallacious division of variables into ‘real variables’ and ‘error.’ [….] A test score is made up of all of these ‘real’ elements, each of which could be perfectly predicted if our knowledge were adequate. Reliability, according to this conception, becomes a measure of our ignorance of the real factors underlying fluctuations of behavior and atypical acts.” (ibid., pp. 6–7). If we had perfect knowledge, we would know exactly what causes “error”; if we know we have strict invariance, then when it comes to group comparisons, we don’t need to.
MI is clearly a very important concept for testing theories and performing psychological comparisons in general. This importance has been recognized in numerous papers (Boer et al., 2018; Jeong & Lee, 2019; Lacko et al., 2022; Lasker et al., 2022; Wicherts, 2016), and yet confusion about (Welzel et al., 2021; cf. Meuleman et al., 2022) and disregard of it (Maassen et al., 2023) remain the norm. This may reflect a lack of appreciation for what MI, or its violation, can reveal to researchers. As Maassen et al. (2023) argued, MI testing should be “an opportunity for researchers to gain a clearer understanding of the substantive and practical meaning of the differences they observe” (ibid., p. 12).
Several recent papers have provided data from situations that have made it possible to empirically test whether MI represents measuring the same things in different groups (Burgoyne et al., 2020; Protzko, 2022 4 ; Schneider et al., 2020). Before directing our attention to them, several things to keep in mind when testing MI must be noted.
The interpretation of MI without strict invariance is not necessarily one of common measurement. It is, at best, one of partially common measurement and is thus not a theoretically applicable test of whether MI involves common measurement. Second, many model fit indices in common use are often insufficient to detect violations of any level of MI. Protzko (2022), for example, exclusively used two approximate fit indices, RMSEA and CFI. These are less powerful for testing omnibus MI violations than χ2 tests or the use of information-theoretic indices like the Akaike and Schwarz’ information criteria (AIC and BIC, respectively). However, they are clearly much more common (Ropovik, 2015, see, esp., the sections “Approximate Fit Indices” and “The Consequences of Disregarding the Model Test”), perhaps because of their weaknesses. Namely, these model fit indices may allow researchers to “null hack” (Protzko, 2018), avoiding finding evidence of misspecification or other group differences in model parameters. Additionally, many purported tests of MI feature liberal fit criteria (see Svetina et al., 2020, Table 1). This is not to say that χ2 tests and indices like AIC and BIC are perfect, but they are known to be more sensitive means of detecting misfit.
Researchers should exercise considerable caution when using approximate fit indices because they are often insensitive, and their cutoffs are usually arbitrary. 5 This should not be taken as a complete condemnation of their use or a statement about their inherent inferiority compared to other tests. χ2 tests and information-theoretic indices have their own downside, in that they are more likely to highlight noninvariance at large sample sizes when there is practically none, forcing MI testing to be a matter of degree.
Another major issue in MI testing is theoretical: to understand when and why MI is expected to hold or fail, theoretical knowledge about our measurements and samples is required. Individuals from a culture valuing toughness may never supply accurate answers to questions about pain. Men may misreport their heights and weights to a greater degree and in a different direction from women because it’s socially desirable for them to be tall and muscular. An English second language learner may fail to show their abilities on a test given in English due less to ability and more to lacking familiarity with the language. Each of these things can be predicted theoretically and verified by testing for MI, but the reasons may not be simple.
An example of a complex theoretical background comes from Protzko (2022). The paper’s explicit aim was to test whether the same constructs were being measured in different groups. This was done by randomly assigning a population-representative sample of 1,500 American adults to answering a questionnaire on the meaning of life or an otherwise-identical one on the meaning of an undefined nonsense concept known as gavagai, and then testing MI with respect to the items of those questionnaires despite them being different. For the purposes of testing whether a “nonsense scale” like the meaning of gavagai scale could have predictive validity, the sample was also administered a questionnaire about their beliefs in free will.
Whatever the “meaning of gavagai” questionnaire measured, it might have been the same, or similar enough, to the construct underlying the “meaning of life” questionnaire. Because of their similarities (the questionnaires use almost the same items excepting the replacement of the words “gavagai” and “life”; the question wording can be found in Table S2), we should expect some comparability in measurement, while their differences ought to elicit widespread measurement noninvariance to the extent the constructs measured truly differ. Based on how much similarity there is in the influences on these tests, we can hypothesize accordingly about the identity of the residual variances. If the responses on the questionnaires were influenced by different things because people consistently interpreted the meaning of life in a way that differed from the meaning of gavagai—despite its possible consistency within individuals—then it should be the case that the residual variances are different.
Finally, the hypothesis of common measurement often has less to do with loadings and intercepts, and more to do with residual variances (i.e., strict factorial invariance). If an investigator’s aim is to see if tests measure the same things, one can never forget to equate all the measurements that have to do with influences—or, in other words, the representations of “the same things.” The residual or error variances are latent variables, and they do represent both systematic and plausibly random causes which could be explored. If we aim to test the hypothesis that MI means the same things are being measured in different groups, we must test the hypothesis suitably, by attempting to equate the residual variances. Since the concept of something like “gavagai” is different from the concept of “life,” its meaning and related responses could plausibly be considered different; if we do not measure properly, we will fail to note that, and if we assume their differences are only in common latent variables rather than the potentially uncommon residual variances, we have lost sight of the objective.
The Present Study
This analysis includes three studies.
The first study is based on the assertion that Protzko (2022) constituted an incorrect test of the important claim that MI indicates that a psychological instrument measures the same thing in different groups. Protzko (2022) utilized insensitive fit indices, in effect ignoring violations of MI; the construct identity and validity of the “meaning of gavagai” questionnaire was improperly considered, and finally; a requisite test of MI—assessing strict factorial invariance—was not conducted.
To address the paper’s methodological problems, I utilized the data from Protzko (2022) to suitably assess MI. The results of this reanalysis are presented didactically, to illustrate the errors in the order they should have been observed and to lay the groundwork for four other analyses of the two other studies.
The second study was based on data from Schneider et al. (2020). These researchers performed two experiments in which they taught participants the rules of figural matrices and then assessed the effects on their test scores. In Experiment 1, 112 German university students were randomized to either a treatment or a control group. Individuals in the treatment group were shown a video in which six rules that were relevant to the subsequent figural matrix tests were explained to them, and they were also given text-based and graphical, step-by-step instructions. The control group was given instructions that were similar, but the content they received information on was about diet and nutrition and was thus not test-relevant. Participants then took four subtests of the Leistungsprüfsystem 2 (LPS-2) involving general knowledge, marking incorrect numbers in numerical sequences, quickly adding numbers, and mental rotation, respectively. They followed this up by taking two variants of the DESIGMA-Advanced figural matrices tests. In Experiment 2, 229 German university students underwent a similar procedure, with the same intelligence test minus its addition subtest, and with only a single variant of DESIGMA-Advanced. Additional information and summary statistics can be found in Schneider et al.’s (2020) study.
In Schneider et al.’s (2020) intervention, instruction was only provided for figural matrices. Since the knowledge required to perform well on matrix tests is not like the knowledge required to do well on the other LPS-2 subtests, there should only have been an effect on matrix test performance, and not on the other subtests. There should also have been bias when comparing the treatment and control groups since the control group lacked the novel influence of explicitly test-relevant information.
The third study was built around the data from Burgoyne et al. (2020). These researchers recruited twins from the Michigan State University Twin Registry study, the Twin Study of Behavioral and Emotional Development in Children, and the Michigan Twins Project registry. They obtained zygosity and demographic information, and measured growth mindset, grit, locus of control, cognitive ability, and a handful of other measures, before randomizing pairs of participants to either a growth mindset intervention or an active control. The growth mindset intervention involved presenting participants with content suggesting that “the brain is like a muscle—it gets stronger (and smarter) when you exercise it,” whereas the active control intervention involved telling participants about the human brain with less aspirational and more factually-focused content like “the parietal lobe is where the brain interprets the sense of touch” (ibid., pp. 5–6).
The effect of Burgoyne et al.’s (2020) intervention was to augment the level of participant growth mindset without seeming to affect their other measures. More importantly, these authors used their twins to fit behavior genetic ACE models to estimate the additive genetic (A), shared environmental (C), and nonshared environmental (E) variance in their mindset measures before and after their intervention, by group (Plomin et al., 2008). Their intervention increased the A variance in their measure of growth mindset in the intervention group, while at the same time, not affecting the A, C, or E variance in a self-determination composite score that they defined in their earlier pilot study (Burgoyne et al., 2018) as a score for a higher order construct called “self-determination,” composed of responses for growth mindset, grit, and locus of control.
Burgoyne et al.’s (2020) finding is perfect for testing whether MI means the same things are being measured in different groups. This is because they found that responding with their mindset measure involved different influences in the experimental and treatment groups: in the experimental group post-intervention, A was more important as a determinant for growth mindset responding. This means that the change in the level of growth mindset likely cannot be interpreted to be due to changes in growth mindset itself. If MI means the same things are measured in different groups, it is not possible for a latent variable whose indicators have different ACE variances to display MI in realistic scenarios. MI will be unattainable unless A, C, and E all have qualitatively indistinguishable effects on responding and the intervention reduced one variance component as much as it increased another one. The first possibility is unlikely. In what world would a genetic effect strongly resemble (when power is moderate) or be identical to (when power is high) the effect of the shared environment? One potential answer is a world in which an A-, C-, and E-influenced common pathway model for a latent variable fully explains the variance in its indicators with and without an intervention (Franić et al., 2013). That does not seem to be our world, as the intervention only affected the mindset measure and not the self-determination composite. Since that is not our world, we know that the intervention must have at least had its effect on the intercept or the residual variance 6 in the mindset indicator, if not both parameters and more. Since the A and C variances definitionally cannot resemble the E variance because it is unsystematic by design, it is irrelevant. Therefore, in scenarios where an AE model (i.e., additive genetics and nonshared environments without shared environmental influence) is strongly confirmed and an intervention affects one or both of those variances, MI must necessarily be violated.
There are certainly more possibilities with biometric variance changes that can be imagined which might be compatible with MI, but they will generally be unfalsifiable, as they’ll involve theorizing about quantitatively distinguishable levels of influences that act on phenotypes through different mechanisms yet are still qualitatively indistinguishable in their manifestations.
These three scenarios provided by Protzko (2022), Schneider et al. (2020), and Burgoyne et al. (2018, 2020) are ones in which, first, a different thing is known to be measured because qualitatively different test instruments were used in different groups; second, a highly-specific effect on a cognitive performance measure was elicited based on a novel source of response variance that is independent from general intelligence; and finally, the variance components involved in just one of a latent variable’s indicators were confirmed to have been altered by an intervention.
Analysis and Results
Section 1. Protzko (2022)
The first step in fitting this data is to fit the structural equation model, and to fit it thrice: with the whole sample, with one group, and with another group. The model used for these analyses is illustrated in Figure 1. Table 2 contains the loadings for both groups, the one given the meaning of life questionnaire and the one given the meaning of gavagai questionnaire. The differences in their unstandardized loadings
7
were computed and a p-value for that difference is provided in the table’s last column. For the free will belief factor, the loadings did not significantly differ (p’s between .109 and .999); however, for the meaning of life/gavagai factor, four of the loadings significantly differed (p’s between 1.7e-7 and .262). This means that metric invariance must fail and model fit indices that fail to detect these differences are failing to capture genuine misspecification. When loadings significantly differ, metric invariance fails; when they do not significantly differ, it will tend to be tenable. The fact that metric invariance must fail does not mean that the difference in loadings is necessarily large, and an approximate fit index should still be expected to change with much more considerable changes. But this is beside the point: clearly, the meaning of life and meaning of gavagai questionnaires must be interpreted differently, which is consistent with them measuring different things. As a result, ω cannot be the same unless the loadings were coincidentally higher and lower to equal degrees across measures, so if that is interpreted as a measure of reliability it will likely indicate measurement that is not in common. The free will belief—meaning of life/gavagai correlated factor diagram. Baseline Model Factor Loadings and Whether They Significantly Differ. *Values of .999 indicate rounding. The configural model was identified by constraining the latent variance to avoid mistakenly constraining noninvariant loadings to equality (Johnson et al., 2009). All results were robust to the use of different identification methods where appropriate. The choice of ML, MLR, or DWLS as the estimator did not qualitatively affect results and loadings were indistinguishable when factors were modeled separately at baseline.
Measurement Invariance, Without Partial Models.
Note. Bolded values indicate evidence of violations of MI and bolded levels indicate a failure at that level with respect to the previous one. For the column “X 2 /df”, this means that the difference between a model and the one it is nested under was significant.
Measurement Invariance, With Partial Models.
Note. Bolded values indicate evidence of violations of MI and bolded levels indicate a failure at that level with respect to the previous one. For the column “X 2 /df,” this means that the difference between a model and the one it is nested under was significant.
Steps of Partial Strict Model Fitting.
Note. Values are X2 for the model if that parameter was freed. Bolded values indicate a parameter was freed which, accompanied by another stage of parameter inspection, meant that the model continued to fit significantly worse than the partial scalar model. Lower X2 values indicate better fit, as the model is closer to nonsignificantly differing from the previous one.
Section II. Schneider et al. (2020, I)
Measurement Invariance, With Partial Models.
Note. Bolded values indicate evidence of violations of MI and bolded levels indicate a failure at that level with respect to the previous one. For the column “X 2 /df,” this means that the difference between a model and the one it is nested under was significant.
In terms of magnitude, the experimental group scored 1.30 g higher on one figural matrices indicator and 1.15 g higher on the other. Using the bias effect size introduced by Lasker and McNaughtan (2022), the amount of this difference between experimental groups that was attributable to bias could be computed. This effect size is a simple extension of the effect size SDI2 introduced by Gunn et al. (2020) and it can be interpreted in terms of Glass’ delta. Briefly, SDI2 is a signed effect size for measuring the impact of noninvariance in continuous outcomes in the framework of an MGCFA, giving by
These effect sizes can be combined to produce an unsigned effect size with the correct magnitude, in the form of SUDI2, like so:
For further information about SDI2 and UDI2, see Gunn et al. (2020). To calculate the proportion of an observed gap due to bias with this metric, SUDI2 can be subtracted from the observed gap if the order of the groups for both effect sizes was consistent.
When SUDI2 is computed from this model of Schneider et al.’s (2020) data, the greater intercept in the experimental group appears to raise that group’s observed scores for either indicator by 1.27 g and 1.17 g, respectively. Although there is some imprecision in these estimates, it is very likely that the entirety of the difference between the randomized treatment and control groups in figural matrices performance was due to psychometric bias and noise.
Section III. Schneider et al. (2020, II)
Measurement Invariance, With Partial Model.
Note. Bolded values indicate evidence of violations of MI and bolded levels indicate a failure at that level with respect to the previous one. For the column “X2/df,” this means that the difference between a model and the one it is nested under was significant.
Section IV. Burgoyne et al. (2020)
The three-indicator self-determination model suggested in Burgoyne et al. (2018) fit their 2020 data perfectly. Because they provided pre- and post-intervention measurements for the treatment and control groups, there were three readily-available means of testing for MI. First, testing the MI of self-determination between the control and treatment groups in the pre- and post-intervention timepoints acts as a test of whether the treatment effect results in any amount of bias. Second, testing the MI of self-determination between the control and treatment groups with pre- and post-measurements modeled simultaneously allows for testing whether MI is violated with stable influences controlled. This allows for more powerful estimation of a heterogeneous treatment effect through controlling away stable influences. Finally, testing for longitudinal invariance within both groups lets us know whether the indicators meant the same things within each group, between measurement occasions. If the treatment effect is biasing, the treatment group should see mindset biased between measurement occasions. If Burgoyne et al. (2020) could maintain their sample’s attention between measurements, then the control group should not show any such bias, and if there is bias, it should be as likely to show up for any indicator as for the mindset indicator. These models were also used in Section V and the diagonally-weighted least squares estimator was used for both studies.
Measurement Invariance, Prior to Intervention.
Note. Bolded values indicate evidence of violations of MI and bolded levels indicate a failure at that level with respect to the previous one. For the column “X 2 /df,” this means that the difference between a model and the one it is nested under was significant.
Section V. Burgoyne et al. (2018)
Measurement Invariance, Prior to Intervention.
Note. Bolded values indicate evidence of violations of MI and bolded levels indicate a failure at that level with respect to the previous one. For the column “X 2 /df,” this means that the difference between a model and the one it is nested under was significant.
Conclusion
MI testing is capable of correctly showing when different things are measured in different groups.
In a reassessment of Protzko (2022), the meaning of life questionnaire was not interpreted the same way as the meaning of gavagai questionnaire. Moreover, the influences on both questionnaires were clearly not wholly shared. The reasons these results are discrepant from those reported with the same data by Protzko (2022) are straightforward. First, Protzko (2022) used insensitive model fit indices that failed to indicate measurement noninvariance where it was tested. Some of this noninvariance had to be apparent on configural model inspection because it was impossible to achieve a different result due to the significance of the loading differences. That it did not appear as indicated with liberal CFI and RMSEA cutoffs indicts those approximate fit indices more than it supports metric or scalar invariance.
Second, Protzko (2022) never fully tested whether the meaning of gavagai/life questionnaires measured the same things in the first place, because strict invariance was not tested. This model must be tested to assess the claim that the same things are being measured because the residual variances represent the influence of things, and the equality of construct reliability is also left unknown. The metric and scalar models test the equivalence of the interpretation of certain model parameters and allow researchers to claim that the latent construct of interest is measured at least partially in common while usually allowing the means of latent variables to be legitimately compared, but these steps do not provide information about the reliability of constructs in different groups, nor do they ensure equal reliability of the indicators, or that the influences on scoring are fully shared. Accordingly, in many scenarios, the proposition that the same thing is measured in different groups is likely to be left untested if the strict model isn’t fitted.
The interpretation of the results of Protzko (2022) as evidence against one of the central functions of MI testing was always spurious. What the experiment would indicate in the negative (i.e., finding MI holds) is that there is a construct validity problem for the meaning of life and gavagai questionnaires because, whatever they are, they would be measuring the same things! In the affirmative (i.e., finding MI does not hold, whatever the reason), it affirms that there are differences in the involved constructs and that the procedure is able to pick up on them. Theoretically, this test was never able to affirm the claims made in Protzko (2022), only to repudiate them.
To confirm what that study wished to is simple enough. One appropriate design would be to provide the same questionnaire to two groups and bias the results in a clear and visible way, such as by giving out answers to or letting people practice for an objective test instrument, or encouraging one group to answer a certain way on a personality questionnaire while the other is left to answer normally. This has been done in the case of stereotype threat, where a sample was primed to feel threat that reduced their performance on an intelligence test. However, MI testing was also sensitive in that case and it was observed that stereotype threat led to biased measurement (i.e., noninvariance; Wicherts et al., 2005). In this paper’s testing with Schneider et al.’s (2020) figural matrix rule-teaching experiment, it was found that bias was observed for the correct parameters and in practically the exact quantities expected. In a further scenario involving a growth mindset intervention, Burgoyne et al.’s (2018, 2020) datasets showed that growth mindset was biased, as one would expect given their biometric results.
Maassen et al. (2023) reviewed the state of MI testing in the journals Psychological Science and PLOS ONE. Out of thirteen experimental within-group comparisons, eight were simply noninvariant and five showed supported for scalar variance. There were a further 128 experimental between-group comparisons, of which 67 were simply noninvariant, twelve only reached configural invariance, thirteen reached metric invariance, and 36 reached scalar invariance. In their codebook, forty MI tests were listed as being empirical, with experimental groups, where the groups reached scalar invariance. The characteristics of these studies are worth noting.
Van Dessel and De Houwer (2019) had three comparisons that reached scalar MI: a comparison over time of people’s stimulus evaluations, which very significantly changed over time, and two comparisons involving evaluations in hypnosis and relaxation conditions. The lack of any apparent difference in measurement parameters between these two conditions is not very remarkable because there was only one statistically significant difference in evaluations between the conditions at the different timepoints, and it was marginal (two-tailed p = .018). Or, in other words, the choice of condition had no apparent effect, so it is unsurprising that measurement wasn’t compromised. Luttrell et al. (2019) had a similar result: measurement of persuasiveness were robust to nonsignificant or extremely marginal interactions on the basis of the type of appeal individuals were exposed to. The results of Zlatev (2019) were quite different. In that study, integrity ratings were comparable between conditions where the target was shown as high or low on caring and when participants and targets agreed or disagreed. There were significant differences in integrity ratings by condition, but the meaning of ratings was MI. A handful of other studies reached scalar MI (Ackerman et al., 2018; Berman et al., 2018; Catapano et al., 2019; Kardas & O’Brien, 2018; Moon et al., 2018; O’Brien & Kassirer, 2019; O’Connor & Cheema, 2018; Sawaoka & Monin, 2018; Srna et al., 2018); their common characteristics seemed to be indicating the robustness of perceptions, ratings, etc. to conditions and MI failing to be violated by effects that were any of nonsignificant, marginal, or otherwise very small.
None of those studies showed MI changes in cognitive capabilities, nor a single instance like in Protzko (2022) or Burgoyne et al. (2018, 2020) where MI is implied or practically implied to be violated and it nevertheless appeared to be confirmed. Those results are reassuring about the possibility of doing psychological research involving attitudes and perceptual ratings without bias in measurement, but they otherwise reveal very little about the theoretical import of MI. Nonetheless, the more revealing finding from Maassen et al. (2023) was the number of studies where MI was not satisfied. As they showed, even when there’s no obvious way that an experiment should cause MI to be violated, it can still be violated, so we cannot take MI for granted and it must always be tested if instruments are going to be interpreted in a way that requires they be unbiased.
At least two other trials have supported the possibility that interventions or formatting differences are detectable through testing MI, one clearly, and one potentially. The clear demonstration came from Arendasy and Sommer (2013). They used a higher order model with crystallized, fluid, and quantitative group factors and showed that variations in the rules involved in figural matrices led to violations of MI in the loadings, intercepts, and residual variances of their figural matrices tasks. Becker et al. (2016) followed similar procedures, but their results were ambiguous because they used conceptually nonsensical models where intelligence, matrices performance, and working memory were separate and correlated group factors, and they ultimately could not all be modeled alongside one another due to convergence problems and they did not adequately test strict invariance in their more limited models. The reanalysis of this data with a model like Arendasy and Sommer’s (2013) would serve as a further test of whether MI means the same things are measured in different groups.
A potential broader implication of these findings is that because interventions aimed at changing traits like growth mindset, spatial ability, or other traits of interest to psychologists usually involve novel sources of potential trait variance, when they cause changes, those changes will tend to be attributable to the resulting noninvariance. As a result, intervention-induced changes in those traits are unlikely to be mediators of the potential beneficial effects of those interventions. In other words, if a causally efficacious trait (e.g., non-cognitive skill, conscientiousness, and grit) predicts success in life and individuals exposed to a trait-boosting intervention become more successful, if the change in the trait was due to psychometric bias, it’s unlikely that change is why the intervention caused people to become more successful. Changes in the variable may be indicate the effects of the intervention, but they are unlikely to have the same implications as cross-sectional variance in that trait and we cannot necessarily generalize the effects of the trait observed in the cross-section to those that follow the trial.
The results provided here make for a clear conclusion: a well-powered finding of MI allows psychometricians to claim that the same constructs are measured in different groups.
Supplemental Material
Supplemental Material - Measurement Invariance Testing Works
Supplemental Material for Measurement Invariance Testing Works by Jordan Lasker in Applied Psychological Measurement
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Data Availability
All code required to replicate the initial reanalyses is available online at https://rpubs.com/JLLJ/PMI22. The same link provides code for a supplementary analysis of randomly generated Likert data that further illustrates issues with commonly used fit indices and the level of testing in this study’s referent. Code for RMSEAD is at https://rpubs.com/JLLJ/SBFNMC. Code for the reanalysis of Schneider et al.’s (2020) results is located at https://rpubs.com/JLLJ/MatMI. Code for the reanalysis of Burgoyne et al.’s (2018; 2020) results is available at
.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
