Abstract
Special measurement effects including the method and testlet effects are common issues in educational and psychological measurement. They are typically covered by various bifactor models or models for the multiple traits multiple methods (MTMM) structure for continuous data and by various testlet effect models for categorical data. However, existing models have some limitations in accommodating different type of effects. With slight modification, the generalized partially confirmatory factor analysis (GPCFA) framework can flexibly accommodate special effects for continuous and categorical cases with added benefits. Various bifactor, MTMM and testlet effect models can be linked to different variants of the revised GPCFA model. Compared to existing approaches, GPCFA offers multidimensionality for both the general and effect factors (or traits) and can address local dependence, mixed-type formats, and missingness jointly. Moreover, the partially confirmatory approach allows for regularization of the loading patterns, resulting in a simpler structure in both the general and special parts. We also provide a subroutine to compute the equivalent effect size. Simulation studies and real-data examples are used to demonstrate the performance and usefulness of the proposed approach under different situations.
Keywords
Introduction
Special measurement effects including testlet and method effects due to different testlet formats, wording, rater characteristic, and so on are common in educational and psychological measurement. Broadly speaking, these special effects can be divided into two groups, each with various psychometric models designed to handle them. Special effects for continuous data are usually analyzed from the perspective of factor analysis (FA, Brown, 2015) with various models for bifactor (Reise et al., 2010) and multi-trait multi-method (MTMM; Campbell & Fiske, 1959) structures available. In contrast, special effects for categorical data are typically addressed from the perspective of item response theory (IRT) with various testlet effect models (Bradlow et al., 1999) available. The two perspectives are usually discussed separately in the literature, although they are statistically communicable by considering uni- or multidimensional IRT as categorical FA (Takane & De Leeuw, 1987).
Recently, a great deal of modeling developments can be found from both perspectives. For special effects with continuous data, various new models for bifactor and MTMM structures have been proposed (e.g., Esra & Atar, 2021; Geiser & Simmons, 2021; Wang et al., 2018). Note that both bifactor and MTMM-type models are usually adopted under the context of method effects. Standard-type bifactor models are confirmatory in nature, with one general factor, multiple orthogonal special factors, and a fully specified loading pattern (e.g., Brown, 2015; Chen et al., 2006; Holzinger & Swineford, 1937). When the loading pattern is partially specified, it becomes an exploratory bifactor model which has enjoyed a rediscovery more recently (e.g., Giordano & Waller, 2020; Jennrich & Bentler, 2011; Reise, 2012; Reise et al., 2010). Although several methods are available, the preferred approach to exploratory bifactor analysis is the classic or derived Schmid-Leiman (SL; Schmid & Leiman, 1957) methods, which are subject to the proportionality constraints and hierarchical bifactor structure (Giordano & Waller, 2020). Taking into account that the general factor is fully specified (e.g., usually measured by all items), these models are partially rather than fully exploratory. Moreover, there is only one general factor and local independence is assumed in all existing bifactor models, which could further limit the applications of this type of models. Compared to bifactor models, models for MTMM are confirmatory only, but have more complex structures, which can be used to model the relationships among multiple traits (e.g., general factors) and methods (e.g., special factors). In its most general form, MTMM structure allows for multiple general factors, correlated special factors, and correlated uniqueness or measurement errors (i.e., local dependence). However, MTMM-type models are only identified under certain conditions. Eid et al. (2006) have detailed the structure and restriction of several identified MTMM-type models, including the correlated trait correlated uniqueness (CTCU) model, correlated trait uncorrelated method (CTUM) model, and correlated trait correlated method (CTCM) model. These three models are most widely used in the MTMM context (e.g., Brown, 2015; Kyriazos, 2018; Marsh & Byrne, 1993). Specifically, CTCU allows for local dependence. CTUM can be considered as an extension of the standard bifactor model with multiple general factors. CTCM can be problematic, especially when the number of special factors is small, and the correlated trait correlated method model with one method less (CTC(M-1)) is usually adopted instead. However, all these models are confirmatory by nature, making them inappropriate for any exploratory settings.
For special effects with categorical data, testlet effect models under the context of IRT are often employed to model the effect of a common stimulus within a bundle of items. The Bayesian random effect model on dichotomous responses proposed by Bradlow et al. (1999) is an early prototype of a testlet effect model. Wang et al. (2002) further extended the random effects model to polytomous responses and released the restriction of constant testlet effects. After that, Wang and Wilson (2005) proposed the Rasch testlet effect model as a special case of the multidimensional Rasch model with the ability to deal with both dichotomous and polytomous responses. Based on Rasch testlet effect model, some multidimensional testlet effect models have been proposed recently (Zhan et al., 2014, 2015). In general, testlet effect models are confirmatory by nature, with only one general factor (i.e., trait) measured by all items and the assumption of local independence.
For the bifactor and MTMM-type models, the maximum likelihood estimation (MLE) implemented through the expectation-maximization algorithm is widely used. In comparison, the Bayesian approach is often employed in testlet effect models, which has several benefits (Fukuhara & Kamata, 2011; Wainer et al., 2007) including its flexibility and scalability to complex scenarios. Recently, a generalized partially confirmatory factor analysis (GPCFA) framework with mixed Bayesian Lasso methods was introduced (Chen, 2021). This Bayesian-based approach provides a regularized and flexible framework for factor analysis, with clear advantages such as making the exploratory and confirmatory approaches two ends of a continuum, handling both continuous and categorical data, and regularizing the complex loading structure and local dependence simultaneously.
In this research, we will modify the GPCFA framework to accommodate special effects for continuous and categorical data with an additional method to measure the effect size. As a result, the revised GPCFA can cover and extend many traditional models for special effects with added benefits. Specifically, bifactor models can be extended with multiple general factors and local dependence, and MTMM-type models can be extended to address partially exploratory settings with partially specified loading patterns, while extended testlet effect models can enjoy all above benefits. It will also be easier to explore and compare different models for special effects within a unified framework and effect size measure. Moreover, all models will automatically inherit GPCFA’s other benefits such as the partially confirmatory design and addressing mixed-type variables with missingness and local dependence.
While building on the GPCFA framework, this research introduces several new contributions. First, we extend the GPCFA framework by dividing the factor structure into the general and special factors to accommodate various special effects. Second, we clarify how the revised GPCFA can incorporate or be connected with different bifactor, MTMM, or testlet structures, providing not only a unified framework to understand different models for special effects but also a chance to specify novel models for potential applications. For example, one can specify new bifactor, MTMM or testlet effect models that can handle data mixed with continuous, dichotomous, and polytomous responses and missingness, with or without regularization or local dependence. Third, we present a method to homogeneously measure and compare effect sizes across various settings. Fourth, we demonstrate the effectiveness of the revised framework and effect size measure through two simulation studies for continuous and categorical data, with illustrations of two real-life examples for applied researchers.
Theoretical Framework
GPCFA for special effects
The GPCFA is a factor analytic model with regularization of the loading pattern and local dependence, which makes it in between the exploratory and confirmatory ends (Chen, 2021). For complete description of this partially confirmatory methodology and related algorithms, readers can refer to the GPCFA literature (e.g., Chen, 2020, 2021; Chen et al., 2021).
Assume there are
Loading estimation structure is conducted through different settings of the design matrix
To accommodate the GPCFA framework to deal with the method and testlet effects, several adjustments of the structure are needed. The factors are separated into two parts: the construct part (general factors) and the special effects part (special factors). We will use subscript “g” and “s” to denote the general and special factors and related loadings, respectively. Here, we use the following notations: total
In practice, the general factor usually refers to latent construct or trait and the special factor usually derived from the study design, test format, or measurement method. Figure 1 illustrates an example of the revised GPCFA model with 9 items and two general and two special factors. In this example, one loading per item is specified in both general and special factors, while the other loadings in the general factors are unspecified, and those in the special factors are set to zero. Additionally, the general factors are correlated while the special factors are orthogonal. Meanwhile, the items can be continuous or categorical and local independence is assumed. In general, one can specify the Q-matrix, factor structure (correlated or orthogonal), local dependence, and type of each item to accommodate various bifactor, MTMM, or testlet structures. An example of revised GPCFA.
The matrix equations and corresponding
For model identification, two constraints are necessary: setting all factorial variances to one, which constrains
Accommodating Bifactor Models
The revised GPCFA framework can accommodate various bifactor models. Specifically, a standard bifactor model with one general factor can be obtained by setting Standard bifactor model.
The model becomes partially confirmatory when the loading matrix for the special factors are partially unspecified, which is similar to an exploratory bifactor model allowing cross-loadings on special factors. It is worth noting that the SL-type exploratory bifactor models are subject to the proportionality constraint due to the SL-transformation procedure (Yung et al., 1999). This constraint requires a linear combination between general factor loadings and special factor loadings. Considering that the general factor is fully specified (i.e., no exploration needed), these models are partially rather than fully exploratory.
In comparison, the revised GPCFA is subject to the identification constraint of a few specified loadings per factor (under the assumption of local independence), rather than the SL constraint. But one can evaluate the constraint post hoc. Under the GPCFA framework, the standard bifactor model can be extended with multiple general factors which can be correlated. In addition, both the loading matrices for the general and special factors can be partially unspecified, making the extended bifactor model exploratory-inclined. Compared to exploratory bifactor models, the revised GPCFA offers more flexibility and scalability by covering multiple general factors, mixed-type data, missingness, and local dependence.
Accommodating MTMM structures
The separation of general and special factors in the revised GPCFA framework allows us to accommodate several MTMM structures with four examples shown in Figure 3. Structures of MTMM Related Models: (a) Correlated Trait Correlated Uniqueness model; (b) Correlated Trait Uncorrelated Method model; (c) Correlated Trait Correlated Method model; (d) Correlated Trait Correlated Method model with One method factor less than methods considered.
The correlated trait correlated uniqueness (CTCU) model can be directly estimated by the original GPCFA model by considering different traits as different factors with local dependence. The observed variable
The residual in CTCU can be correlated between items within the same method effects. Namely, the correlated residual is equivalent to the special factor substantively, both of which refer to the method effect. However, the effect size is incalculable unless one can summarize multiple pairs of correlated residuals within the method reasonably.
The correlated trait uncorrelated method (CTUM) is equivalent to the extended bifactor model, with multiple general factors that are correlated and multiple special factors that are uncorrelated:
Correlated trait correlated method (CTCM) and CTC(M-1) are more generalized than CTUM by allowing all special factors to be correlated. In CTCM, the covariance matrix of special factors,
Accommodating Testlet Effect Models
This section introduces the transformation of GPCFA to accommodate several testlet effect models under the IRT context which are mainly for dichotomous responses. With the local independence assumption in IRT, the residual covariance matrix
It leads to the item response function with the normal ogive model, namely, using the cumulative function of the standard normal distribution
Alternatively, one can make model identified by setting the
The
Different testlet effect models can be obtained as follows: The general testlet effect model (Li et al., 2006) can be transformed from GPCFA by restricting the number of general factors to one and keeping the residual covariance matrix
The two-parameter normal ogive (2PNO) testlet effect model proposed by Bradlow et al. (1999) can be written for each latent response
The transformation between GPCFA and the extended 2PNO testlet effect model is:
The constant
In the extended 2PNO testlet effect model, if we restrict the discrimination parameters
In addition, GPCFA can also estimate the extended Rasch testlet effect model, like within-item multidimensional testlet effect model (Zhan et al., 2014, 2015) by releasing the constraint of special factor allocation (i.e., adding the number of special factors). Moreover, polytomous responses or mixed-type formats, with missingness and local dependence, can be readily addressed within the GPCFA framework. However, the GPCFA cannot be transformed into the three-parameter testlet effect model with guessing parameter yet.
Equivalent Effect Size
Typically, the effect size refers to the amount of variance due to the random effect when the general factor or trait is standardized (Wang & Wilson, 2005). Since all factors are standardized under GPCFA, we define the equivalent effect size measure D, the average of the square loadings on the random effect (i.e., special factor). Moreover, due to the partially confirmatory setting, we separate the random effect into two parts. For specified loadings in the loading vector of the special factor, all loading estimates will be included. For unspecified loadings, we only consider those loading estimates with absolute values greater than the cutoff of 0.1, which is typically used under the Lasso scenarios. We denote the
Simulation Studies
Two simulation studies were conducted to evaluate the performance of the proposed model on continuous and categorical data. Based on previous research (Chen, 2021) and real-life examples, a sample size of N = 1000 was used, and two effect sizes (D = 0.1 and 0.2) were simulated. Different conditions of local dependence and other settings were manipulated as shown in the following studies. The true diagram for two simulation studies is given in Appendix C. For each condition cell, 200 replications were simulated and analyzed. The performance assessment for each parameter includes the bias (BIAS), the mean of the standard error (SE), the root mean square error (RMSE) between the estimates and the true values, and the percentage of estimates that differed significantly from zero (
To stabilize most Markov chains (i.e., the estimated potential scale reduction (EPSR) <1.1 (Gelman et al., 2014)), 20,000 iterations of burn-in draws were performed, followed by additional 20,000 iterations. All studies used the LAWBL package (Chen, 2022) in the R (R Development Core Team, 2021) computing environment. The sim_lvm and pcfa functions in the package were used to generate and analyze the data, respectively. A sample implementation code is provided in Appendix D.
Study 1: Model Performance and Comparisons of Special Effects for Continuous Data under Local Dependence
In this study, we evaluated the performance of the proposed models for special effects for continuous data under local dependence. We set the number of general factors and special factors as K
g
= K
s
= 3, with six items per special factor (i.e., J = 18). The true loading matrix was:
Two factorial correlations between general factors were investigated:
Simulation Results for Study 1.
Note. λ
g1
~
Study 2: Model Performance and Comparisons of Special Effects for Categorical Data
In this study, we evaluated the performance of the revised GPCFA for special effects with categorical data under the assumption of local independence, which are common in testlet effect models. We investigated two levels of the number of categories, M = 2 and 4, for all items. The number of general factors and special factors were set as Kg = 1 and Ks = 3, respectively, with six items per special factor, namely, J = 18. The true loading matrix was set as:
General factor loadings were set from 0.5 to 0.75 with an interval of 0.05, repeated three times. All 6 nonzero loadings for each special factor were determined by the effect size, as in study 1. All these loadings with nonzero true values were set as specified loadings which were freely estimated. Other loadings were set as
Simulation Results for Study 2.
Note.
Empirical Examples
Study 1: Humor Styles Questionnaire: Special Effect for Continuous Data
It is common to encounter special effects such as method effects due to wording, item formats, or reverse items in multidimensional psychological scales. In this study, the Humor Styles Questionnaire (HSQ) (Martin et al., 2003) was used to test if there’s a method effect for the reverse items using the GPCFA model. The HSQ consists of 32 items with four general factors, including 11 reverse items (Appendix E). The public dataset from 1070 respondents can be found at https://openpsychometrics.org/_rawdata/, which include 130 missing values.
Significant Loadings and Residual Estimates for the Humor Styles Questionnaire.
Note. Fg1 = affiliative humor; Fg2 = self-enhancing humor; Fg3 = aggressive humor; Fg4 = self-defeating humor; Fs1 = special effect for reverse items; LD = local dependence (only significant terms were presented); D: effect size; only significant or above .1 loadings are presented; underscored are unspecified loadings.
aIndicates non-significant loadings.
Different CFA Models’ Fitness for the Humor Styles Questionnaire.
Note. RMSEA = root mean square error of approximation; CFA = confirmatory factor analysis; M0 = baseline (no cross-loading or residual covariance); M1 & M2 = all significant loadings identified in the GPCFA; M1 = GPCFA with local independent; M2 = GPCFA with all significant residual covariance.
Factorial Correlation for the Humor Styles Questionnaire.
Note. All correlation estimates are significant.
Study 2: PISA Reading Assessment: Special Effect for Categorical Data
Educational assessments with testlet effects are common in psychometrics. In this study, we used the PISA reading assessment for UK in 2000 to explore the testlet effects by GPCFA (Chen & de la Torre, 2014). The dataset contains 1039 responses with 26 released items from six independent articles from booklet 8 and 9. For the baseline model, there was one general factor as reading literacy and 5 special factors as 5 independent articles (the last article was excluded due to too few items). The correspondent articles with specified items can be found in Appendix F (there is no cross-loading).
Parameter Estimates of the PISA Reading Assessment.
Note. Fg1: Reading literacy; Fs1 - Fs5: 5 different articles; D: effect size; only significant and above .1 loadings are presented; underscored are unspecified loadings.
aIndicates non-significant loadings.
Different CCFA Models’ Fitness for the PISA.
Note. RMSEA = root mean square error of approximation; CCFA = categorical confirmatory factor analysis; M0 = baseline; M1 = GPCFA suggested.
Discussion
Special effects including the method and testlet effects are common issues in psychological and educational measurement. This research extends the GPCFA framework to accommodate special effects for continuous and categorical data by modifying the factor structure. The revised GPCFA can be related to different bifactor, MTMM-type, and testlet effect models by setting different constraints. A useful indicator D was produced to measure and compare the special effect sizes. Models for special effects under the revised GPCFA framework offer multidimensionality for both the general and special factors (or traits) and automatically inherit GPCFA’s benefits to accommodate the partially confirmatory design with regularizations, local dependence, mixed-type formats, and missingness jointly. As a result, it provides a chance to easily specify novel models for potential applications within one framework.
Compared with traditional bifactor, MTMM and testlet effect models, the revised GPCFA framework is more flexible in three ways. First, the proposed model allows for multiple general factors and special factors with different constraints on factorial correlation and local dependence flexibly. Second, different from traditional rotation and MLE, the Bayesian Lasso method in GPCFA can estimate loading matrix and local dependence at the same time while dealing with mixed types of variables and missingness in a unified framework. Third, the regularization of loading structure covers a wide range of the substantive continuum. This partially confirmatory approach allows for regularization of the loading patterns, resulting in a simpler structure in both the general and effect parts. Unspecified loadings for both the general and special factors also offer us more flexibility to incorporate uncertainty (e.g., addressing cross-loadings) during the modeling process. Moreover, one can analyze both the testlet effect in IRT and method effect in factor analysis with the equivalent effect size in a unified way.
Two simulation studies and corresponding real-life studies were adopted to evaluate and illustrate how the revised GPCFA framework can address special effects for both continuous and categorical cases under different conditions. Specifically, the small effect size (∼.1) achieved better model estimation in simulation studies and large effect size (∼.2) might lead to overestimating for some parameters. The real-life examples based on the Humor Styles Questionnaire and PISA reading assessment illustrate how GPCFA can be used to test different special effects and calculate the effect size. In practice, the special effects are usually around 0.1, and we can consider random effects with an effect size less than that value as negligible. Finally, the R package LAWBL (Chen, 2022) is free and powerful in implementing GPCFA with different types of constraints.
There are a few limitations in this research. From the estimation perspective, Bayesian Lasso is time-consuming and requires raw data to estimate the procedure. In contrast, the frequentist approach based on MLE is faster, and only needs summary statistics for typical models. Future research can explore alternative algorithms that can combine the flexibility of the MCMC and the efficiency of the MLE. The variational inference (e.g., Anderson & Peterson, 1987; Hinton & Van Camp, 1993) based on the Bayesian approach is promising, which can inherit many of MCMC’s benefits and balance computational efficiency and accuracy at the same time. Recently, this algorithm had been introduced under the FA and structural equation modeling context (Dang & Maestrini, 2022; Khan et al., 2010), which can provide a basis for its implementation under the revised GPCFA framework. From the structure perspective, extensions of the GPCFA to accommodate more variants such as the three-parameter testlet effect model are worth exploring. Further research can also empower the revised GPCFA framework for research design in complex settings by incorporating both structural and measurement components. By regularizing different structural and measurement parametric matrices flexibly, one is allowed to create many more partially confirmatory designs that can be used for different purposes. Finally, more empirical evidence across different real-life scenarios is still desirable to demonstrate the capacity of GPCFA to accommodate various special effects in practice.
Supplemental Material
Supplemental Material - Accommodating and Extending Various Models for Special Effects Within the Generalized Partially Confirmatory Factor Analysis Framework
Supplemental Material for Accommodating and Extending Various Models for Special Effects Within the Generalized Partially Confirmatory Factor Analysis Framework by Yifan Zhang and Jinsong Chen in Applied Psychological Measurement.
Footnotes
Acknowledgments
We wish to thank the editor and reviewers for their helpful comments on the manuscript.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by a General Research Fund Grant (17603022) from the Hong Kong Research Grants Council.
Supplemental Material
Supplemental material for this article is available online.
