The linear composite direction represents, theoretically, where the unidimensional scale would lie within a multidimensional latent space. Using compensatory multidimensional IRT, the linear composite can be derived from the structure of the items and the latent distribution. The purpose of this study was to evaluate the validity of the linear composite conjecture and examine how well a fitted unidimensional IRT model approximates the linear composite direction in a multidimensional latent space. Simulation experiment results overall show that the fitted unidimensional IRT model sufficiently approximates linear composite direction when correlation between bivariate latent variables is positive. When the correlation between bivariate latent variables is negative, instability occurs when the fitted unidimensional IRT model is used to approximate linear composite direction. A real data experiment was also conducted using 20 items from a multiple-choice mathematics test from American College Testing.
It is generally recognized that a test, despite the original intention to assess a single targeted construct, often inadvertently includes secondary constructs that are peripheral to the assessment but could influence the interaction between the test taker and the test items. Humphreys (1982) for example, coined the term “systematic heterogeneity” to describe the characteristic of a test that is sufficiently broad to include minor contents in order to become a meaningful and valid test. As argued by many others (Hambleton & Swaminathan, 1985; Reckase, 1979, 1985; Stout, 1987; Ackerman et al., 2003; Del Rosario Basterra, Trumbull, & Solano-Flores, 2011), in educational and psychological testing, a test that is designed to measure a single construct—psychological, educational, or otherwise, is likely to involve other minor constructs in practice. Thus, there exists a dilemma between the psychometric desire for assessing a single construct versus the need for a test to function meaningfully as a valid instrument. Ip (2010) called this the dimensionality versus validity dilemma.
Yet another scenario the dimensionality versus validity dilemma could arise is when a test is designed to measure multiple constructs. However, because of practical constraints, it is desirable to report a unidimensional score. Researchers have to determine if the test is sufficiently unidimensional for reporting purpose.
One of the most commonly used psychometric models for measuring a single construct is the unidimensional item response theory (UIRT). Because of the recognition about the potential presence of multiple dimensions in a test, multidimensionality IRT (MIRT) has also been increasingly used to analyze educational and psychological data. The trend of MIRT is also partly driven by the emergence of stable MIRT software programs (Han & Paek, 2014). With the increasing availability of mature software tools in MIRT to analyze data, the balance in the discussion about the dimensionality versus validity dilemma is shifting. As simulation and analytic tools become more efficient and less costly, this dilemma now can be thoroughly investigated from both MIRT and UIRT perspective.
then the fitted UIRT to the data represents a linear composite that points in the direction of the first eigen vector of the matrix (Wang, 1987, p.30). Here, represents the vector of (dichotomous) response from the ith person (i=1,..,I) to the jth item (j=1,..,J), the vector of abilities in the m-dimensional latent space (), the matrix of discrimination parameters in a compensatory MIRT, the vector of intercept item parameters, and the cumulative normal distribution function. This conjecture is important because if it is valid it will allow a simplified mechanism for understanding the fitted UIRT to multidimensional data. A researcher will be able to anticipate the behavior of the UIRT with varying underlying MIRT parameters. As a result, sensitivity analysis can be efficiently conducted for assessing how specific changes to candidate test items could affect the composite direction. The information derived from the LCC can also be used to exclude items that do not fit into what Ackerman (1992) called the validity sector, which refers to the limit of the extent to which the secondary dimensions are allowed into the test. Despite its importance, the LCC has not been comprehensively evaluated, although a substantial literature about the robustness of UIRT to model misspecification does exist (Drasgow & Parsons, 1983; Harrison, 1986; Ackerman, 1994, Ackerman, 1989; Junker & Stout, 1994; Zhang & Stout, 1999; Reise et al., 2014). In an unpublished doctoral dissertation, Kim (1994) compared Wang’s LCC with some other conjectures of composites—all linear in nature, using simulations. However, the conditions for the simulation were limited. It is interesting to note that while authors expressed doubts about linearity of the LCC (Junker, 1994), preliminary simulation work from Junker’s group (Kim, 1994, p. 97) seemed to show consistency between the fitted UIRT and Wang’s linear composite direction. We will further discuss the linearity issue in the Discussion section.
In the current paper we evaluated the LCC using both simulated and real data. In our simulation experiment, the evaluation of the conjecture follows these steps: 1. Generate response data from an MIRT; 2. Compute the composite direction based on the true MIRT item parameters per Wang’s LCC; 3. Mathematically derive the conjectured fitted UIRT item parameters and person ability using multivariate normal theory (Zhang & Wang, 1998); 4. Fit UIRT to the data; and 5. Compare both item parameter and ability estimates of the fitted UIRT to the conjectured model in Step 3. In the real data analysis, we used a confirmatory MIRT model as the reference model. Estimated MIRT parameters were used to obtain the LCC model, and model parameters between the LCC model and the fitted UIRT were directly compared.
The remainder of the paper is organized as follows. First, we give background about the computation of the linear composite, then we describe the simulation experiment and the real data analysis. We focus on the two-dimensional MIRT. Finally, we provide a discussion.
Computing a Linear Composite
Assume that the item response function can be denoted by the two-dimensional compensatory logistic model (Reckase, 2009) and ignoring the person index
where and are slope parameters for the jth item and is a scalar parameter related to intercept of the jth item. It is also assumed that the latent variables and follows a multivariate normal distribution. Without loss of generality, it’s assumed that and are standardized, that is, , where is a 2×2 positive definite matrix with unit diagonal elements. Under these two assumptions, the true score of a single reported score may be related to the latent variables through a linear composite (Zhang & Wang, 1998). A linear composite of the latent variables is defined to be a standardized linear combination of and such that
where is a vector of weights with non-negative ’s and the following constraint must be satisfied
where and the vector α represents the direction of composite in the two-dimensional latent space. To obtain the elements of vector , let’s first denote as the matrix of two-dimensional item discrimination parameters. The eigen decomposition of is computed to obtain the eigenvector of the first eigenvalue . Once is obtained the first weight is computed as follows
With algebraic manipulation, can be solved via equation (4) which results in the following quadratic equation
where the coefficients of the quadratic equation are , , and . If , the weight can simply be obtained by:
The marginal item response function with respect to a linear composite is represented as a unidimensional response function under the assumptions of multivariate normal and and two-dimensional compensatory logistic model
where
and
Recall that is assumed. Figure S1 in supplementary materials illustrates the linear composite direction in a two-dimensional latent space.
Simulation Experiment
The purpose of the simulation experiment was to examine how similar the linear composite scale, , obtained from the generated two-dimensional MIRT (2D-MIRT) was to the estimated unidimensional IRT (UIRT) scale, . A fully crossed factorial simulation design was ran in R software (R Core Team, 2017). The mirt package (Chalmers, 2012) was the software of choice for generating response data, item calibrations, and scoring respondents. Several factors were manipulated in the simulation experiment including sample size, test length, correlation for and , and correlation for and . A sample size of and were randomly generated from a multivariate normal distribution with and covariance matrix which consisted of unit diagonal elements and correlation, , between the latent variables and . The between the latent variables was manipulated at different levels including −.8, −.4, .0, .4, .8, and 1 (for UIRT model). A test length of 40, 80, 100, and 150 items were randomly generated where and were either positive or negatively correlated. The and were generated from separate uniform distributions such that when correlating the item vectors of and (across items) the empirical correlation (r) was either negative (e.g., the and item vectors computed correlation was r −.90) or positive (e.g., the and item vectors computed correlation was r −.90). For the positively correlated a-parameters level (), first half of the items a-parameters were randomly generated from and second half of the items a-parameters were randomly generated from. For the negatively correlated a-parameters level (), the first half of the items were set to be large (Unif (1.2, 1.6)) for the first dimension, and for the second dimension, a-parameters were set to be small (Unif (.2, .6)). For the second half of the items, the a-parameters were set to be small (Unif (.2, .6)) for the first dimension, and for the second dimension, a-parameters were set to be large (Unif (1.2, 1.6)). So, essentially the positively correlated case was approximate unidimensionality and the negatively correlated case was approximate simple structure. To generate the model intercept, , the multidimensional difficulty , was first randomly generated from . Once and a-parameters were obtained for each item , the following equation was used to obtain (Reckase, 2009): where which represent the overall multidimensional discrimination of the item. The vector plots (Ackerman, 1996) located in Figure 1 show the item locations in a two-dimensional latent space when the a-parameters are either positively (a) or negatively (b) correlated. Once the generated 2D-MIRT model was obtained, the procedure discussed in the previous section was used to compute the linear composite scale, , along with corresponding item parameters and . Dichotomous response data generated from the 2D-MIRT model was fitted using UIRT 2PL model to obtain the estimated unidimensional scale, .
Simulation Experiment: Item vector plots illustrating item locations in a two-dimensional latent space where and are either positively (a) or negatively (b) correlated.
An expectation-maximization (Bock & Aitkin, 1981) algorithm was used to estimate the item parameters and of the UIRT model with a convergence criterion set at .0001. For the 2D-MIRT, the first item’s was fixed to zero to resolve rotational indeterminacy issues. Once the item parameter estimates were obtained, an expected a-posteriori (Embretson & Reise, 2000) estimator was used to obtain the estimated latent variable parameters, . To allow for a fairer comparison between the two scales, and , both a scaling and linear transformation procedure was implemented. The scaling procedure was first implemented to standardize the and scales using a simple z-score conversion formula. Once and were rescaled, a linear transformation method was implemented as follows (Kolan & Brennan, 2014)
where and . A similar linear transformation procedure was done to make the results between item parameters (, ) and (, ) comparable using the same metric. Once all model parameters were rescaled, metrics such as mean absolute difference (MAD) and correlation, , were used to evaluate the results. The MAD was computed as follows
In the case of computing the MAD across item parameters (, ) and (, ), total sample size was replaced with total test length in the equation above. There was a total of joint conditions with each joint condition being replicated 100 times. Table S1 under supplementary materials provides a summary of the simulation design.
Results from Simulation Experiment
MAD results from the simulation experiment are presented in Figure 2–4. The results from the simulation experiment are located in supplementary materials under Figures S4, S5, and S6. The values in the tables represent the average MAD and with corresponding standard errors of MAD and across the 100 replications within each joint condition. Results showing how similar the estimated UIRT scale is to the linear composite scale obtained from the generated 2D-MIRT model are located in Figure 2 Results showed that on average, as the number of items increased, the MAD decreased and results increased between and . Additionally, sample size N had minimal impact on the MAD and between and when and were independent or positively correlated. However, when both and and and were negatively correlated, the MAD decreased and increased as sample size increased.
MAD results between and from simulation study. The points in the graph represent the average MAD value across 100 replications while the bars represent the standard errors. Corsl in the legend is an abbreviation for “correlation of slopes.”
MAD results between and from simulation study. The points in the graph represent the average MAD value across 100 replications while the bars represent the standard errors. Corsl in the legend is an abbreviation for “correlation of slopes.”
MAD results between and from simulation study. The points in the graph represent the average MAD value across 100 replications while the bars represent the standard errors. Corsl in the legend is an abbreviation for “correlation of slopes.”
Compared to when correlation between and was negative, the MAD and results generally improved when the correlation between and was positive, regardless of correlation between and . When and were independent, the MAD and values between and were similar under the positively and negatively correlated and conditions. Figure 2 also shows that the results for ρ=−.4 were significantly better than for ρ=−.8. Indeed, the results for ρ=−.4 had only small differences from the results for positive ability correlation, especially for the three longer test lengths. Note that the baseline scenario for from the UIRT model is when . In summary, comparing the results from the two simulation experiments for positively and negatively correlated , , the MAD and results for were comparable when the correlation between the latent variables and in the simulation experiment were independent or positive. In contrast, when the correlation between the latent variables and in the simulation experiment were negative, especially when ρ=−.8, the MAD and results in the simulation experiment were poorer than the results from the baseline scenario.
Results showing how similar the estimated slope parameters in the fitted UIRT model to the linear composite model obtained from the generated 2D-MIRT are located in Figure 3. Results showed that on average, as the sample size N increased, the MAD noticeably decreased and results slightly increased between and . Indeed, r remained close to 1 when sample size is high. Results indicated that test length had minimal impact on the MAD and between and when and were orthogonal or positively correlated. For negatively correlated and , MAD decreased from test length=40 to test length=80 but showed no further tendency to decrease with further increases in test length. When both sample size and test length were high, the MAD was low and r was high even when and were highly negatively correlated (e.g., ρ=−.8). The MAD result tended to improve when the correlation between and was positive. When ρ was negative, the MAD and results between and were better under the positively correlated and condition than the negatively correlated and condition. The baseline results for from the UIRT model is when . These results indicated that the correctly specified UIRT model (ρ=1), values for r for recovered -parameters were similar to the incorrectly specified UIRT model (ρ=0).
Results showing how similar the estimated intercept parameters in the fitted UIRT model to the linear composite model obtained from the generated 2D-MIRT are located in Figure 4. Results show that as the sample size increased, the MAD decreased and remained close to 1 between and . Results show that test length and correlation between both and and and had minimal impact on the MAD and results between and . The baseline results for from the UIRT model are when . These results indicated that the incorrectly specified UIRT model recovered -parameters equally-well as the correctly specified UIRT model.
To further illustrate the results presented in this study, Figure S7 plots the relationship between the expected test score under UIRT and MIRT when (a) and (b) . The results from the plots show a positive relationship between the expected test scores when , but an independent relationship between the expected test scores when . Figure 5 illustrates the relationship between the estimated from UIRT and generated and from MIRT at The results illustrate a positive relationship between the three latent variables. Figure S4 in supplementary materials illustrates the positive relationship between estimated from UIRT and .
Scatterplot between generated from MIRT, generated from MIRT, and estimated from UIRT; In this example, , , , and the correlation between and is negative.
Real Data Experiment
A real data experiment was conducted using a 20-item multiple-choice mathematics test obtained from American College Testing (ACT). The purpose of this experiment was to examine how well a fitted UIRT model approximates the linear composite direction, in a multidimensional latent space when real data is presented. A sample size of 4000 was used for the analysis. The responses to the 20 multiple-choice items were dichotomously scored where 0 was coded for an incorrect response and 1 was coded for a correct response. A UIRT model and 2D-MIRT were both fitted to the observed response data. In the 2D-MIRT analysis, the 20 items were either substantively identified as “pure math” (i.e., 10 only loaded on ) or “math and verbal” (i.e., 10 loaded on both and ). An expectation-maximization algorithm in the mirt package (Chalmers, 2012) was used to estimate the item parameters and for the UIRT model and , , and for the 2D-MIRT. The convergence criterion for the UIRT model and 2D-MIRT was set at .001. Once the item parameter estimates were obtained, an expected a-posteriori estimator was used to obtain the estimated latent variable parameters, in the 2D-MIRT and in the UIRT model. The fitted 2D-MIRT was used to compute the linear composite direction in equation (3) with corresponding item parameters and computed in equations (11) and (12). The estimated factor correlation between latent variables and was .42. After rescaling, the fitted unidimensional scale, and the linear composite scale, obtained from the fitted 2D-MIRT were compared using the MAD and metrics. Results comparing to and to were also examined using the MAD and metrics. Results from the real data experiment showed that the fitted UIRT model was sufficient (i.e., acceptable MAD values) in approximating the linear composite direction based on the fitted 2D-MIRT. The MAD and between and was .08 and 1, respectively. Results for the slope parameters show that the MAD and between and was .02 and 1, respectively. Results for the intercept parameters show that the MAD and between and was .03 and 1, respectively. Item parameter estimates including and in the UIRT model, , , and in the 2D-MIRT, and and in the marginal response function with respect to a linear composite are provided in Table S2 under supplementary materials. Vectors for the estimated item parameters with the corresponding linear composite direction is in Figure 6.
Item vectors for the estimated item parameters obtained from the mathematics test dataset. The linear composite direction is denoted by .
Discussion
The purpose of this study was to investigate through empirical means how well the fitted UIRT model approximates the linear composite direction in a multidimensional latent space, as defined by Wang’s LCC (1987). Theoretical development on the issue is considered highly challenging, if not impossible (Junker, 1994; Junker & Stout, 1994). Our extensive simulation results overall show that the fitted UIRT model sufficiently approximates when correlation between and is positive. When the correlation between and is negative, instability occurs when the fitted UIRT model is used to approximate . In the real data ACT example, correlation between latent traits is moderately positive and the LCC appears to be valid. Comparing results between (a) positive correlation, and (b) negative correlation between (a1, a2), MAD and r are generally slightly worse for (b). In summary, the linear approximation works best when the latent distributions are positively correlated and the discrimination parameters across dimensions are also positively correlated.
Our results suggest that the LCC is not universally true. For example, the linear composite is a rather poor approximation when the latent distributions are negatively correlated at a high level and at the same time sample size and test length are low. However, in most practical applications the correlations between abilities are positive—a phenomenon that has long been observed. Historically, Spearman (1927) used the term positive manifold to characterize different cognitive abilities that are positively correlated and attributed positive manifold to a general underlying factor, or the well-known g-factor in intelligence testing. More recent work seemed to suggest that positive manifold could also emerge purely by positive beneficial interactions between cognitive processes during development and that the underlying factor played no role (Van der Maas et al., 2006). In any case, when positive manifold is present, the LCC should still approximately hold. Additionally, the extent of accuracy of the approximation will depend on many factors including, most importantly, the multidimensional structure of the item. Negative correlated latent traits are present in some psychological measurement (e.g., negative, and positive effects, Watson et al., 1988) but are not common.
Although we have investigated how different patterns of item loadings in the multidimensional space (see Figure 2), a limitation of our investigation is that the issue of possible nonlinearity composite has not been thoroughly covered. One example of nonlinearity is when the mean values in () do not follow a linear trend when condition on an overall measure of ability such as total score. Specific patterns of item structures may lead to such nonlinearity. Ip et al. (2019) provided an example of such condition, which was termed non-proportional ability requirement (NPAR). Using PISA data as an example, the authors described a situation under which disproportionate cognitive demand across the dimensions is required for items that reside at certain regions (e.g., high ends) of the latent abilities. Another example is vertical scaling (Carlson, 2017; Strachan et al., in press). It is therefore possible that in some cases a nonlinear approximation of the composite will better represent the fitted UIRT. Indeed, Junker (1994) expressed skepticism about the linearity of the fitted unidimensional model for representing the multidimensional latent space (also see Kim, 1994), and Ip et al. (2013) offered a nonlinear version of the approximation, which was based on solving a system of nonlinear equations. Their preliminary results based on simulations showed that the nonlinear approximation tracked closely how multidimensional responses is represented by a single dimension, which they called the functional dimension.
The current study is also limited to the study of MIRT of two dimensions. If the number of dimension m is more than two and the dimensions still form a positive manifold, then we expect the LCC continues to hold. The degree to which the approximation holds is likely also to depend on the factors we discussed above. When m>2 in the positive manifold, analogous to variance explained in principal component analysis, the first eigen value of could also provide information of how well the fitted UIRT represents the composite direction. Another limitation of the real-data example is the use of 2PL model which ignores guessing effect. Finally, the sample size of 250 may also be too small for MIRT to function well.
By understanding the conditions under which the LCC works best, one can utilize the results of this study in different ways. Reise et al. (2014) argued that many psychological constructs have substantive breath and heterogeneous contents that result in multidimensional responses and the standard approach of using goodness-of-fit indexes to assess fit of UIRT cannot guarantee that “the common target latent ability is identified correctly or that estimated item parameters properly relate the relation between the item responses and the common latent trait.” The authors went on to suggest the approach of comparing slope parameters in UIRT and MIRT alternatives. The result in this study can be used to inform this kind of approach. For example, the LCC can be used to quickly see how parameters in a fitted UIRT change with the addition or deletion of items by examining the first eigen vector of , the angle of the linear composite, and the corresponding slope parameter of the fitted UIRT. The procedure can then be used to assess if an item is acceptable in terms of its deviation from the linear composite direction.
Supplemental Material
sj-doc-1-apm-10.1177_01466216221084218 – Supplemental material for Evaluation of the Linear Composite Conjecture for Unidimensional IRT Scale for Multidimensional Responses
Supplemental material, sj-doc-1-apm-10.1177_01466216221084218 for Evaluation of the Linear Composite Conjecture for Unidimensional IRT Scale for Multidimensional Responses by Tyler Strachan, Uk Hyun Cho, Terry Ackerman, Shyh-Huei Chen, Jimmy de la Torre and Edward H. Ip in Applied Psychological Measurement
Supplemental Material
sj-docx-2-apm-10.1177_01466216221084218 – Supplemental material for Evaluation of the Linear Composite Conjecture for Unidimensional IRT Scale for Multidimensional Responses
Supplemental material, sj-docx-2-apm-10.1177_01466216221084218 for Evaluation of the Linear Composite Conjecture for Unidimensional IRT Scale for Multidimensional Responses by Tyler Strachan, Uk Hyun Cho, Terry Ackerman, Shyh-Huei Chen, Jimmy de la Torre and Edward H. Ip in Applied Psychological Measurement
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research,authorship,and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research,authorship,and/or publication of this article: This work was supported by the Institute of Education Sciences (R305D150051).
ORCID iD
Tyler Strachan
Supplemental Material
Supplemental material for this article is available online.
References
1.
AckermanT. A. (1989). Unidimensional IRT calibration of compensatory and noncompensatory multidimensional items. Applied Psychological Measurement, 13(2), 113–127. https://doi.org/10.1177/014662168901300201
2.
AckermanPL (1992). Predicting individual differences in complex skill acquisition: Dynamics of ability determinants. Journal of Applied Psychology, 77(5), 598–614. https://doi.org/10.1037/0021-9010.77.5.598
3.
AckermanT. (1996). Graphical representation of multidimensional item response theory analyses. Applied Psychological Measurement, 20(4), 311–329. https://doi.org/10.1177/014662169602000402
4.
AckermanT. A.GierlM.WalkerC. M (2003). Using multidimensional item response theory to evaluate educational and psychological tests. Educational Measurement: Issues and Practice, 22(3), 37–51.
5.
AckermanT. A. (1994). Using multidimensional item response theory to understand what items and tests are measuring. Applied Measurement in Education, 7, 255–278.
6.
AckermanT. A.MaY.IpE.H. (In press). Comparison of three unidimensional approaches to represent a two-dimensional latent ability space. In Proceedings of the International Meeting of the Psychometric Society, New York City, July 2018.
7.
BockR. D.AitkinM. (1981). Marginal maximum likelihood estimation of item parameters: Application of an EM algorithm. Psychometrika, 46(4), 443–459. https://doi.org/10.1007/bf02293801
8.
CamilliG. (1992). A conceptual analysis of differential item functioning in terms of a multidimensional item response model. Applied Psychological Measurement, 16(2), 129–147. https://doi.org/10.1177/014662169201600203
9.
CarlsonJ. E (2017). Unidimensional vertical scaling in multidimensional space (ETS RR-17-29). Educational Testing Service.
10.
ChalmersP. R (2012). Mirt: A multidimensional item response theory package for the R environment. Journal of Statistical Software, 48(6), 1–29. https://doi.org/10.18637/jss.v048.i06
11.
Del Rosario BasterraM.TrumbullE.Solano-FloresG (2011). Cultural validity in assessment. Routledge.
12.
DrasgowF.ParsonsC. K. (1983). Application of unidimensional item response theory to multidimensional data. Applied Psychological Measurement, 7(2), 189–199. https://doi.org/10.1177/014662168300700207
13.
EmbretsonS. E.ReiseS. P (2000). Item response theory for psychologists. Erlbaum.
HarrisonD. A. (1986). Robustness of Irt Parameter Estimation to Violations of The Unidimensionality Assumption. Journal of Educational Statistics, 11(2): 91–115.
16.
HanK. T.PaekI. (2014). A review of commercial software packages for multidimensional IRT modeling. Applied Psychological Measurement, 38(6), 486–498. https://doi.org/10.1177/0146621614536770
17.
HumphreysL.G (1982). Systematic heterogeneity of items in tests of meaningful and important psychological attributes: A rejection of unidimensionality. In: Unpublished manuscript. University of Illinois at Urbana-Champaign.
18.
IpE. H. (2010). Empirically indistinguishable multidimensional IRT and locally dependent unidimensional item response models. The British Journal of Mathematical and Statistical Psychology, 63(2), 395–416. https://doi.org/10.1348/000711009X466835
19.
IpE. H.StrachanT.FuY.ChenS.RutkowskiL.LayA.WillseJ.AckermanT. (2019). Bias and bias correction method for non-proportional abilities requirement (NPAR) tests. Journal of Educational Measurement, 56(1), 1–22. https://doi.org/10.1111/jedm.12204
JunkerB. W (1994). Inference, robustness, and essential unidimensionality. Talk given at the 10th annual workshop on item response theory modeling. University of twente.
22.
JunkerB. W.StoutW. F (1994).Robustness of ability estimation when multiple traits are present with one trait dominant. In LaveaultD.ZumboB. D.GessaroliM. E.BossM. W. (Eds), Modern theories of measurement: Problems and issues (pp. 31–61). University of Ottawa.
23.
KahramanN.ThompsonT. (2011). Relating unidimensional IRT parameters to a multidimensional response space: A review of two alternative projection IRT models for subscale scores. Journal of Educational Measurement, 48(2), 146–164. https://doi.org/10.1111/j.1745-3984.2011.00138.x
24.
KimH.R (1994). New techniques for the dimensionality assessment of standardized test data. In: Doctoral thesis. University of Illinois.
25.
KolanM. J.BrennanR. L (2014). Test equating, scaling, and linking: Methods and practices. Springer.
26.
LuechtR. M.MillerT. R. (1992). Unidimensional calibrations and interpretations of composite trait for multidimensional tests. Applied Psychological Measurement, 16(3), 279–293. https://doi.org/10.1177/014662169201600308
27.
McDonaldR. P (1997). Normal-ogive multidimensional model. In HambletonR.K.van der LinderW. (Ed.). Handbook of modern item response theory (pp.257-269). Springer. https://doi.org/10.1007/978-1-4757-2691-6_15
28.
R Core Team (2017). R: A language and environment for statistical computing. In: R foundation for statistical computing. URL. https://www.R-project.org/
29.
ReckaseM. D. (1979). Unifactor latent trait models applied to multifactor tests: Results and implications. Journal of Educational Statistics, 4(3), 207–230. https://doi.org/10.3102/10769986004003207
30.
ReckaseM. D. (1985). The difficulty of test items that measure more than one dimension. Applied Psychological Measurement, 9(4), 401–412. https://doi.org/10.1177/014662168500900409
31.
ReckaseM. D (2009). Multidimensionality item response theory. : Springer-Verlag.
32.
ReckaseM.D.CarlsonJ.E.AckermanT.A.SprayJ.A (1986). The interpretation of unidimensional IRT parameters when estimated from multidimensional data. In: Paper presented at the annual meeting of the Psychometric Society, Toronto.
33.
ReiseS.P.CookK.F.MooreT.M (2014). Evaluating the impact of multidimensionality on unidimensional item response theory model parameters. In ReiseS.P.RevickiD.A. (Eds.), Handbook of item response theory modeling: Applications to typical performance assessment. Routledge. (pp. 13–40).
34.
SpearmanC. (1927). The abilities of man, New York: Macmillan.
35.
StoutW. (1987). A nonparametric approach for assessing latent trait unidimensionality. Psychometrika, 52(4), 589–617. https://doi.org/10.1007/bf02294821
36.
StrachanT.ChoJ.ChenS-H.AckermanT.KimK.Y.WillseJ.WeeksJ.IpE.H (in press). Using a projection IRT method for vertical scaling when construct shift is present. Journal of Educational Measurement.
37.
van der MaasH. L. J.DolanC. V.GrasmanR. P.WichertsJ. M.HuizengaH. M.RaijmakersM. E. (2006). A dynamical model of general intelligence: the positive man-ifold of intelligence by mutualism. Psychological Review, 13, 842–60. DOI: 10.1037/0033-295X.113.4.842.
38.
WangM.M. (1987). Fitting a unidimensional model to multidimensional item response data (ONR Rep. 042286). Iowa City, IA: University of Iowa.
39.
WatsonDClarkLATellegenA (1988). Development and validation of brief measures of positive and negative affect: the PANAS scales. Journal of Personality and Social Psychology, 54(6), 1063–1070. https://doi.org/10.1037//0022-3514.54.6.1063
40.
ZhangJ.StoutW. F. (1999). The theoretical DETECT index of dimensionality and its application to approximate simple structure. Psychometrika, 64, 213–249.
41.
ZhangJ.StoutW. (1999a). The theoretical DETECT index of dimensionality and its application to approximate simple structure. Psychometrika, 64(2), 213–249. https://doi.org/10.1007/bf02294536
ZhangJ.WangM. (1998, April). Relating reported scores to latent traits in a multidimensional test. Paper presented at the annual meeting of the American Educational Research Association. San Diego, CA.
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.