Abstract
It has been argued that item response theory trait estimates should be used in analyses rather than number right (NR) or summated scale (SS) scores. Thissen and Orlando postulated that IRT scaling tends to produce trait estimates that are linearly related to the underlying trait being measured. Therefore, IRT trait estimates can be more useful than summated scores when examining relationships between test scores and external variables. Also, when the model holds, IRT trait estimates possess an interval scale that is a property assumed for dependent variables by most statistical procedures used in educational research. The objective of this study was to use Monte Carlo methods to compare the performance of IRT trait estimates and SS scores in predicting outcome variables in the context of health and behavioral assessment. The use of scores based on the graded-response model versus summated scores was compared. Results indicated that IRT-based scores and summated scores are comparable when evaluating the relationships between test scores and outcome measures. Thus, applied researchers could use summated scores in predictive studies and circumvent evaluating the assumptions underlying use of IRT-based scores.
Introduction
The study of relationships between estimates of individuals’ traits or behaviors and outcome variables are prevalent in health and behavioral sciences research. The estimates of the traits or behaviors are typically derived from set(s) of items with ordinal response scales (e.g., Likert items). Based on classical test theory (CTT), item responses for a scale or subscale are aggregated or summed to obtain scale or subscale scores (i.e., summated scores), and these scores may be used to assess the relationship between the test scores and outcome measures. Alternatively, item response theory (IRT) models can be used to scale the items and obtain estimates of the underlying trait being measured. These trait estimates may then be used to assess the relationship between test scores and outcome measures. However, the literature does not clearly indicate which type of scores should be used in practice by researchers.
Some researchers have discussed that the two types of scores are not always linearly related (Harwell & Gatti, 2001; Thissen, Nelson, Rosa, & McLeod, 2001). Specifically, summated scores are usually observed to be linearly related to the IRT trait estimates under the Rasch model but are less similar to the trait estimates obtained under other commonly used IRT models such as the graded-response (GR) model (Harwell & Gatti, 2001) or the combined three-parameter logistic and GR model (Thissen et al., 2001).
Thissen and Orlando (2001) further discussed that since IRT scaling tends to produce trait estimates that are linearly related to the underlying trait being measured, IRT trait estimates may be more useful than CTT-based scores when examining the linear relationship between test scores and external variables (e.g., outcome measures). Theoretically, if the underlying trait being measured is linearly related to an external variable, the IRT trait estimates also have a linear relationship with that variable. Therefore, CTT-based scores may be less accurate in predicting an outcome measure. Thissen and Orlando (2001) also discussed that a nonlinear relationship between summated scores and the underlying trait typically results from ceiling and/or floor effects, which can be accounted for by IRT scaling.
Harwell and Gatti (2001) argued that there are also theoretical advantages to IRT trait estimates over CTT-based (number right [NR] or summated scale [SS]) scores. NR and SS scores are essentially a summation of ordinally scaled variables and therefore do not possess an interval scale. There are two problems with the use of these scores in statistical analyses such as regression analyses: nonnormality and “incoherence” between the test score scale and the latent trait scale (Harwell & Gatti, 2001). Both problems may yield biased parameter estimates and misleading results. Alternatively, IRT can be used as a scaling model to transform item responses to trait estimates that possess an interval scale (Harwell & Gatti, 2001). If the IRT model is appropriate for the data, it is reasonable to expect that using IRT trait estimates would lead to improved prediction of outcome variables. Ferrando and Chico (2007) further argued that maximum likelihood (ML) trait estimates based on IRT models better reflect the characteristics of test items and therefore provide more test information than either unweighted or weighted raw scores. The increased test information afforded by IRT models versus CTT applications implies the possibility of greater accuracy or precision in trait estimates versus CTT-based scores.
Finally, IRT models offer more flexibility in scaling sets of items with the same response scale (summated scales). By summing item responses to create total or subscale scores, the CTT model weights each item equally, an assumption necessary to estimate internal consistency reliability for a scale (Li & Wainer, 1998; Raykov, 1997). Equal weighting of items also implies each item is equally related to the underlying trait being measured. Last, since the same response scale is used across the set of items, the CTT model assumes the response options have the same meaning across the set of items. Alternatively, a variety of IRT models may be applied to summated scales that address some of these constraints. The rating scale model discussed by Andrich (1978) most closely approximates the summated score approach used with health and behavioral assessments. This model incorporates a common discrimination parameter and a common response scale across the set of items. Other models can be estimated that relax the assumption of a common discrimination parameter and/or relax the assumption of a common response scale (e.g., GR IRT model). The different IRT models can be estimated and compared to evaluate the extent to which these assumptions are met. Note that relaxing the assumption of a common discrimination parameter is equivalent to the case where factor loadings vary across items measuring the same trait.
Despite the theoretical advantages attributed to IRT traits estimates over CTT-based scores, these advantages are only derived if the IRT model is appropriate. IRT models are based on several assumptions (e.g., dimensionality, functional form of the model, and local independence), and these assumptions must be evaluated to validate the IRT application. This complicates the application of IRT models to summated scales since applied researchers must be familiar with methods used to evaluate the underlying assumptions. This is in contrast to the use of CTT-based scores since the “correctness” of the CTT model and inferences based on the use of raw scores is generally not evaluated (Ferrando & Chico, 2007).
Other arguments for the use of CTT-based scores are rooted in the scaling properties of summated scales and robustness of statistical analyses using summated scale scores. While each item (e.g., Likert item) reflects an ordinal scale, researchers have discussed that the aggregation of these items into a scale score measuring a construct may yield or closely approximate an interval scaled score (e.g., Carifio & Perla, 2007). Furthermore, many analyses using summated or raw scores are robust to the assumption of interval data (Glass, Peckman, & Sanders, 1972), although other researchers have discussed that the use of summated scores from Likert items can influence the assessment of relationships (e.g., Russell & Bobko, 1992).
To address the contradictory recommendations surrounding the use of CTT- versus IRT-based scores, previous research has compared parameter estimates from IRT- and CTT-based applications and found that IRT trait estimates and CTT scores are highly correlated for tests composed of dichotomous items (e.g., Fan, 1998; Lawson, 1991; MacDonald & Pauonen, 2002; Ndalichako & Rogers, 1997). For example, based on the data from the Texas Assessment of Academic Skills tests, Fan (1998) compared the performance of the NR scores under CTT and the IRT scores under the one-, two-, and three-parameter IRT models for low and high ability groups, male and female samples, as well as random samples. The results indicated that the correlations between the two ability measures were very high (>.97) under most conditions that were studied.
The focus of prior research was on the relative accuracy of the IRT- and CTT-based parameter estimates. However, there is a lack of empirical studies evaluating the robustness of using IRT trait estimates versus CTT-based scores in prediction studies. The results of a few empirical studies found that validity coefficients based on IRT trait estimates were not consistently better than those of CTT-based scores (Dodd, Koch, & De Ayala, 1989; Ferrando, 1999; McBride & Martin, 1983; Young, 1995). For example, McBride and Martin (1983) found that the use of IRT ability estimates was associated with slightly higher validity coefficients than the use of NR scores when the number of items was equal to or smaller than 10. They also found that validity coefficients based on IRT scores were no better than those based on NR scores when the number of items was more than 30. Dodd et al. (1989) compared the validity coefficients from IRT scores and those from CTT-based scores for a test composed of both multiple choice and essay items. The results showed that the IRT approach yielded slightly lower validity coefficients when compared with the CTT approaches. Finally, Ferrando and Chico (2007) examined the external validity of scores based on the two-parameter logistic (2PL) model. Validity coefficients based on IRT trait estimates (2PL) were found to be similar to those based on NR scores. These researchers also found that the presence of “strong” floor or ceiling effects may provide an advantage to analyses using IRT-based scores since nonlinearity in the relationships exists in the tails of the score distributions. They suggested that further studies should be conducted and consider other factors such as different item distributions and different IRT models (e.g., GR models).
The purpose of this study was to extend previous research and evaluate the use of IRT- versus CTT-based scores in the context of health and behavioral assessments that incorporate polytomous items. Little research has been conducted examining the relationship between IRT- and CTT-based scores for tests based on these types of applications. Furthermore, much of the previous research has involved comparisons based on real data applications. The present study used simulation methods to compare the accuracy of the different scores, as well as the recovery of predicted outcomes or validity coefficients based on IRT trait estimates versus SS scores. Conditions were introduced in the study to evaluate the robustness of the different trait estimates under violations of CTT model assumptions: (a) equal item weights across items and (b) constant response scale across items. The advantages of a simulation study approach are that the true relationships between variables are known and factors that could affect the methods can be manipulated directly.
Method
The manipulated factors of the simulation study included the following: (a) the number of items: 10, 20, and 40; (b) correlation of scores with the outcome measure: 0.3, 0.4, and 0.6; (c) range in item slopes: Uniform [1.2, 2.2] and Uniform [0.5, 2.9]; and (d) sample size: 250, 500, and 1,000. Number of items was manipulated since the number of items affects the reliability or precision of scores and this factor has been found to affect the comparison of IRT- and CTT-based scores (e.g., Ferrando & Chico, 2007). The range of item slopes was varied to relax the assumption that items were equally related to the trait being measured. Response scales for items were fixed at five categories to reflect scales typically used in health and behavioral assessments.
Trait estimates and outcome measures were generated from a standard bivariate normal distribution based on the specified correlation values. For the five-category response items, the GR model (Samejima, 1969) was used to simulate ordinally scaled item responses:
where x represents a response in or above a given category (x = j = 1, . . ., mi), m is the number of thresholds, θ is the person trait parameter, α i is the slope parameter for item i, and bij is the threshold parameter for item i and response category j.
For the varying item slope condition, item slope values were randomly chosen from a distribution that reflected a wide range—Uniform [0.5, 2.9]. This condition yielded items that varied extensively in terms of their association with the underlying trait being measured and reflected a condition that violated the assumption of equal item weights underlying CTT-based scores. For the “constant” item slope condition, item slopes were randomly varied based on a narrow distribution—Uniform [1.2, 2.2]. Slopes were randomly varied over a narrow range as opposed to fixed at a common value to reflect more realistic conditions.
With regard to the threshold parameters for the GR model, the first threshold parameter for the GR model was obtained from a Uniform [−2.5, 0] distribution, and subsequent threshold parameters were obtained by adding a constant from the set [0.5, 1, 1.5]. The value of the constant was randomly determined based on the discrete probability distribution [1/3, 1/3, 1/3].
Table 1 presents an example set of randomly generated item slope threshold parameters under the GR model. To illustrate the different response scales for the different items, these parameters were used to simulate item responses and a GR model was estimated in PARSCALE (Muraki & Bock, 1997). The model estimated in PARSCALE reflects a slightly different parameterization, where bij under the GR model is decomposed into bj − cjk. Under this parameterization, bj is an item threshold parameter and cjk is a category parameter for item j and category k. This alternative parameterization allows for isolating the item location along the trait scale from the category parameters associated with the response scale. The parameter estimates from PARSCALE for the constant slope condition are presented in Table 2. As can be seen, the category parameter estimates, although similar for some items, vary considerably across the set of items.
Simulated Item Parameter Values for a Condition With Test Length of 10
Estimated Item Parameter Values From PARSCALE for a Constant Slope Condition With Test Length of 10
Note that item responses based on a common response scale across items could also be simulated with a rating scale model to compare the results with the model reflecting different response scales. However, as will be presented, the use of CTT-based scores was very robust to the condition where different response scales were simulated. Thus, it was not necessary to compare results under a condition where a common response scale was simulated.
Using the item and ability parameters, item responses were simulated and analyzed using MULTILOG to obtain IRT-based ML trait estimates (
The outcome measures included root mean square deviations (RMSDs) comparing predicted values for the outcome measure (z-score estimates) versus true values (z scores), and these values were calculated at high- (
Note that in this study the empirical validity coefficients all reflected disattenuated coefficients or coefficients that were not corrected for unreliability. While some predictive studies incorporate coefficients corrected for attenuation, in particular, predictive studies in employment testing contexts, most predictive validity studies in health and behavior assessment do not correct for attenuation in outcome measures. Furthermore, correcting for attenuation would not affect the relative comparison of conditions under study.
Results
Table 3 presents average RMSDs comparing predicted values for the outcome measure versus true values across 500 replications for the varying item slope condition and sample size equal to 250. As noted before, this condition reflected items that varied in terms of their association with the underlying trait being measured and reflected a condition under which a model with item slope parameters may be preferred over a model with a single slope parameter for items. Table 4 reports average RMSDs for predicted versus true outcome measure values across 500 replications for the constant slope condition in which item slopes reflect a relatively narrow range.
RMSDs for Predicted Outcome Values (Varying Item Slopes; 250 Persons)
Note. RMSD = root mean square deviation; IRT = item response theory.
RMSDs for Predicted Outcome Values (Constant Item Slopes; 250 Persons)
Note. RMSD = root mean square deviation; IRT = item response theory.
When comparing Tables 3 and 4, it is apparent that there was no practical difference in the results for varying versus constant slope parameter conditions. There was essentially no effect of the number of items on the accuracy of the predicted outcome for either condition. The finding of no effect of number of items was surprising given that there is increased accuracy in scores as the number of items increases. In addition, there was essentially no effect of the type of score (IRT-based score vs. summated score) on the accuracy of the predicted outcome for either the varying slope condition or the constant slope condition. Furthermore, the level of the outcome variable (low, medium, or high) also had no effect on the accuracy of the predicted outcome. The only factor affecting the accuracy of the predicted outcome was the level of the correlation between the predictor and predicted outcome variables. As the correlation increased, the accuracy of the predicted outcome also increased. RMSDs for predicted versus true outcome measure values decreased from around .95 to .80 as the correlation increased from .30 to .60.
It should be noted that the size of the RMSDs given a z-score metric appear rather large. However, these values are consistent with an expectation based on the standard error of estimate (SEE) under the population model. One way of expressing the SEE is in terms of the Pearson correlation between the predictor and outcome measure:
where ρ is the Pearson correlation between the predictor and outcome measure and
Table 5 illustrates the comparison of estimated correlations between the predictor and predicted outcome variables with true values for the varying item slope condition and the constant item slope condition. The values are compared for correlations based on the use of summated scores versus IRT-based scores with a sample size of 250. Table 5 demonstrates that RMSDs for the correlations increased as the true correlation increased, a finding that is not surprising. In general, use of summated scores as a predictor yielded smaller RMSDs for the correlations than using the IRT-based scores, but this effect decreased as the number of items increased. However, the differences between the RMSDs were more noticeable under the varying slope condition, a condition that was hypothesized to favor the use of IRT-based scores.
RMSDs for Correlations (250 Persons)
Note. RMSD = root mean square deviation; IRT = item response theory.
The RMSD results for predicted outcomes for sample size of 500 and 1,000 were virtually equivalent to those for N = 250 and so are not presented. Thus, sample size does not affect this outcome measure. Tables 6 and 7 present the RMSD results for correlation based on a sample size of 500 and 1,000 for the varying and constant slope conditions. As for the results based on N = 250, the RMSDs for the correlation increased as the true correlation between the predictor and outcome increased. Also, the summated score as a predictor yielded slightly smaller RMSDs for the correlation than using IRT-based scores, and as for N = 250, the effect decreased as the number of items increased. In comparing the results for N = 250 versus 1,000 (Table 4 vs. Table 7), RMSDs are somewhat smaller for N = 1,000 under the constant and varying slope conditions. In addition, the differences observed between the varying slope and constant slope conditions when N = 250 were not found for N = 1,000.
RMSDs for Correlations (500 Persons)
Note. RMSD = root mean square deviation; IRT = item response theory.
RMSDs for Correlations (1,000 Persons)
Note. RMSD = root mean square deviation; IRT = item response theory.
Discussion
When studying relationships between estimates of traits and outcomes in the social and behavioral sciences, researchers have the option of using estimates of the variables based on CTT or IRT. While researchers have considered the different scaling properties of CTT- versus IRT-based scores, the debate surrounding the use of ordinal or interval scaled scores in analyses does not appear resolved.
The current study used a Monte Carlo approach to compare the behavior of trait estimates obtained from CTT versus IRT measurement frameworks and their use as predictor variables in validity studies. Conditions were specifically manipulated to evaluate the robustness of the different trait estimates under various conditions and assumptions underlying the CTT model: (a) CTT model weighs each item equally, which in turn implies each item is equally related to the underlying trait being measured; (b) CTT model assumes the response options have the same meaning for all items since the same response scale is used across the set of items.
Results indicated that IRT- and CTT-based scores (SS) were very comparable in terms of predicting outcomes (e.g., validity coefficients) and that the use of CTT-based scores appear robust to violations in assumptions underlying the use of summated scores. While the findings of comparability between scores from the two frameworks are consistent with previous research, they were somewhat surprising given the explicit manipulation of the assumptions underlying CTT-based scores. Results further indicate that SS scores may have an advantage over IRT-based scores for scales of short length (10 items), particularly for smaller sample sizes (N = 250). Thus, for applied researchers, there may be no practical disadvantage to using CTT-based scores in prediction studies, and the one major advantage is that researchers do not need to engage in analyses required to evaluate the assumptions underlying the use of IRT-based scores. Researchers should be reassured that many types of analyses using summated scores are also robust to the assumption of interval data (e.g., Dowling & Midgley, 1991; Glass et al., 1972).
Despite the finding of no practical advantage to IRT-based scores with ordinal scaled items, it may be possible to introduce conditions that generate more nonlinearity in the relationship between summated scores and the underlying trait being measured (Ferrando & Chico, 2007). Under such conditions, IRT-based scores may now have an advantage over summated scale scores. For example, items reflecting extremely high discrimination or sets of items measuring extreme levels of the trait should exhibit more nonlinearity. To test this prediction, a simulation condition was examined for highly discriminating items (values fixed at 6). For a 20-item test, N = 250, and a true validity coefficient equal to .3, RMSD values for predicted outcomes and correlations were very similar to the corresponding values in Tables 4 and 5. While nonlinearity may be a potential problem with binary items (e.g., Ferrando & Chico, 2007), nonlinearity appears to become even less relevant as the number of response categories increase. Finally, although it may be possible to introduce more extreme conditions to create an advantage to IRT-based scores, it is important to consider that such conditions are not typically associated with instruments used in practice.
In any simulation study, the results generalize only to the conditions that were studied and the context of the modeled population. Results from this study are intended to generalize to applications involving scales consisting of polytomous items (five-choice items). Further in that context the intent was to manipulate factors that would have practical value to applied researchers. While the conditions were intended to reflect practical values, it is possible that results do not generalize to other contexts.
Footnotes
Acknowledgements
The authors wish to thank the reviewers and editor for their comments and help in revising the article.
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
