Abstract
In my reaction to the major contribution by Mallinckrodt, Miles, and Recabarren, I endorse their recommendation that the use of item response theory (IRT) be increased. Advantages of IRT are numerous, with most resulting from the fact that IRT models typically take a more realistic view of how item responses are related to underlying traits than classical test theory. However, I raise concerns regarding their advocacy of the Rasch IRT model, which arguably takes an overly simplistic view of how items and traits are related. Many alternative IRT models exist, including ones based on the ideal-point measurement philosophy (most IRT models use the dominance model). I recommend that researchers seek to determine which IRT model best fits their items—including the larger question of whether a dominance versus ideal-point approach is preferable—and avoid assuming that one model will always perform the best.
In addition to the views presented in the major contribution by Mallinckrodt, Miles, and Recabarren (2016 [this issue]) to which I am offering this reaction, numerous researchers have advanced the position that we should move beyond methods based on classical test theory (CTT), and instead adopt methods based on item response theory (IRT) when developing assessment instruments (e.g., Drasgow & Hulin, 1990; Embretson, 1996; Hambleton, Swaminathan, & Rogers, 1991; Harvey & Hammer, 1999; Hulin, Drasgow, & Parsons, 1983; Lord, 1980; Samejima, 1979; Thissen & Steinberg, 1985; Wright, 1977). Clearly, Mallinckrodt et al. are to be commended for their efforts to help move the field of counseling psychology in particular—and the entire field of psychology in general—toward this important goal.
Despite the fact that researchers have been advocating a change to IRT for decades (e.g., Lord, 1980; Wright, 1977), progress toward achieving this goal has been slow. One reason is likely the fact that IRT models tend to be considerably more mathematically complex than CTT, and until relatively recently, the software needed to apply these models was complex to configure and difficult to use. Indeed, I have long wondered whether I will see the day when IRT finally surpasses CTT as the dominant approach used in psychological measurement.
On the positive side, there are numerous advantages to making the switch to IRT. These advantages include the ability to (a) use computer-adaptive testing techniques to enhance item security while reducing administration time and maintaining precision (e.g., Waller & Reise, 1989); (b) detect the presence of inappropriate assessment scores (due to faking, lack of adequate language skills, etc.) via quantitative appropriateness indices (e.g., Drasgow, Levine, & McLaughlin, 1987); (c) determine the degree to which individual test items operate in a biased fashion with respect to subgroups of examinees via differential item functioning (DIF), or to detect differential test functioning (DTF) at the total-score level (e.g., Hambleton et al., 1991; Hulin et al., 1983); (d) obtain more precise estimates of examinees’ trait scores that make use of more of the information contained in the item responses than is possible using CTT’s number-right scoring (e.g., Hambleton et al., 1991; Hulin et al., 1983); and (e) in general, use measurement models that more accurately reflect the potentially complex relationships that exist between the latent constructs we seek to measure and the item responses we can actually observe (e.g., Carter et al., 2014; Embretson, 1996; Samejima, 1979).
This last issue is especially significant in light of the ongoing debate regarding the merits of the widely used dominance model of measurement (e.g., Likert, 1932; which forms the conceptual basis for the Rasch and older IRT models) versus the alternative ideal-point approach (e.g., Carter et al., 2014), on which more recent IRT models are based (e.g., Roberts, Donoghue, & Laughlin, 2000; Stark, Chernyshenko, Drasgow, & Williams, 2006). A growing body of research (e.g., Carter et al., 2014; Drasgow, Chernyshenko, & Stark, 2010) suggests that ideal-point methods may consistently be superior to methods based on the dominance approach (including the process recommended by Mallinckrodt et al., 2016) when measuring the types of noncognitive individual differences traits (e.g., personality, interests, values, attitudes) that are of widespread interest to counseling psychologists.
Rasch Limitations
Given that this article represents a reaction to the major contribution of Mallinckrodt et al. (2016), one might correctly expect that I will offer at least a somewhat different view of the topic at hand than the one they advanced. Let me first stress that I strongly agree with their overall takeaway point (i.e., we should abandon CTT and move to IRT-based methods when developing and revising new assessment instruments), as well as their suggestions regarding the usefulness of the focus-group methodology for item writing, and the importance of searching for DIF on the basis of relevant demographic factors.
However, I raise concerns regarding two issues: (a) the technical limitations of Rasch IRT models, which have been widely criticized in terms of adopting an overly simplistic view of how latent traits are related to observed item responses and (b) the more general tendency of many researchers to dogmatically favor a single philosophy of measurement and apply it in all assessment-development contexts, paying little or no attention to determining the degree to which it provides a realistic representation of how latent traits are related to assessment items in that situation. These issues are considered in the following sections.
Binary Rasch Model
The Rasch model (e.g., Wright, 1977) was initially developed to deal with binary item responses (e.g., the right/wrong scoring of items on ability and achievement assessments). Binary item response data can also be produced in noncognitive assessments from multiple-choice ratings that are not scored in a right/wrong fashion. For example, many personality and interests assessments (e.g., the Myers–Briggs Type Indicator [MBTI]; see Harvey & Hammer, 1999) present respondents with pairs of words or statements, and ask them to pick the one that most accurately describes them. Such ratings can be dummy coded (0/1) such that a “1” response is scored if the alternative chosen reflects the keyed (high) pole of the trait (e.g., in the word pair “prefer quiet vs. love meeting strangers,” giving a “1” if the scale is one in which Extraversion is the keyed direction, and the person chose the “love meeting strangers” option).
Although it is not difficult to find enthusiastic proponents of the Rasch model (e.g., Wright, 1977), and to identify certain areas in which it still remains popular (such as for some types of educational achievement and licensure-based assessment; for example, Wu, West, & Hughes, 2008), many psychometricians view the binary-scored Rasch model as being something that is largely of historical interest, due to its suffering from significant limitations as a result of the overly simplistic view it takes regarding the form of the trait–item relationship (e.g., Drasgow & Hulin, 1990; Embretson, 1996; Hambleton et al., 1991; Hulin et al., 1983; Lord, 1980). That is, the binary Rasch model only estimates a single index on which items differ: difficulty or b parameter.
As Figure 1 in Mallinckrodt et al. (2016) illustrates, every item in an assessment that is calibrated using this type of Rasch IRT model has an identically shaped item characteristic curve (ICC), and items differ only in their left–right location on the trait axis (higher difficulty as you move to the right). This implies two things that many psychometricians find to be fundamentally untenable in many assessment situations: namely, that (a) all items are identical with respect to how strongly they are related to the underlying trait (i.e., discrimination, or the slope of the ICC at its point of maximum inflection, denoted by the a parameter) and (b) no processes such as guessing (for a right/wrong item) or social desirability (for a personality or interest item) would cause some people who have very low true scores on the trait to still get the item right (via successful guessing) or endorse the keyed pole (in the socially desirable direction for a personality or interest item), denoted by the “pseudo-guessing” nonzero-lower-asymptote parameter c.
Obviously, successful guessing often occurs on multiple-choice right/wrong items, and many individuals choose to present themselves in a more socially desirable fashion than others when completing self-report personality inventories. Likewise, no matter how hard you try, in real-world situations, the items in an assessment usually vary considerably in terms of their discriminating power and strength of relationship to the underlying trait (e.g., as seen in item-total correlations computed in CTT or factor loadings in a multifactor inventory). In such cases, the Rasch model is by definition fundamentally misspecified, and it cannot provide an accurate description of the true ways in which observed items relate to the underlying latent trait.
The Mallinckrodt et al. (2016) article illustrated the use of a Rasch IRT model for polytomous items (i.e., ordered-category responses, as in a Likert-type agree–disagree scale), so it is important to note that the concerns noted previously specifically focus on the binary model. However, they are discussed here given that (a) the binary Rasch model is seen in practice at a higher rate than the polytomous variant, (b) Mallinckrodt et al. discussed the claimed advantages of the binary Rasch model at length, and (c) they effectively took an advocacy position (p. 157) in “the Rasch Wars” (McNamara & Knoch, 2012) by recommending the Rasch model over more realistic IRT models that include discrimination and asymptote parameters. Accordingly, I conclude that if researchers are given the choice of using a model that is more complex to estimate (but that accurately describes the item–trait relationships) or easier to estimate (but that forces an overly simplistic view of how items function onto the data), they should avoid the misspecified model when developing new assessment instruments.
Rasch proponents do suggest a strategy for dealing with the fact that their model does not fit many actual assessment items: throw out the items that do not fit their theory of how items should operate (a strategy included in the process recommended by Mallinckrodt et al., 2016). Although in one sense this does address the problem, I and many others hold the opposing view that when theories are shown to be clearly inconsistent with the data, you do not ignore the data, you modify the theory to be able to explain all of the existing empirical data relevant to it.
Polytomous Rasch Model
Regarding the Mallinckrodt et al. (2016) recommendations for using the graded-response Rasch model (see Appendix A, and pp. 162-167), although their procedures are appropriate as far as they go, the complexity of the process (including the discarding of misfitting items) begs the question as to whether a simpler method could be achieved by choosing a more general IRT model. That is, many of the iterative steps they describe in which items are removed from the pool (e.g., based on factor loading differences) are things that might not be necessary if a more flexible or realistic IRT model had been chosen instead of the polytomous Rasch model.
For example, multidimensional IRT (MIRT) models exist (e.g., see Brown & Maydeu-Olivares, 2013; Hartig & Höhler, 2008) to deal with instruments that measure multiple traits. Rather than discarding potentially large numbers of items due to the multidimensionality seen in the iterative exploratory and confirmatory factor analyses recommended by Mallinckrodt et al. (2016), the more complex nature of the traits (using all the items) could be modeled in a MIRT solution. In contrast, if the goal is to develop a unidimensional scale, researchers could select from more general polytomous IRT models that offer the ability to represent more complex item–trait relations, without the need for discarding misfitting items (e.g., Roberts, 2001; Roberts et al., 2000). Constrained special cases of the generalized graded unfolding model (GGUM; Roberts et al., 2000) exist to deal with many different types of items, including models based on constant-unit, multiple-unit, rating scale, and partial-credit assumptions.
In sum, I am not questioning the fact that the relatively complex process recommended by Mallinckrodt et al. (2016) may produce acceptable results. However, I am suggesting that by adopting a more general and flexible IRT model, researchers might be able to make a valuable trade-off in terms of reducing the complexity of the process and having less of a need to discard items due to their failure to adhere to the restrictive assumptions of the Rasch model.
Avoiding Dogma When Selecting Measurement Models
No Point in Taking Rigid Sides in the “Rasch Wars”
Mallinckrodt et al. (2016) correctly noted that researchers tend to be highly polarized regarding the Rasch model (e.g., McNamara & Knoch, 2012). That is, some enthusiastically extol its virtues and recommend it in all situations, whereas others conclude that its model is so inconsistent with the way real-world items operate that it should be avoided at all costs.
However, rather than taking the Rasch “side” and advocating the use of the simplest model available (binary or polytomous Rasch), as Mallinckrodt et al. (2016) did—or by advancing the opposite dogma that Rasch should never be used in any setting—I recommend that researchers instead focus on selecting the best tool for the given assessment situation. Start with a general IRT model that is capable of representing a wide range of ways in which latent traits relate to items (e.g., three-parameter for binary data, GGUM for polytomous), and then move to simpler IRT models if the results indicate that the added complexity is not needed.
Clearly, if by chance one happens to create an item pool where all items have identical discriminating power, and all ICCs have a zero lower asymptote, then by all means use a Rasch model. However, as the “discard all items that do not fit the theory” approach used by Rasch advocates suggests, finding such a situation in practice is unlikely.
Given that for most of the types of assessments used by counseling psychologists we would expect to encounter items that differ considerably in their discrimination parameters, and that these items may exhibit nonzero lower ICC asymptotes (e.g., due to varying social desirability), my recommendation is to “hold the Rasch” when picking an IRT model to use, and start the process by using the most general IRT model appropriate to the types of items in the assessment (e.g., see Roberts et al., 2000; Thissen & Steinberg, 1985). Simply put, it is a lot easier to start with a general model and add restrictions to simplify it (while still accurately modeling the true item–trait relationships) than it is to start with a model that is a gross oversimplification of reality and iteratively throw out all the items that do not fit it.
If at the end of the process one is left with the Rasch model, so be it. However, especially for the types of noncognitive assessments that are popular among counseling psychologists, in my instrument-development experience (e.g., Harvey & Hammer, 1999), it would be highly unusual to find that the items written to form new assessments exhibit identical discrimination and zero lower asymptotes.
Dominance Versus Ideal-Point Models
A final area of concern regarding the need to avoid dogmatism involves an even bigger measurement-model question than the “Rasch wars” issue. Specifically, should we develop assessments based on dominance model assumptions (e.g., Likert, 1932) that form the conceptual basis for the Rasch and many IRT models, or do we embrace the ideal-point model (e.g., Carter et al., 2014) that reflects the Thurstonian approach (e.g., Thurstone, 1927, 1928, 1929)? These two models embody fundamentally different ideas as to how latent traits and item responses are related.
The dominance model embodies the logic that underlies many IRT models (e.g., see Figure 1 in Mallinckrodt et al., 2016) with respect to the shape of the ICC. That is, as the scores of examinees on the latent trait (x axis) increase, we see a monotonic increase in the probability that people who score at that level will answer the item correctly (right/wrong) or in the keyed direction. This function starts low (at zero, in the Rasch model) and increases until it reaches upper asymptote (where all people get the item correct, or give the keyed response).
In contrast, the ideal-point model (for details, see Carter et al., 2014; Roberts et al., 2000) is based on the position that a person will rate the degree to which a statement is accurate in describing them (for binary or polytomous Likert-type scales) based on how closely their true score on the trait matches the location parameter of that item (i.e., its left–right location on the trait scale). The bigger the difference between the location of the item and the person’s true location on the trait (in either direction), the lower they will rate its accuracy in describing them.
For items formed as statements that describe the high or low extremes of a bipolar scale (a common practice in noncognitive assessments), IRT methods based on the ideal-point and dominance models will produce similar results (i.e., the monotonically increasing ICC seen in Figure 1 of Mallinckrodt et al., 2016). However, for items that are less-extreme in nature—that is, describing things that people in the intermediate region would endorse—the methods diverge, with the ideal-point approach producing an item response function that exhibits an inverted U (i.e., increasing at first, peaking, and then decreasing at higher scores). For example, on an Introversion–Extraversion scale, an item such as “I like to socialize, but there are times that I prefer to be alone” may receive high accuracy ratings for people who score in the intermediate range, but lower ratings from people whose scores lie toward the high or low extremes (i.e., because high introverts may not like to socialize at all, and high extraverts want to socialize all the time).
A growing body of research (e.g., Carter et al., 2014; Drasgow et al., 2010; Stark et al., 2006) indicates that the ideal-point model may offer a consistently better way to measure noncognitive psychological traits than the dominance model. For example, using the GGUM model, Zimmermann et al. (2015) found that the ideal-point approach was effective in measuring dimensions in the Diagnostic and Statistical Manual of Mental Disorders (5th ed.; DSM-5; American Psychiatric Association, 2013) “Criterion A” impairments in the personality functioning domain. Carter et al. (2014) found that the ideal-point method was superior with respect to measuring aspects of conscientiousness, especially when using such scores to predict performance outcome variables.
In sum, I endorse the recommendation offered by Mallinckrodt et al. (2016) that we hasten the move from CTT to IRT-based instrument development, and that we should routinely search for (and discard) items that exhibit significant demographic DIF. However, I take issue with their conclusion that we should rely on Rasch models when doing so and that we should follow a complex, iterative process to remove items that fail to conform to that model’s restrictive view of how items should relate to latent traits. Rather, the binary or polytomous Rasch models should only be used if the results of fitting a more general IRT model indicate that their additional complexity is not needed to faithfully model how the items actually perform.
In addition, in light of recent research (e.g., Carter et al., 2014; Drasgow et al., 2010; Stark et al., 2006) indicating that in many situations we can obtain better measurement by moving beyond traditional dominance models (on which most IRT methods are based, including Rasch) and adopting the ideal-point approach, we should actively pursue this alternative method when developing new assessments. Significantly, this choice has fundamental implications for the ways in which we write items. That is, when writing items under the ideal-point approach, we seek to produce items that individuals who score in the “average” range of the scale will rate as being highly accurate in describing themselves. In contrast, when writing items using the dominance model, we seek to identify statements that people who lie at the high or low extremes of the scale will endorse as being highly accurate.
Footnotes
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
