Abstract
It is known that sum score-based methods for the identification of differential item functioning (DIF), such as the Mantel–Haenszel (MH) approach, can be affected by Type I error inflation in the absence of any DIF effect. This may happen when the items differ in discrimination and when there is item impact. On the other hand, outlier DIF methods have been developed that are robust against this Type I error inflation, although they are still based on the MH DIF statistic. The present article gives an explanation for why the common MH method is indeed vulnerable to the inflation effect whereas the outlier DIF versions are not. In a simulation study, we were able to produce the Type I error inflation by inducing item impact and item differences in discrimination. At the same time and in parallel with the Type I error inflation, the dispersion of the DIF statistic across items was increased. As expected, the outlier DIF methods did not seem sensitive to impact and differences in item discrimination.
DIF and Type I Error Inflation
Differential item functioning (DIF) and DIF identification methods have received much attention in the past decades. In the item response theory (IRT) framework, an item is said to function differently (or shortly to be DIF) if respondents from different groups but with the same value of their underlying latent trait (e.g., their ability or proficiency level) have nevertheless different probabilities of answering this item correctly. In other words, at equal ability levels the success probability of a DIF item differs among the groups of respondents. Such items may jeopardize the principle of fair measurement across the respondents and should therefore be identified. The topic is addressed in detail in several books and book chapters (Camilli & Shepard, 1994; Holland & Wainer, 1993; Osterlind & Everson, 2009; Penfield & Camilli, 2007) and in numerous journal articles (Bolt, 2002; Cohen & Kim, 1993; Clauser & Mazor, 1998; Finch, 2005; Finch & French, 2007; Magis, Béland, Tuerlinckx, & De Boeck, 2010; Millsap & Everson, 1993; Narayanan & Swaminathan, 1994; Penfield, 2001; Rogers & Swaminathan, 1993; Roussos & Stout, 1996).
A rather large number of methods has been developed to identify DIF items. These methods can be broadly classified in two categories: methods based on test scores (i.e., number-correct) as a matching variable and methods based on item response theory (IRT) models (Millsap & Everson, 1993). In the former category, the most well-known methods are Angoff’s Delta plot (Angoff & Ford, 1973), standardization (Dorans & Kulick, 1986), the Mantel–Haenszel (MH) method (Holland & Thayer, 1988), logistic regression (Swaminathan & Rogers, 1990), and SIBTEST (Shealy & Stout, 1993). The most well-known IRT-based DIF methods include Lord’s chi-squared test (Lord, 1980), Raju’s area approach (Raju, 1988, 1990), and the likelihood ratio test (Thissen, Steinberg, & Wainer, 1988).
Once an appropriate DIF detection method is selected, DIF analyses are most often conducted in two successive steps. First, a DIF statistic is computed on the basis of either the item responses and test scores (for score-based DIF methods) or calibrated item parameters (for IRT-based DIF methods). These DIF statistics are then compared with a DIF detection threshold and items are flagged as DIF whenever their related DIF values exceed that threshold. This first step relies mostly on statistical aspects and, therefore, well-established DIF statistics with known asymptotic distributions.
The second step consists in computing effect size measures. DIF effects can be statistically significant without the size of the effect is large enough to have practical relevance. For instance, large samples of respondents might lead to the statistical detection of tiny, yet significant DIF effects. Most DIF methods (but not all) have related effect size measures. For instance, the MH method has an effect size measure derived from the common log-odds ratio of the partial tables (Holland & Thayer, 1985), and the existence of the so-called ETS Delta scale permits to classify effect sizes as “negligible,”“moderate,” or “large” (see, e.g., Dorans & Holland, 1993; Holland & Thayer, 1988).
In this article, we focus on the first part of the DIF identification process, that is, on the statistical tests without further examination of the effect size measures. Ideally, statistical-based DIF methods have proper Type I error rates (i.e., false alarm rates). The reason for our interest is that several studies have shown that test score-based methods exhibit Type I error inflation (Candell & Drasgow, 1988; Clauser, Mazor, & Hambleton, 1993; Magis & De Boeck, 2012; Wang & Su, 2004), also in ideal situations when DIF is absent from the data (Meredith & Millsap, 1992; Zwick, 1990), which is of course undesirable for a statistical test. We want to investigate this inflation phenomenon in its purest conditions, namely, when there is no DIF, in other words given the null hypothesis is true, which is the proper way to determine the Type I error rate. Testing the Type I error rate in the presence of DIF items is of practical relevance indeed but it can be a complicating and even distorting factor for a good understanding of what is the case. Our primary purpose is to explain the Type I error inflation.
An obvious reason for something going wrong is that matching on the basis of test scores does not imply that there is matching on latent trait (or ability) levels (DeMars, 2010). Only if the test score is a fair proxy (or a sufficient statistic) for the latent trait, or if the true latent trait distributions are the same in all groups of respondents, can one expect that Type I error inflation does not occur. In other words, when the test score is not a sufficient statistic for the latent trait (such as when the items differ in discrimination) and the group ability distributions differ (such as with item impact and thus different group averages), then can one expect Type I error inflation from test score-based DIF methods, including MH (see also Bolt & Gierl, 2006). This explanation, however, does not uncover the underlying mechanism.
Our primary aim is precisely to understand the underlying mechanism and to test the implied explanation. Our secondary aim is to present and evaluate a solution. The presented solution is based on a different concept of DIF and a better understanding of the phenomenon. When the mechanism is better understood it is possible to propose a solution along the same lines.
Controlling Type I Error Inflation
Several approaches were proposed to overcome, or at least limit, the Type I error inflation effect. First, effect size measures (discussed earlier) can limit this inflation by discarding items that were flagged as DIF but have small effect sizes. Unfortunately, while effect size measures are available for some methods (such as MH or Lord’s methods), for some other approaches (such as Angoff’s Delta plot) there are not yet specific effect size measures. Another problem is that effect size criteria are not based on statistical testing although we want to follow a statistical reasoning approach given the statistical nature of the Type I error level criterion.
The second approach is often referred to as item purification (Candell & Drasgow, 1988). It consists in an iterative process wherein items previously flagged as DIF are removed from the computation of the test score and the DIF method is then reapplied with these newly, purified scores. The process ends whenever two successive iterations return identical classifications of items as DIF or non-DIF. Item purification was shown to be efficient in controlling Type I error inflation in absence and in presence of DIF (Clauser et al., 1993; Fidalgo, Mellenbergh, & Muniz, 2000; Lautenschlager & Park, 1988; Wang & Su, 2004; Wang & Yeh, 2003). Purification is applicable also for IRT-based methods (Candell & Drasgow, 1988). However, the iterative process is a heuristic and does not guarantee that it will converge. Purification is of practical value and it often converges in a few steps, but as mentioned it is a heuristic.
Interestingly, Kim and Oshima (2013) mentioned that Type I error inflation could also occur with long tests, due to the effect of not adjusting the process to multiple testing. They evaluated several alternatives to control for this inflation by multiple-testing adjustments, including the usual Bonferroni correction, Holm’s procedure (Holm, 1979), or Benjamini–Hochberg false discovery rate (Benjamini & Hochberg, 1995). These approaches basically rely on an accurate selection of DIF items based on the original DIF statistics and using a sequential procedure to select only items with lowest p values. The approach consists of a post hoc analysis of the DIF detection results arising from the original method (the latter being potentially influenced by Type I error inflation). All these three aforementioned methods consist of more than one step and they are not based on an explanation of the inflation.
Finally, based on the conceptual idea that DIF items are outliers in the space of all test items, Magis and De Boeck (2012) developed an outlier DIF approach consisting of two steps: (a) computing DIF statistics that are normally distributed under the null hypothesis of absence of DIF; (b) flagging items as DIF when their DIF statistic is an outlier compared with the other items. For Step (a) the MH log-odds ratio statistic was considered, as it is known to be normally distributed under the null (Magis & De Boeck, 2012; Penfield & Camilli, 2007). Among other advantages of this approach, the study highlighted that the Type I error rates were not inflated in possibly troubling situations as described above, and in situations with and without DIF items. It will be explained why this method takes into account the mechanism at the basis of the inflation.
Aims and Goals
The purpose of this article is to explain Type I error inflation with the common score-based DIF methods (without effect size measures, item purification, or post hoc analyses) and to evaluate this explanation as well as the outlier DIF approach as a solution. This is done by the variation of two factors that are assumed to be at the basis of the inflation. We focus on score-based DIF-methods, because Type I error inflation was mostly discussed in the context of such methods. Moreover, we will make use of the MH method, the gold standard for DIF (Wainer, 2010). We expect that the inflation shows for the MH method but not for the outlier DIF methods.
The explanation is that the item response functions (IRF) cross when there are differences in discrimination, such that the successes of lower ability respondents stem relatively more from low discriminating items than the successes of higher ability respondents. As a consequence, when the reference group and the focal group differ with respect to their mean level ability, the posterior success probabilities of items with a flat IRF is higher in the group with the lower ability than in the group with the higher ability, and the opposite is true for items with a steep IRF. As a consequence, conditioning on the sum score, some items seem relatively easier in the low group than in the high group and other items seem relatively more difficult in the low group.
The specific resulting expectation is that not the mean of the DIF statistic increases but that the variance does instead. This means that even in absence of DIF, the absolute values of the DIF statistic will become much larger than expected, so that based on a fixed detection threshold (which is commonly used with traditional score-based DIF methods), more items will be flagged as DIF. In other words, the Type I error inflation stems from the increased variance which in turn stems from the simultaneous realization of two conditions: differences in discrimination and a difference in average ability. The variance and thus the inflation should increase when one of the two increases while keeping the other constant, on the condition that the other condition is met to at least some degree. The outlier DIF methods should not show the effect because they do not depend on the variation of the DIF statistic but only on the outlyingness of its values.
The next section sketches succinctly the two DIF methods under consideration, the MH and the outlier DIF approaches, before describing the design of a confirmatory simulation study and presenting the related results.
Mantel–Haenszel Method and Outlier Approach
The following framework is considered. Differential item functioning is investigated among two groups of respondents, the reference group R and the focal group F. Respondents from these groups were assigned a test of n dichotomously scored items. Possible test scores therefore range from 0 to n, and let
Mantel–Haenszel DIF Statistic
The DIF statistic to be considered in this article is a variant of the so-called Mantel–Haenszel estimate of the common odds ratio (MH-OR)
The logarithm of this statistic (1), referred to as the MH log-odds ratio (or MH log-OR) DIF statistic
is asymptotically normally distributed under the null hypothesis of absence of DIF (Holland & Thayer, 1988; Penfield & Camilli, 2007). More precisely, the standardized version of the MH log-OR statistic, simply referred to as the MH-DIF statistic
has a standard normal distribution under the null hypothesis. The denominator of (3) is the square root of the variance of the
Hence, the standard MH method that is considered in this article is applied as follows: (a) compute the partial tables for each observed test score and with respect to item
Outlier Approach
Steps (a) to (c) are the same for the outlier DIF approach, making use of the same DIF statistic that is normally distributed under the null hypothesis, such as the MH-DIF Statistic (3). The outlier method differs in that the criterion for identification is not fixed but varies with the spread of the statistic across items and depends on the outlyingness of its value. DIF items are outliers compared with the other items. This idea rejoins the definition of DIF that underlies the Delta Plot method (Angoff & Ford, 1973). In practice, the MH-DIF statistics are transformed into standardized z scores as follows:
where
with
A possible problem is that the sample mean and standard deviation are not robust (in its statistical sense) with respect to outliers. Magis and De Boeck (2012) therefore proposed to replace the estimates of the mean and standard deviation by robust alternatives, the median
This outlier DIF approach, based on both the non-robust Rule (5) and the robust Rule (6), was shown to be unaffected by Type I error inflation (Magis & De Boeck, 2012), neither in the presence nor in the absence of DIF, in contrast with the common MH Method (3).
Simulation Study
To test the expected effects for the common MH method and the absence of the effects for the outlier DIF equivalents, a simulation study is conducted. Tests of 20 items were generated for 1,000 respondents in the reference group and 1,000 respondents in the focal group. Large samples of respondents were generated to make sure that asymptotic conditions related to the null distribution of DIF statistics can be assumed to be satisfied (remember that we focus on the statistical testing aspects). Item parameters were generated under a two-parameter logistic (2PL) model. Two conditions were systematically varied: the size of item impact and the variability in item discrimination levels. More precisely, item difficulties of the first 10 items were selected from the regular sequence from −2 to 2 with steps of 0.444, and the same sequence was used for the last 10 item difficulties. Two different sequences of item discriminations were selected. For the first 10 items, the discrimination levels were obtained as the regular sequence from
Item impact was introduced as follows. Ability levels in the reference group were drawn from the standard normal
For a selected couple of
Results
First, the logit of the Type I error rates of the standard MH method are displayed in Figure 1, for increasing values of the

Logits of Type I error rates of the Mantel–Haenszel method, for various sizes of item impact (
The Type I error rates were obtained from the 2,000 MH-DIF statistics per simulation setting. These DIF statistics are displayed in Figure 2 as boxplots, with each of the four panels corresponding to one value of item impact.

Distributions of MH-DIF statistics per size of item impact (
Three main conclusions can be drawn. First, the distributions are all very symmetric around their central values. Second, as expected the means are almost not affected, the central values do not depart much from zero, which is the expected average value under the null hypothesis of absence of DIF. This is true for all three values of impact different from zero. Third, the MH-DIF statistics become more dispersed (i.e., the variances of the statistics increase) with increased item impact and increased variability in item discriminations.
Table 1 summarizes the sample means and standard deviations of all 24 boxplots in Figure 2. The values reported in Table 1 obviously confirm the last two conclusions above. One can observe a slight increase in the sample means when the variability in item discriminations increases, and this increase is more marked with larger item impact. The slightly increasing means are due to the fact that the degree of discrimination is a multiplicative parameter and the geometric mean deviates from the arithmetic mean, which is of course not the case if the degrees of discrimination are equal (strictly speaking an assumption of the MH method). Nevertheless, these mean values remain quite stable around their theoretical value of zero. On the other hand, the sample variances are close to one when either
Sample Means and Standard Deviations of the Simulated MH-DIF Statistics, for Various Sizes of Item Impact (γ Parameter) and Dispersions of the Discrimination Levels (τ Parameter).
Finally, the Type I error rates of the outlier DIF approach are reported in Table 2, both for the non-robust rule (5) and the robust rule (6). These Type I errors are remarkably stable across the different item impact sizes and differences in item discriminations, separately for the two rules. No effect on the Type I error rate can be observed with these methods. The robust rule yields larger rates than the non-robust rule, and for the latter the rates are very close to the nominal significance level. This was already observed by Magis and De Boeck (2012). The larger error rates for the robust rule can be explained by the fact that the robust estimator of dispersion (the MAD) often returns smaller values than the sample standard deviation, yielding smaller detection thresholds and hence, larger Type I error rates. This is especially the case when there are no outliers, as in this simulation study without DIF items.
Type I Error Rates of the Outlier DIF Approach, With the Non-Robust and the Robust Rules, for Various Sizes of Item Impact (γ Parameter) and Dispersions of the Discrimination Levels (τ Parameter).
Discussion
The simulation study yields two major conclusions. First, as expected, the standard MH method is affected by Type I error inflation in situations where both the discriminations vary across items and the ability levels are different in the two groups of respondents. In line with the explanation of the inflation, the variance of the DIF statistic increases consistently with the two conditions. Second, and again as expected on the basis of the explanation, the outlier DIF variant is unaffected by Type I error inflation even in extreme cases of large item impact and large variability in the item discriminations. In sum, the explanation of the Type I error inflation is confirmed by the simulation study, and a method that is based on the explanation does not show the inflation.
It is clear from Figure 2 what the diagnosis is for the Type I error inflation with the MH method. The MH method makes use of a constant DIF detection threshold, the quantile of a standard normal distribution, which depends on the significance level but neither on the test length nor on the sample sizes. One can easily understand the effect of this “fixed” threshold on the computation of related Type I error rates, as was illustrated by Figure 2. With the usual 5% significance level, all items whose DIF statistics lie outside of the range [−1.96, 1.96] are flagged as DIF by the standard MH-DIF statistic. Due to the increase in dispersions of the MH-DIF statistics, the proportions of items lying outside this range increase, and so do the Type I error rates.
In contrast with the standard MH method, the outlier detection method based on the same statistic adjusts for the sample location and dispersion of the MH-DIF statistic. Taking for instance the non-robust rule (5), it is obvious that in case of larger variability in the DIF statistic, the sample variance
This study suggests that DIF methods can be classified in two categories: (a) methods with fixed DIF detection thresholds based on a prespecified statistical distribution (such as MH and logistic regression) or based on some empirical argument (such as standardization or Delta Plot); and (b) methods for which the thresholds are derived from the sample distribution of the item DIF statistics (such as the outlier DIF variant). These two categories are further referred to fixed threshold and data-driven threshold methods, respectively. In the present framework, data-driven methods outperform the fixed threshold methods for the reasons described in the previous sections. The other methods we have discussed for controlling the Type I error inflation are fixed threshold methods. They adjust the threshold based on effect size, on a heuristic (purification) or on a multiple-test criterion. They are not based on data-driven thresholds.
It is worth mentioning that the stability of the Type I error rates was also established in at least one other recently improved DIF method starting from the Delta plot method. The Delta plot method (Angoff & Ford, 1973) was among one of the first proposed DIF methods and is based on the computation of pairs of Delta scores (i.e., standardized item scores per group of respondents). These pairs, referred to as the Delta points, are displayed in a scatter plot and form an ellipsoid, the Delta plot, for which the major axis can be easily computed. The related DIF statistic is the perpendicular distance between each Delta point and the major axis of the Delta Plot. The larger the distance, the further the Delta point lies from the major axis and the higher the probability is for an item to exhibit DIF. Angoff and Ford (1973) did not suggest any precise threshold for these distances, but many different proposals have been formulated (see Magis & Facon, 2012, for a short review). All those thresholds, however, are fixed values derived from empirical results or personal choice. Magis and Facon (2012) proposed an improvement by considering the set of Delta points as arising from a bivariate normal distribution, in line with the shape of the Delta Plot. This leads to a data-driven threshold, involving both the sample variances of the Delta scores in each group and the covariance between them (Magis & Facon, 2012). This modified Delta plot method was compared with its original form and the simulation study highlighted very stable Type I error rates in absence of DIF for the modified Delta plot, whatever the setting of the study. Again, this improvement is based on deriving a DIF detection threshold from the set of DIF statistics themselves, in line with the outlier DIF approach. In sum, the modified Delta plot is another example of data-driven DIF method, whereas its original form, the Delta plot, is a fixed-threshold DIF method.
In the present framework, data-driven DIF methods obviously outperform the MH method for the reasons discussed in the previous sections. This does not imply that data-driven methods should be preferred overall. For instance, the modified Delta plot was shown to be most effective in case of small samples, whereas its global performance was not better than the MH method with larger samples (Magis & Facon, 2012). The outlier DIF method was shown to become slightly conservative in the presence of DIF, exhibiting smaller Type I error rates than the significance level in conjunction with an increase of power (Magis & De Boeck, 2012). Thus, beyond their intrinsic assets, the data-driven DIF methods are of primary interest for controlling Type I error inflation in absence of DIF.
Data-driven DIF methods are not the only solution to controlling for Type I error inflation, but it is an elegant approach. It is a one-step procedure, it does not need further assumptions, and it makes use of existing DIF statistics. Effect size measures require a second step and are based on consensus. Item purification is an iterative approach. Multiple testing corrections such as Holm’s or Benjamini–Hochberg procedures are post hoc analyses relying on the output of the test-score methods. Data-driven methods, as characterized in this article, can provide stable Type I errors in a single run of the method. Although it would be highly valuable from a practical point of view to compare all available methods, this is beyond the scope of the present study, not just because this would require a very large simulation study. Our scope was conceptual in the first place. The simulation study was conducted to test a theoretical explanation and a method that is based on the explanation. Finally, this article only focused on Type I error inflation in absence of DIF. Some studies (Candell & Drasgow, 1988; Clauser et al., 1993; Magis & De Boeck, 2012; Wang & Su, 2004) also highlighted a similar phenomenon when DIF is introduced in the data, although the related effect on power is not clearly established. It would then be very useful to perform a similar investigation by introducing DIF in the generated data sets and examine for potentially interesting effects on the power values, since Type I error inflation is expected to occur also in this framework.
Footnotes
Acknowledgements
The authors wish to thank George A. Marcoulides, Editor, and two anonymous reviewers for their helpful comments.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was funded by a postdoctoral research grant “Chargé de recherches” of the National Funds for Scientific Research (FNRS, Belgium), the IAP Research Network P7/06 of the Belgian State (Belgian Science Policy), and the Research Funds of the KU Leuven, Belgium.
