Abstract
The purpose of this simulation study was to establish general effect size guidelines for interpreting the results of differential bundle functioning (DBF) analyses using simultaneous item bias test (SIBTEST). Three factors were manipulated: number of items in a bundle, test length, and magnitude of uniform differential item functioning (DIF) against the focal group in each item in a bundle. A secondary purpose was to validate the current effect size guidelines for interpreting the results of single-item DIF analyses using SIBTEST. The results of this study clearly demonstrate that ability estimation bias can only be attributed to DIF or DBF when a large number of items in a bundle are functioning differentially against focal examinees in a small way or a small number of items are functioning differentially against focal examinees in a large way. In either of these situations, the presence of DIF or DBF should be a cause for concern because it would lead one to erroneously believe that distinct groups differ in ability when in fact they do not.
Keywords
The issue of fairness in educational testing has been a concern for researchers, practitioners, and the public for some time now. An unfair test is one that disadvantages one group of examinees over another group of examinees, although both groups are equal in ability. One way in which researchers and practitioners have addressed the issue of fairness in educational testing is through differential item functioning (DIF) analyses. These analyses allow one to determine whether equal ability examinees from two distinct groups, referred to as the focal (i.e., female) and reference (i.e., male) groups, differ in their probability of answering a single test item correctly (Angoff, 1993). DIF studies are often conducted in an exploratory fashion where each item is assessed in turn using the remaining test items as the matching subtest. Under this approach, substantive explanations as to why particular items function differentially for two distinct matched groups are postponed until after the items have been statistically flagged. An alternative approach is to conduct differential bundle functioning (DBF) analyses based on meaningful explanations for why certain bundles of items function differentially for distinct groups of examinees with equal ability (Douglas, Roussos, & Stout, 1996). In this approach, substantive hypotheses are developed a priori to explain why bundles of items are functioning differentially and these hypotheses can then be tested using simultaneous item bias test (SIBTEST; Shealy & Stout, 1993). This approach has great potential for helping psychometricians understand cultural and cognitive reasons for the occurrence of DIF (K. Banks, 2006; Roussos & Stout, 1996; Walker & Beretvas, 2001; Walker, Zhang, & Surber, 2008).
Specifically, the use of SIBTEST in DBF research has led to a greater understanding of DIF because of differences in accommodation status (Finch, Barton, & Meyer, 2009), culture (K. Banks, 2006), gender (Douglas et. al., 1996; Gierl, Bisanz, Bisanz, & Boughton, 2003; Gierl, Bisanz, Bisanz, Boughton, & Khaliq, 2001; Mendes-Barnett & Ercikan, 2006; K. E. Ryan & Chiu, 2001; K. E. Ryan & Fan, 1996), language spoken (Abbott, 2006, 2007; Gierl & Khaliq, 2001), reading ability (Walker et al., 2008), and writing ability (Walker & Beretvas, 2001). SIBTEST has emerged as a quintessential method for conducting DBF research because of its (a) ability to detect DIF or DBF based on a researcher’s hypotheses, (b) ability to distinguish between genuine DIF/DBF and DIF/DBF due to impact (true mean ability differences between groups), and (c) use of a regression correction technique that deals with the problem of Type I error due to impact (Shealy & Stout, 1993).
However, one issue that permeates the DIF research where SIBTEST is concerned is that, while there are guidelines for interpreting the size of DIF in individual items (see Roussos & Stout, 1996), there are presently no guidelines for interpreting the size of DIF in bundles of items. Researchers have been calling for the development of effect size guidelines in SIBTEST for some time now, and no one has taken up the challenge. Fifteen years ago, Douglas et. al. (1996) stated,
a still further research question concerns how to judge the size of DIF displayed by a bundle. It is not clear how the typical criteria used in judging the level of item DIF could be extended to judging the level of bundle DIF or whether totally new criteria should be developed. (p. 466)
Four years ago, Abbott (2007) stated, “since researchers do not agree upon how to distinguish statistical from practical significance in DBF research, future research investigating and developing guidelines for interpreting bundle effect size measures is necessary” (p. 27).
It seems logical, given the impetus for DIF and DBF analyses, that any effect size guidelines for DIF detection procedures should take into consideration the impact of DIF and DBF on ability estimation bias which is independent of the power of the test statistic that is influenced by sample size. Today, test items are routinely examined to determine if they function differentially against certain groups of examinees prior to their use in any high-stakes large-scale testing endeavor. Items that are found to function differentially are typically revised or removed from the test altogether. However, despite attempts to ensure that tests are fair to all examinees, certain groups continue to lag behind others in their performance on high-stakes large-scale tests. Consider the National Assessment of Educational Progress (NAEP) as an example. English language learners (ELLs) lag behind non–English language learners (non-ELLs) on NAEP at Grades 4, 8, and 12 in reading, math, and science. In 2009, the largest difference in average scale score between ELLs and non-ELLs occurred at Grades 8 and 12. At Grade 8, the average scale score for ELLs was 50 points less than that for non-ELLs in science. At Grade 12, the average scale score for ELLs was 50 points less than that for non-ELLs in reading. Likewise, Blacks and Hispanics lag behind Whites on NAEP at Grades 4, 8, and 12 in reading, math, and science. However, the difference in average scale scores is more pronounced for Blacks. In 2009, the largest difference in average scale score between Blacks and Whites occurred at Grade 8 in science where the difference was 36 points. 1 These findings inevitably beg the question of whether such disparities are caused by a series of items that are functioning differentially against the disadvantaged group, reflect true differences in achievement, or are the result of something else entirely.
Obviously, there is not one definitive reason why the achievement gap persists, but DIF analyses were first undertaken by the psychometric community in the 1960s in response to the public’s concern that achievement tests were unfair to minority examinees (Angoff, 1993). The logic behind this effort was that if some test items focused on content that minority examinees were unfamiliar with then these examinees would be less likely to answer such items correctly even though they possessed the content knowledge measured by the test. The fact that the achievement gap continues despite meticulous test construction by test developers and item writers has resulted in a myriad explanations for the lingering achievement gap, other than the presence of DIF/DBF.
For example, when considering the achievement gap that persists for different racial groups it has been argued that “stereotype threat,” which exists when a person’s social identity is attached to a negative stereotype, may cause minority examinees to underperform in a manner consistent with the stereotype (Jordan & Lovett, 2007; K. E. Ryan & Ryan, 2005; Sackett, Hardison, & Cullen, 2004; Steele, 1997). It has also been argued that the cultural mismatch between teachers, the majority of whom are middle-class White females, and students contributes to the achievement gap (Irvine, 2002, 2003; Irvine & Collison, 1999; Lee & Majors, 2003). Also, teachers’ pedagogical practices can work to exacerbate or narrow the achievement gap (Cochran-Smith, 2003; Ladson-Billings, 1994, 2000; A. M. Ryan, 2006; Sleeter, 2001). Still another argument is that the very nature of the curriculum, which is believed to reflect the values, ideals, and perspectives of White middle-class Americans, has served to widen the achievement gap. Scholars who ascribe to this line of thinking posit that the achievement gap can be reduced through the adoption of a multicultural curriculum (J. A. Banks, 1995, 2006; Gay, 2003).
When considering the achievement gap that persists for ELL students, several studies have examined whether large-scale content area tests administered in English are valid measures of academic performance for ELL students (see Mahoney, 2008; Martiniello, 2008, 2009; Ockey, 2007; Snetzler & Qualls, 2000; Young et al., 2008). In each of these studies, although the mean score performance for ELL students was always significantly lower than that for non-ELL students, no systematic DIF against ELL students could be found. This suggests that the construct-irrelevant variance is external to the test. ELL students may face difficulty on large-scale tests because of their limited English vocabulary and inexperience with the English language. Research has shown that if a reader is to have any success at comprehension, he or she must know 95% of the printed words (Laufer, 1989; Liu & Nation, 1985). This 95% includes 2,000 high-frequency items (87% of a given text) plus 800 terms from the University Word List (8% of a given text). Such a criterion is hard to attain for ELL students whose English vocabulary tends to be much smaller than that of native English speakers (August, Carlo, Dressler, & Snow, 2005). However, rather than urge test developers to simplify the language on large-scale tests, educators must find ways to increase the English vocabulary of ELL students. Educators must also work hard to teach ELL students how to navigate through the linguistic features they are likely to face on large-scale tests which include but are not limited to passive voice constructions, conditional clauses, long noun phrases, and relative clauses.
Given that there are a plethora of explanations for the lingering achievement gap, including the presence of DIF/DBF, coupled with the fact that the current effect size guidelines for interpreting the magnitude of DIF in single items when using SIBTEST is based on rules developed at ETS more than 20 years ago (Zwick & Ercikan, 1989), it is imperative to establish more current effect size measures for SIBTEST that take into consideration the impact of DIF/DBF on ability estimation. Therefore, the primary purpose of this simulation study was to establish general effect size guidelines for interpreting the results of DBF analyses using SIBTEST. A secondary purpose was to validate the current effect size guidelines for interpreting the results of single-item DIF analyses using SIBTEST.
Overview of SIBTEST
SIBTEST (Shealy & Stout, 1993) uses a nonparametric statistical procedure to detect unidirectional DIF or DBF in dichotomous items. Unidirectional DIF occurs when the direction of DIF is constant across the ability scale. However, it should be noted that a version of SIBTEST known as Crossing-SIBTEST can be used if it is hypothesized that there is an interaction between ability and DIF such that an item functions differentially in favor of one group at lower levels of ability and functions differentially in favor of the other group at higher levels of ability.
DIF detection methods compensate for impact, or true mean ability differences between the two groups for which an item is thought to function differentially, by matching examinees on some conditioning variable and summarizing differences in item or bundle performance across the different levels of the conditioning variable. However, the presence of impact can still inflate the Type I error rates of DIF detection methods because the expected value of the reference group’s target ability will tend to differ from the corresponding expected value for the focal group (Shealy & Stout, 1993). SIBTEST uses a nonlinear regression correction to correct for this inflated Type-I error rate (Jiang & Stout, 1998).
Specifically, SIBTEST is based on the assumption that DIF occurs when
where Y denotes the score on the studied item or bundle of items, T denotes the true score on the conditioning variable, and the subscript R or T refers to either the reference or the focal group, respectively. The conditioning variable used in SIBTEST is a matching subtest that consists of a set of items hypothesized to be unidimensional and therefore only measuring the primary construct of interest. The studied item is hypothesized to be multidimensional, measuring a secondary dimension on which the two matched groups are believed to differ, in addition to the primary construct of interest. The assumption outlined in Equation (1) implies that if ER [Y|T]−EF [Y|T] ≈ 0, then the studied item or set of items does not exhibit DIF or DBF. Denoting this difference by B(T), a global index of DIF can be obtained by
where fF(T) denotes the probability density function of T in the focal group. Therefore, the statistical procedure used by SIBTEST is H0: β = 0 versus H1: β ≠ 0. The global index of DIF is approximated by
where n is defined to be the number of items in the matching subtest, pk denotes the proportion of examinees who obtained a score of k on the matching subtest, and
where
The test statistic corresponding to this global DIF index is defined by
where
where
There has been extensive prior research investigating the influence of sample size on power and Type I error rates when using SIBTEST to detect DIF/DBF. However, it should be noted that this line of research is not particularly relevant to the current study because effect size measures should not be influenced by sample size. Whereas previous research, pertaining to the impact of sample size on power and Type I error rates, has focused on counting up the proportion of times the test statistic correctly (or incorrectly) rejected the null hypothesis, the present study focuses on how the presence of DIF/DBF affects the actual magnitude of
Roussos and Stout (1996) found that even with as few as 250 examinees in the reference and focal groups, Type I and Type II errors were not inflated to a large degree, when impact was not present. With such small sample sizes, the Type I error rate was inflated slightly only in the presence of impact; however, this issue was only found with certain types of items (e.g., those with high discrimination and low difficulty). In terms of the effect of sample size on power, previous research has demonstrated that SIBTEST displays adequate power (slightly less than 0.8) when there are 500 examinees present in each group (Gierl, Gotzmann, & Boughton, 2004). Finch (2005) found that the greatest influence on power, when using SIBTEST to detect DIF, was the size of the focal group; smaller focal groups yielded lower power than focal groups with larger sample sizes.
Method
Given the fact that the focus of this study was on creating effect size guidelines for DBF detection when using SIBTEST, as opposed to studying Type I error rates and power when doing DBF analyses, ideal conditions were simulated. Specifically, as previous research has demonstrated that impact may increase the Type I error rate, no impact was simulated. This ensured that the most ideal conditions were simulated and that the results would not be influenced by the well-known effect of impact on increasing the Type I error rate. No impact was simulated by randomly generating the true ability for both reference- and focal-group examinees from a standard normal distribution.
Similarly, sample size was also not considered in this study because prior research has shown that an increased sample size may increase the power of SIBTEST, whereas a decreased sample size may decrease the power of SIBTEST (Furlow, Ross, & Gagne, 2009; Gierl et al., 2004; Roussos & Stout, 1996). Likewise, the ratio of focal-group examinees to reference-group examinees was not considered in this study because prior research has clearly indicated that having an unequal ratio of focal- and reference-group examinees is only an issue when the overall sample size is small, less than 250 in the reference and focal groups (Furlow et al., 2009). Therefore, this study only included ideal sample size conditions of 500 examinees in both the reference and focal groups. A sample size of 500 in each group is considered more than adequate for controlling the Type I error and power rates (Clauser & Mazor, 1998).
Factors in this simulation study included the manipulation of (a) the number of items in a bundle (1, 3, or 5) that were simulated to function differentially, (b) test length (20 or 40 items), and (c) the magnitude of uniform DIF against the focal group in each item in a bundle. Of particular interest was the sum of the magnitude of DIF across items when testing for DBF. It was hypothesized that the magnitude of
Design of Simulation Study Conducted
Note. DIF = Differential item functioning.
Item parameters were randomly generated for each replication to ensure that the results were as generalizable as possible. Specifically, the discrimination parameter (a) was randomly generated from a log-normal distribution with a mean of 1.0 and a standard deviation of 0.5, the difficulty parameter (b) was randomly generated from a standard normal distribution, and the lower asymptote (c) was randomly generated from a beta distribution with a mean of approximately 0.16 and a standard deviation of approximately 0.004. When an item was simulated to function differentially, the item was always simulated to be more difficult for the focal group by adding the magnitude of the DIF to the difficulty parameter used for the reference group (e.g., bfoc = bref + magnitude of DIF).
SIBTEST was used to test for DIF/DBF. To prevent from biasing the results, the matching subtest was always created by using only those items that were not simulated to function differentially. In other words, the matching subtest was always DIF-free (Ackerman & Evans, 1994; Gierl, 2005). One-tailed tests were conducted in all cases to examine whether the studied items were functioning differentially against the focal group, which is reflective of the direction for which DIF was simulated.
An additional step was included in this study. The additional step consisted of obtaining estimates of ability for the reference and focal groups in an effort to better understand the effect of DIF/DBF on ability estimation. This step was taken to determine how much DIF or DBF an item or bundle of items needed to contain before the ability estimates for the focal examinees became seriously biased. The ability estimates were obtained for the reference and focal group members simultaneously. This ensured that the estimates were on the same scale for each group. T tests were conducted to determine if the ability estimates were statistically significantly different because of the presence of DIF/DBF. The t tests were conducted on the average ability estimates obtained for the reference and focal groups for each of the various conditions examined. In addition, bias, defined as the average signed difference between true and estimated ability, was used to summarize the impact of DIF/DBF on ability estimation.
Results
One-Item DIF Analyses
Figure 1 illustrates the results obtained from the simulations conducted to examine the behavior of

Average

Power when using SIBTEST to test for DIF for all one-item conditions studied
Multi-Item DBF Analyses
Figure 3 illustrates the average

Average beta-uni statistics obtained for all DIF/DBF sums studied, grouped by two different test lengths studied
Impact of DIF/DBF on Ability Estimation
Figure 4 illustrates ability estimation bias as a function of the sum of DIF, under all conditions studied. As expected, test length has an effect on ability estimation bias, with more bias associated with shorter tests. However, what are more interesting are the different patterns of ability estimation bias obtained for the reference and focal groups. It was expected that as the presence of DIF or DBF increased there would be increased negative ability estimation bias for focal-group examinees. However, what Figure 4 illustrates is that as the sum of DIF increased there was increased negative ability estimation bias for focal-group examinees, as well as increased positive ability estimation bias for reference-group examinees. In other words, whereas the ability of the focal group was expectedly underestimated, the ability of the reference group was surprisingly overestimated. This effect is more pronounced for shorter length tests and for larger bundles of items that are functioning differentially.

Estimation bias obtained when using SIBTEST to test for DIF and DBF across all conditions studied
So, does this estimation bias result in statistically significant differences in ability estimation for reference and focal groups? This question seems to get at the heart of establishing DIF or DBF effect size guidelines for SIBTEST because, after all, it is overall performance differences on achievement tests which initiated the development of DIF detection procedures (Angoff, 1993) and, despite all our attempts to ensure that tests are fair to all examinees through DIF and DBF analyses, the achievement gap still persists today. Figure 5 shows the rejection rates obtained when comparing the average estimated ability for the reference and focal groups under various DIF/DBF conditions studied using simple t tests. As the figure illustrates, the relationship between finding statistically significant differences in the average ability estimates for the reference and focal groups is highly dependent on two things: (a) the sum of DIF or DBF present in the items and (b) test length. In other words, it is the proportion of DIF or DBF, defined as the sum of DIF divided by the number of test items, that is indicative of whether mean ability estimates will statistically differ for reference- and focal-group examinees, not simply the presence of DIF or DBF. In addition, as the sum of DIF increases, having more items in a bundle will lead to finding statistically significant differences in average ability estimates for the reference and focal groups more often, as long as the proportion of DBF is high.

T test rejection rates obtained when comparing the average estimated ability for the reference and focal groups under various DIF/DBF conditions studied
So how can these results be used to establish effect size guidelines for SIBTEST? In the conditions explored in this study,
then
Regression Model Obtained When Regressing Sum of DIF and Number of DIF Items on Beta-Uni Statistic From SIBTEST
Note. DIF = differential item functioning; SIBTEST = simultaneous item bias test; SE = standard error. Dependent variable: Beta-uni statistic.
Dividing the Sum of DIF by the total number of items on the test will provide an estimate of the proportion of DIF which can be used as an estimate of how much the DIF or DBF will affect ability estimation. However, this estimate needs to be interpreted according to the number of items that are found to be functioning differentially, given that greater ability estimation bias is associated with having more items in a bundle (see Figure 3). According to the results obtained from this study, when the proportion of DBF is 0.15 (i.e., 3/20, where 3 = the sum of DIF and 20 = the total number of items on the test, which is the highest proportion of DIF/DBF considered in this study) then statistically significant ability differences will be found between reference- and focal-group examinees about 60% of the time for a five-item bundle and about 40% of the time for a three-item bundle, which is likely cause for concern. Moreover, this trend seems to increase exponentially, as illustrated in Figure 5.
How do these results compare with the Roussos and Stout (1996) effect size guidelines for SIBTEST which state that
Real-Data Examples
To demonstrate how these results might be applied in a practical situation, the results of two previous studies that used a DBF analytic framework were considered. In the first study, published by Walker and Beretvas (2001), Poly-SIBTEST was used to determine if open-ended items on a mathematics test that required examinees to explain their answers in writing functioned differentially in favor of examinees who were proficient in writing. In this study, the average
In the second study, published by Walker et al. (2008), SIBTEST was used to determine if mathematics items that required a higher reading load functioned differentially in favor of proficient readers. In this study, the highest
Discussion
The primary purpose of this study was to determine effect size guidelines when using SIBTEST to conduct DBF analyses. This was accomplished by linking the original impetus for conducting DIF analyses—the fact that many believed that large-scale tests were unfair to minority students—to ability estimation since the achievement gap has lingered despite attempts by psychometricians to ensure that tests are fair to all by conducting DIF/DBF analyses. Obviously, there is not one definitive reason as to why the achievement gap persists, but the results of this study clearly demonstrate that only with very high levels of DIF or DBF, in terms of the proportion of DIF/DBF with respect to the total number of items on the test, can these disparities be caused by a series of items that are functioning differentially against the disadvantaged groups. This finding suggests that scholars must begin to look outside the statistical properties of the test in an effort to develop explanations for the lingering achievement gap. Ability estimation bias can only be attributed to DIF or DBF when a large number of items in a bundle are functioning differentially against focal examinees in a small way or a small number of items are functioning differentially against focal examinees in a large way. In either of these situations, the presence of DIF or DBF should be a cause for concern because it would lead one to erroneously believe that distinct groups differ in ability when in fact they do not. Given that large-scale test items are routinely examined for DIF, it is unlikely that this would occur in practice. It may be that the differential test performance between distinct groups is signaling a much broader and dire problem with the American educational system that extends far beyond the test.
It should be noted that the findings of this study are somewhat contingent on an adequate sample size and equal reference- and focal-group examinees. Future research is needed, in terms of investigating how overall sample size and the ratio of reference- and focal-group examinees affect the magnitude of
Finally, it is unclear why the well-known effects of amplification were only partially supported. The results of this study suggest that amplification only occurs with a large magnitude of DIF/DBF. Although this was unexpected, it was not the primary focus of this study. Therefore, future research is needed in this area to determine under what conditions amplification will occur in practice.
Footnotes
Acknowledgements
The authors would like to thank the reviewers for their helpful feedback and Mark Gierl for reading an earlier version of this article.
Earlier versions of this article were presented at the Seventh Conference of the International Test Commission, Hong Kong, and the 2011 Annual Conference of the National Council for Measurement in Education, New Orleans, LA.
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
