Abstract
This study extends the multilevel measurement model to handle testlet-based dependencies. A flexible two-level testlet response model (the MMMT-2 model) for dichotomous items is introduced that permits assessment of differential testlet functioning (DTLF). A distinction is made between this study’s conceptualization of DTLF and that of differential bundle functioning (DBF) with an emphasis on the practical utility of decomposing differential functioning into item- and testlet-specific components. A small-scale simulation study is also conducted to assess estimation of the MMMT-2 model and its measures of DIF and DTLF when compared with SIBTEST’s estimation of DIF and DBF. Results support use of the MMMT-2 model for assessing impact, DIF, and DTLF for tests that include testlet-based dichotomous items with equal item discriminations.
Given the focus on assessment that has been increasing since the institutionalization of the No Child Left Behind policy and the evidence-based practice movement, it is of the utmost importance to use models that can evaluate the potential noninvariance of item and testlet scores. Scores used to make decisions about, for example, students and teachers must be fair and measure the same construct in the same way across groups. One of the first steps in assessing the fairness of scores entails the assessment of differential item functioning (DIF), differential bundle functioning (DBF), and differential testlet functioning (DTLF).
Identification of DIF in a dichotomous item indicates that, for examinees of the same ability, the probability of a correct response differs by group (e.g., gender, ethnicity, etc.) and thus there is the potential that the item is biased against one of the groups. Similarly, DBF indicates that, for a set of (here) dichotomous items, the probabilities of a correct response to the items in the bundle differ by group, after conditioning on ability. If DBF is identified, then substantive analysis is necessary to identify whether the differential functioning of the bundle is construct relevant or construct irrelevant in order to infer bias or not. The differential functioning associated with DBF is summarized across the items that form the bundle. Thus, DBF contains differential functioning that is common across items in a bundle summed together with differential functioning that might be unique to individual items in a bundle.
There are different reasons to bundle items together. The reason of interest here is that items might be bundled together within a testlet. Thus, the testlet would constitute the bundle. The current study is in part designed to introduce a conceptualization of DTLF that is distinct from DBF. Like DBF, DTLF for a set of dichotomous items in a testlet indicates that, conditional on ability, the probabilities of a correct response to the items contained in a testlet differ depending on group membership. However, although DBF represents the sum of the direction and degree of differential functioning of each item and across the bundle’s items, DTLF (as defined here) represents only the differential functioning that is common across the set of items in the testlet. Under this conceptualization, when the bundle of interest is a testlet, then DBF is a function of the testlet’s (potential) DTLF plus each item’s (potential) DIF. Breaking down a testlet’s DBF into the part that is common in direction and magnitude across the testlet’s items (i.e., DTLF) and the parts that are unique to each item (i.e., DIF) permits a more nuanced assessment of the source of the differential functioning in a bundle (here, testlet) and the potential for item and/or testlet bias.
For example, it is possible that DTLF might be found favoring girls over boys (after controlling for math ability) in their performance on a math word problem that requires particularly strong reading comprehension skills for the common stimulus. It would be up to those interpreting the DTLF to decide whether this particular differential functioning results in “biased” item scores or not. The analyst might consider the reading comprehension skill as complementary to the math ability being assessed (i.e., as construct relevant), and therefore although DTLF is identified, the testlet would not be considered as biased against boys. Alternatively, it is possible that the intent of the math test was to derive a pure math ability score. If this were the case then the word problem’s interference in the testlet’s item functioning might be interpreted as a nuisance dimension and thus the testlet items would be considered biased.
Imagine further a situation in which the same DTLF was detected for a testlet that consistently favored girls over boys given the (assumed here) better reading comprehension skills of girls and yet DIF was found in one item within the testlet that favored boys over girls. For example, if the item exhibiting DIF contained some information concerning content of more interest to boys than to girls (such as football content) then performance on that particular item might favor boys over girls of the same math ability. It seems likely that most test developers would not consider knowledge of football as a construct that should contribute to a measure of math ability and therefore the item would be considered biased although the reading comprehension skill-based DTLF might not be inferred as unfair. However, if DBF were used to quantify the differential functioning of item scores in the testlet, then (depending on the magnitudes of the DIF and DTLF) it might appear that there is no DBF because of the cancellation of the DIF by the opposite direction of the DTLF effect. Remember that DBF will be the sum of the unique and common differential functioning and the two might cancel each other out. Similarly, estimation of DIF using any of the previous DIF identification methods would also lead to an inference of no differential functioning because again the DIF effect could be canceled out by the DTLF. Estimation of the model suggested here that separates out DTLF from DIF will quantify both the DIF and the DTLF and will identify that the effects are working in opposite directions. Only if differential functioning is identified can decisions be made about whether each of the (item and testlet) sources of the differential functioning has resulted in bias or not. The current study is designed to make the case for the use of a measure of DTLF and to distinguish it from DBF.
Although a lot of research has led to the introduction and assessment of several different DIF indicators (e.g., Finch, 2005; Holland & Thayer, 1988; Raju, 1988; Rogers & Swaminathan, 1993; Shealy & Stout, 1993), there has been less such research on measures of DTLF. One commonly used indicator of DTLF is SIBTEST’s (Shealy & Stout, 1993) measure of DBF. A lot of research has supported use of SIBTEST for this purpose (see, e.g., Banks, 2006; Gierl, Bisanz, Bisanz, & Boughton, 2001; Gierl & Khaliq, 2001; Ryan & Chiu, 2001; Walker & Beretvas, 2003; Walker, Zhang, & Surber, 2008). However, the idea that DBF is the sum of DIF and DTLF does not seem to have been sufficiently clarified in previous research (with the exception of Wainer, Sireci, & Thissen, 1991). Thus, although the primary focus of the current study is to introduce a new measure and conceptualization of DTLF, the current study also includes a small-scale evaluation of the estimation of the suggested DTLF measure and empirically demonstrates the distinctions between DTLF and DBF (assessed using SIBTEST).
The new measure of DTLF is parameterized using the multilevel measurement model (Beretvas & Kamata, 2005) and thus incorporates the added flexibility possible under the multilevel modeling framework. Various researchers have noted how the generalized linear mixed model can be used to obtain Rasch item response model parameters (e.g., Adams & Wilson, 1996; Adams, Wilson, & Wu, 1997; Cheong & Raudenbush, 2000; Fischer, 1995; Kamata, 2001). When used in this way, the resulting generalized linear mixed model has been termed the multilevel measurement model (MMM; Beretvas & Kamata, 2005). Extensions to these MMMs have already been introduced that permit researchers to assess item and person parameters for either dichotomous or polytomous items on unidimensional or multidimensional measures while simultaneously assessing multiple sources of DIF (e.g., Cheong & Raudenbush, 2000; Kamata, 2001; Luppescu, 2002; Meulders & Xie, 2004; Van den Noortgate & De Boeck, 2005; Williams & Beretvas, 2006).
Another primary benefit of the MMM is that the model can be extended to include additional levels to handle dependencies resulting from clustered data structures. For example, Pastor and Beretvas (2006) added a level to the MMM to handle the dependency resulting from repeated measures (Level 1) on items (Level 2) within people (Level 3). Dependency can also be a direct result of a dataset’s structure. For example, if a data set consists of multiple students per sampled school (e.g., students’ clustering within schools), then a three-level MMM can be modeled with item scores (Level 1) clustered within students (Level 2) within schools (Level 3; Kamata, 1999). Alternatively, the cross-classification of students by, for example, middle and high school could be modeled using a three-level cross-classified MMM (Beretvas, Meyers, & Rodriguez, 2005). However, the MMM can also be used to handle another source of dependence, specifically dependence that results from the clustering of items within testlets. And although an extension to the MMM has already been suggested for use with handling testlet-based dependence (Jiao, Wang, & Kamata, 2005), the current study suggests an alternative MMM parameterization for handling testlet data that permits concurrent assessment of testlet-specific DTLF.
Testlet Response Models
Conventional testlet response models
If a set of items is linked by a common stimulus (such as by a reading comprehension passage or through a common mathematics data set), then the items are considered to be a part of a “testlet” (Wainer & Kiely, 1987). Responses to items within a testlet can no longer be assumed locally dependent, thereby violating a fundamental assumption made with standard unidimensional item response models. A number of conventional testlet response theory (TRT) models have been suggested that each include testlet effect parameters to account for the dependency. Much like the corresponding dichotomous item response models (e.g., the one-, two-, and three-parameter logistic models), these TRT models are differentiated by the number of item parameters used.
Bradlow, Wainer, and Wang (1999) introduced a two-parameter TRT model. Wainer, Bradlow, and Du (2000) extended Bradlow et al.’s (1999) model and suggested use of a three-parameter TRT model that included a pseudoguessing parameter for each item. Li, Bolt, and Fu (2006) extended the model further and derived a more flexible parameterization of the three-parameter TRT model that permits unique item discriminations on both the general ability factor and on each testlet ability factor.
W. Wang and Wilson (2005) proposed the use of a Rasch version of Wainer et al.’s (2000) three-parameter TRT model in which item discriminations are assumed equal and the pseudo-guessing parameter was constrained to zero:
where pij represents the probability of a correct response to dichotomous item i for person j, bi is the item discrimination, and θ j is the person ability parameter. γ d ( i ) j represents the testlet effect of item i in testlet d for person j. In TRT models, there are typically two person factors or abilities being modeled. These factors include the general ability, θ j , and the testlet-specific ability, γ d ( i ) j , factors. Inclusion of the testlet parameter in the model in Equation 1 models the dependence of items that are linked by a common testlet. W. Wang and Wilson (2005) listed the benefits of the Rasch version of the model including observable sufficient statistics as well as pointing out that the less parameterized (Rasch) model can be well estimated with smaller sample sizes.
All these TRT models involve the assumption that the general ability, θ, and testlet-specific abilities, γs, are independently and normally distributed. Under Bradlow et al.’s (1999) two-parameter TRT model, the variances of each testlet-specific ability distribution are assumed constant across testlets. Thus, under Bradlow et al’s model, only two random effects variance components need to be estimated—one for the variance of the θs and one representing the common variance of the γ d ( i ) j s for each testlet. For the one-parameter TRT model (see Equation 1) and the three-parameter TRT model, however, a unique variance component is estimated for each testlet-specific ability factor (for each γ d ( i ) j ) in addition to the variance of examinees’ general ability (i.e., the variability in the θs).
Three-level multilevel measurement model for testlets
In addition to these conventional TRT models, Jiao et al. (2005) suggested an extension to the MMM for handling testlet-based dependencies. The authors suggested a Rasch-based multilevel measurement model for testlets for dichotomous items that consisted of three levels (MMMT-3). Item scores (Level 1) are modeled as clustered within testlets (Level 2) that are nested within examinees (Level 3). More specifically, the model is at Level 1:
where pidj is the probability of a correct score on item i (for i = 1, 2, . . ., k) in testlet d for person j and Xqidj is item q’s indicator. Xqidj is dummy-coded with a value of negative one if i = q and zero, otherwise. (Note that Jiao et al.’s, 2005, model has been slightly modified here such that it no longer includes the fixed intercept term). At Level 2, the model is
and at Level 3 the model is
(Note that to facilitate explaining the correspondence between parameters in the multilevel measurement models with those in conventional TRT models, the conventional use of γ and β in Raudenbush and Bryk’s (2002) levels formulation has been reversed in the multilevel model in Equations 3 and B.)
Combining the levels’ formulations for Jiao et al.’s (2005) model into a single equation, the probability of a correct response to item i in testlet d for person j is
where u00 j and γ i 00 correspond to the person ability θ j and item difficulty bi in W. Wang and Wilson’s (2005) conventional Rasch-based TRT model (see Equation 1). And the conventional Rasch TRT model’s testlet effect, γ d ( i ) j (see Equation 1), corresponds with the Level 2 residual term, r0 dj . In addition, like conventional TRT models, Jiao et al.’s model also involves the assumption that the testlet and general ability factors are independent. Jiao et al.’s MMMT-3 model, however, involves the assumption that the testlet effects’ variances are the same across testlets.
As noted, there are two primary benefits associated with using any of the suggested MMMs and including the MMMT-3 for obtaining item and person parameters. First, extra levels can be added to the model to handle potential dependencies in the data structure. Thus, if the data set of interest consisted of item scores on a test that incorporated testlet items for students in schools, then a fourth level could be added to the MMMT-3. Clearly, however, estimation of a model with four levels becomes onerous in terms of both interpretation and estimation. The second benefit of the MMM results from the ease with which variables representing potential sources of DIF and impact can be added to the model (see, e.g., Beretvas & Kamata, 2005). This latter benefit also applies to the MMMT-3. Addition of a person-specific predictor to the equation (see Equation 4) for the intercept (e.g.,
Two-level multilevel measurement model for testlets
Consider a test consisting of mq items with m testlets each consisting of q items. The Level 1 MMM equation modeling the log-odds of a correct response to item i for person j for the mq-item test consisting of m testlets is then
which looks very like the conventional MMM (see Equation 2) where Xij is a dummy-coded item indicator and Tij as a dummy-coded testlet indicator with both X and T coded with “−1” for the relevant item and testlet, respectively. For each testlet of q items, there are only (q − 1) dummy-coded item indicators, and for the m testlets, there are m testlet indicators. Note that in Equation 6 each qth item (i.e., items q, 2q, . . . , mq) is used as the reference indicator for each testlet and thus the associated indicator does not appear in the equation. (It should be emphasized that while the MMMT-2 model’s parameterization provided here requires that every item be dichotomous, the model does not require that every item be associated with a testlet, nor that each testlet consist of the same number of items. In addition, the equal number of items per testlet example that is provided here is used solely to facilitate description of the model.)
At level two, the simplest MMMT-2 model is as follows:
Matching the assumption made in some other TRT models (e.g., Li et al., 2006; Wainer et al., 2000; W. Wang & Wilson, 2005), testlet and person abilities are assumed normally and independently distributed with means of zero and the following covariance formulation:
Note, however, that it is possible to model nonzero covariances among any of these effects and that provides an interesting extension for further research (see, e.g., Paek, Yon, Wilson, & Kang, 2009). However, conventional TRT models (see Equation 1, for example) assume that the general ability and testlet ability factors are uncorrelated and thus the same assumption was made here.
Combining Equations 6 and 7, the probability of a correct response to nonreference indicator item i of testlet d for person j is
Thus, the probability of a correct response is modeled as a function of the item’s actual difficulty,
However, here is one of the distinctions between this TRT model’s parameterization and previous parameterizations. Under the MMMT-2, the testlet effect for examinee j on testlet d is decomposed into the person-specific testlet random effect,
A nonzero fixed testlet effect,
Differential Testlet Functioning
In addition to decomposing the effect of a testlet into the testlet’s difficulty and the examinees’ ability on the construct specific to a testlet, the MMMT-2 model, unlike the MMMT-3 model (Jiao et al., 2005), can be easily extended to include an assessment of testlet-specific DTLF. Just as excessive variability across examinees in an item’s difficulty (after controlling for θ) indicates potential DIF, so excessive variability in a testlet’s difficulty (after controlling for θ and γ
d
(
i
)) indicates potential DTLF. Thus, if
with
Another benefit of using the MMMT-2 model’s parameterization is that it can be easily extended to simultaneously assess impact, DIF, and DTLF. For example, to model gender-based impact, DIF, and DTLF in Item 1 and Testlet 1, Equation 9 becomes
The coefficient, γ01, represents the degree of gender-based impact. Gender-based DIF in Item 1 is captured by γ11 and gender-based DTLF that affects all items in the first testlet is described by γ T 11. As noted, the MMMT-2 model’s parameterization permits separation of differential functioning that is common across items within a testlet (e.g., γ T 11 in Equation 10) from the differential functioning affecting a single item (e.g., γ11 in Equation 10). Use of this parameterization permits a more specific assessment of what might be at the source of differential functioning in an item in terms of whether it is occurring for an individual item (or some items) or from the common stimulus, or both.
It is possible to combine the MMMT-2’s DIF and DTLF parameters’ values to obtain an overall measure of the differential functioning across a testlet’s items corresponding with current conceptualizations of DBF. Thus, for a testlet consisting of q items, the amount of DBF for that bundle (where the testlet constitutes the bundle of interest in the current study) would be q times the testlet’s DTLF plus the sum of the amount of DIF identified for each item in the testlet. In the example contained in Equation 10 in which only the first item exhibited DIF and there was DTLF in its (the first) testlet, then the amount of DBF for that first testlet (bundle) would be quantified by
Conceptualization of DBF partly originated to address DIF amplification concerns (Douglas, Roussos, & Stout, 1996; Nandakumar, 1993). The idea of DBF matches exactly what was termed “differential testlet functioning” by Wainer et al. (1991). Specifically, some small degree of true DIF might exist for each item within a bundle; however, the magnitude of this DIF could be so small that there is insufficient power to detect the differential functioning at the item level (Douglas et al., 1996). DBF provides the sum of the DIF in the items that are being bundled together. Therefore, if the source of the bias is consistent in its direction across a bundle of items then pooling the items together into a bundle and examining DBF results in more power to detect the differential functioning. However, if the direction of the DIF in items within a bundle is not consistent then cancellation of the differential functioning may occur. In a differential functioning cancellation scenario, DIF favoring one group for one set of items in a bundle could be “canceled” by DIF favoring the other group on other items in the bundle (Nandakumar, 1993; Wainer et al., 1991). A resulting assessment of DBF might reveal that overall the items are not functioning differentially in the bundle. And although this might hold summatively for the bundle, this might not be true of individual items within the bundle.
Conceptualizing differential functioning cancellation in the testlet context, it is also possible that there might be DTLF across a testlet’s items favoring one group (say, girls) whereas within the same testlet some items’ content might favor the other group (e.g., boys). The resulting DTLF (favoring girls) might be canceled out by the items’ DIF (favoring boys) resulting in no DBF being detected. However, it is the source of the differential functioning (be it DTLF or DIF) that needs to be interpreted when decisions are being made about potential score bias. It is possible that the DTLF is construct relevant (or “benign”) whereas the items’ DIF might be construct irrelevant (“adverse”) or vice-versa. If the source of the differential functioning is construct irrelevant, then the item (or testlet) score should be considered biased. However, if the DBF assessment leads to inferences of no differential functioning, then no judgments can or would be made and no items could be appropriately edited to remove bias. Thus, it would seem useful for those interested in test construction and validation to understand more fully the contribution of each component of a test (e.g., the passage common across the testlet’s items and the structure and content of each item itself) to each item’s functioning even if these effects might cancel each other out.
Other measures of DBF do not permit this decomposition of an item’s functioning into the unique item versus common-across-the-testlet components. For example, SIBTEST’s DBF statistic, β
DBF
(Douglas et al., 1996), is one of the more commonly used tests of DBF. The statistic,
where
where propk is the proportion of focal group examinees with a score of k (out of K scores) on the ability measure and
Given SIBTEST’s DBF provides a summary of the DIF across items within the bundle (here, testlet), it would be expected that, when there is differential functioning cancellation, SIBTEST’s estimate of DBF will mask items’ opposing (positive for some and negative for others) patterns of DIF. And if the testlet itself is providing the cancellation effect (i.e., an item might favor boys and the associated reading passage’s content favors girls), then the DBF again will not indicate that there is any differential functioning within the testlet and, in addition, SIBTEST’s DIF will also not capture the canceled differential functioning.
Last, it is possible to request that the SIBTEST software program produce the results of multiple analyses at once. For example, test statistic results for DIF and DBF in multiple items and bundles can be provided in a single SIBTEST output. However, although this might give the impression to the naive user that DIF and DBF are simultaneously being modeled and assessed by SIBTEST, in fact each analysis is being conducted separately.
As demonstrated in Equation 11, the MMMT-2 model can provide a summary of the overall differential functioning in a testlet (i.e., the DBF); however, the contribution of the MMMT-2 model is that it can be parameterized to provide a breakdown of the differential functioning in an item’s score into the parts specific to each item (DIF) and the part common across the items (the DTLF). The MMMT-2 model then permits a more specific decomposition of the (item and testlet) factors that influence an item’s overall difficulty. If cancellation were occurring in the directions of the differential functioning, this should be manifested by differing directions of the MMMT-2 model’s parameter estimates for DIF and DTLF although this would not be clear from SIBTEST’s DBF test. In addition, while the current study is primarily designed to introduce the MMMT-2 model to simultaneously test for DIF and DTLF without being confounded by cancellation, a secondary objective of this study was to assess how well the MMMT-2 model’s parameters are recovered under various scenarios. Thus, a small-scale simulation study was conducted that included situations in which DIF only and in which DIF and DTLF were present in the simulated data. In some of the conditions, the differential functioning was generated to be in the same direction to assess amplification scenarios. In other conditions, the differential functioning was simulated to be in opposing directions. This design permitted demonstration of the differential functioning captured by SIBTEST’s
Method
A small-scale simulation study was conducted to assess DTLF and DIF identification using the MMMT-2 model and to demonstrate the distinction between SIBTEST’s βDBF and βDIF statistics and the relevant MMMT-2 parameters. Data were generated to fit each of the seven scenarios listed in Table 1. The scenarios differed in terms of the number of items for which DIF was generated, whether DTLF was generated, and the direction of the DIF and/or DTLF. Positive DIF or DTLF (represented as positive values for the relevant γ value) meant that the item (or testlet) was easier for the reference group than for the focal group, after controlling for ability differences. Negative DIF or DTLF meant that item functioning favored the focal group. Only items in the first testlet were generated to have DIF in scenarios in which DIF was simulated. Only zero, one or two items were generated to exhibit DIF. Only the first testlet was modeled as exhibiting DTLF (for conditions with DTLF).
Generating Values for Parameters Manipulated in Simulation Study by Condition
Note. DIF = differential item functioning; DTLF = differential testlet functioning.
The scenarios that were examined included a baseline condition in which no DIF or DTLF was generated (Conditions 1-4). Another set of four conditions (Conditions 5-8) were generated in which only one item (arbitrarily, Item 1) was generated to have a moderate amount of positive DIF with no DTLF generated. Conditions 9 through 12 were simulated to have positive DIF in two items with no DTLF. In conditions 13 through 16, data were generated to have positive DIF in Item 1 and positive DTLF for the first testlet (see Table 1). Conditions 17 through 20 had positive DIF for the first two items and positive DTLF for their testlet. To assess recovery of cancellation DIF, in Conditions 21 through 24, positive DIF was simulated for Item 1 and negative DIF for Item 2 with no DTLF generated. Conditions 25 through 28 were generated in which positive DIF was generated for Item 2 and negative DTLF for the associated testlet providing differential functioning cancellation where an item’s DIF is in the opposite direction to its testlet’s DTLF (rather than in a direction opposite to another item’s DIF). Thus, this last pair of scenarios provided an explicit direct test of how well SIBTEST’s βDIF and βDBF and the MMMT-2 formulations of DIF and DTLF capture two forms of differential functioning cancellation.
In conditions in which DIF or DTLF was simulated, the differential functioning was generated to have a magnitude of 0.4 (see Table 1). This mimicked a moderate amount of differential functioning reflecting a difference of 0.4 in the reference and focal groups’ item (or testlet) difficulties. Within each scenario listed in Table 1, two design conditions were explored. The two design factors were fully crossed resulting in four types of data sets within each scenario. The first factor was the variance of Testlet 1’s effects (
Only one test length was examined, here, consisting of 50 items with 5 items within each of the 10 testlets. This matches values used in subsets of conditions investigated in other TRT simulation research (Bradlow et al., 1999; Demars, 2006). Only one total sample size (n = 2,000) with equal reference and focal group sample sizes was investigated here for this small-scale simulation study.
Data were generated assuming the MMMT-2 model with the conditions’ values for γ11, γ21, and γ T 11 (see Table 1) substituted into Equation 10. Values for the reference and focal groups’ item difficulties for Items 6 through 50 (the valid subtest across conditions) were equal and randomly sampled from a standard normal distribution. Item difficulties for Items 3, 4, and 5 were set to (γ T 10), (γ T 10 − 1), and (γ T 10 + 1) for both the reference and focal groups. Item difficulties for Items 1 and 2 were set to the condition’s value of γ T 10 for the reference group across conditions. Item difficulties were set to the relevant condition’s value of (γ11 + γ T 10 + γ T 11) and (γ21 + γ T 10 + γ T 11) for Items 1 and 2, respectively, for the focal group. Examinee ability (on the primary dimension assessed by the test), u0 j , was sampled from a standard normal distribution for each examinee j. Person-specific testlet ability, uTdj, for each testlet was also sampled from an independent normal distribution with a mean of zero and variance of 0.25, 0.5, and 0.75 for Testlets 2 through 4, 5 through 7, and 8 through 10, respectively. These testlet effects’ variance values mimicked those simulated in other TRT simulation research (e.g., Bradlow et al., 1999; Jiao, Wang, & He, 2008; Li et al., 2006; W. Wang & Wilson, 2005). Person-specific testlet effects for testlet one, uT1 j , were sampled from a normal distribution with a mean of zero and a condition-specific variance of τ T 1.
Values for the testlet and person residuals and for the condition’s parameters were then combined to obtain the probability of a correct response to the relevant item. This expected probability was then compared with a number sampled from a uniform distribution with a [0, 1] range for each simulee and item. If the expected probability exceeded the randomly sampled value, then a score of one was assigned for that item and person. This was done for each of the 50 items and 2,000 examinees in the replication’s data set.
A full MMMT-2 model was estimated to assess DIF in Items 1 and 2 as well as DTLF for the first testlet. (Note that this model matched Conditions 17 through 20 as listed in Table 1. However, the model was overparameterized for the remaining conditions.) SIBTEST was also used to assess potential DIF in each of Items 1 and 2 and for DBF in Testlet 1 for each of the 28 conditions. Thus, for each data set, three SIBTEST statistics were calculated including
The average parameter estimates from the MMMT-2 model’s Item 1 DIF, Item 2 DIF, and Testlet 1 DTLF (i.e., the mean of the
Data were generated using SAS (SAS Institute, 2006). One hundred data sets were generated for each combination of conditions. SAS PROC GLIMMIX was used to estimate the MMMT-2 using residual pseudo-likelihood method (RSPL) estimation. SIBTEST (Shealy & Stout, 1993) was used to obtain estimates of βDIF and of βDBF for each replication’s data set. SAS was also used for summarizing results.
Results
Average DIF, DTLF, and DBF Estimates
Table 2 contains the average amount and direction of DIF in Items 1 and 2 estimated using the MMMT-2 model’s coefficients (
A Comparison of True Values With Mean Estimates of DIF and DTLF Using the MMMT-2 and SIBTEST by Condition
Note. DIF = differential item functioning; DTLF = differential testlet functioning. SIBTEST’s
MMMT-2 model estimates
When there was no DIF in any of the items (see Conditions 1 through 4), the average
The same results were found for the MMMT-2 estimates of DTLF. In conditions when the true γ
T
1 was zero, the average
SIBTEST estimates
In the four conditions of the scenario where there was no DIF or DTLF, SIBTEST’s average estimate of DIF and DBF was essentially zero (see Table 2). In Conditions 5 through 12, where true DIF and no DTLF was generated, SIBTEST’s average estimate of DBF,
When DTLF was introduced and paired with one item’s DIF (Conditions 13 through 16) and two items’ DIF (Conditions 17 through 20), the amount of DBF identified by SIBTEST increased accordingly (average
For the item-based cancellation DIF scenario (Conditions 21 through 24), the positive and negative DIFs generated for Items 1 and 2 cancelled each other at the bundle (testlet) level. Specifically, the average DBF (or total differential functioning across items in the testlet) was zero (combining the 0.4 with the negative 0.4 DIF effect for the two items). The same result was identified for the testlet-based cancellation DIF scenario (Conditions 25 through 28). In this scenario, positive DIF and negative DTLF (of the same magnitudes) were generated. This led to cancellation of the DIF (with an average SIBTEST
For both SIBTEST and the MMMT-2 coefficient estimates, the effect of a fixed testlet effect, γ T 10, did not seem to have a substantial effect on their estimation of DIF and DTLF. The value of the studied testlet effects’ variance had a very slight effect. When there was no variability in testlet abilities across examinees (i.e., for conditions when τ T 1 = 0), the magnitude of the DIF and DBF was slightly larger than when there was variability in testlet abilities (when τ T 1 = 1).
Relative Standard Error Bias
MMMT-2 model estimates
Using Hoogland and Boomsma’s (1998) cutoff of 10% as reflecting substantial relative standard error (SE) bias, a large proportion of the MMMT-2 DIF (
Relative Standard Error Bias of DIF and of DTLF Estimates Using the MMMT-2 and SIBTEST by Condition
Note. DIF = differential item functioning; DTLF = differential testlet functioning. Relative standard error bias of magnitude ≥10% was interpreted as substantial bias (Hoogland & Boomsma, 1998) and is highlighted in bold and italics.
SIBTEST estimates
SIBTEST’s estimates of DIF were associated with substantial SE bias in far fewer conditions than for the MMMT-2 model’s SE estimates (see Table 3). As was found for the MMMT-2 estimates, SE estimation was worse for Item 1 than for Item 2. Substantial SE bias was found in SIBTEST’s estimates of DTLF in only 4 of the 28 conditions examined in this study. However, there was no discernible pattern in the DTLF bias for both the MMMT-2 model and SIBTEST estimates.
Discussion
The primary purpose of this study was to introduce the MMMT-2 TRT model and to demonstrate is utility for assessing both DIF and DTLF. In particular, this model permits decomposition of sources of differential functioning into the component specific to an item and the part that might be in common with other items in a testlet. This distinguishes the MMMT-2’s formulation of DTLF from that used in more conventional measures, including, for example, SIBTEST’s DBF.
Another purpose of this study was that a clear distinction can and should be made between DTLF and DBF. The conceptualization of DTLF introduced here refers to the differential functioning of a testlet’s set of items that favors one group over another (after controlling for ability). This differential functioning, however, is the part that is common across items and does not include item-specific DIF. Thus, unlike DBF, the MMMT-2 parameterization of DTLF is not (and was found here not to be) affected by differential functioning cancellation nor amplification.
However, under the conceptualization of DBF used in SIBTEST, DBF includes the differential functioning that is in common across items within the bundle (here, a testlet) as well as the part that is unique to each item. Thus, DBF measures provide a total amount of differential functioning unique to and common across items in a testlet. Thus, although it is well known that DBF is commonly used because of its potential to amplify DIF, the effect of differential functioning cancellation on DBF is less commonly recognized. As the results of this study demonstrated, if the direction of an item’s DIF is opposite to that of another item’s DIF (or of the testlet’s DTLF), then the magnitude of the total differential functioning of the testlet’s items, as measured by DBF, will be less than if the DIF (and DTLF, where relevant) were in the same direction. A case can be made for the resulting test score not functioning differentially across examinees in the presence of differential functioning cancellation. However, test developers might be interested in recognizing these opposing patterns of DIF (and DTLF) and using that information with further test development and revision at the item and testlet level.
It is important to remember that detection of DIF, DTLF, or DBF does not mean that an item or testlet is biased. It is up to the researcher or practitioner to interpret substantively what might lie at the root of the differential functioning. One source of DIF and/or DTLF might be interpreted as unacceptable (adverse) bias and another source might be considered acceptable (benign). Use of the MMMT-2 permits partitioning of differential functioning into the item-specific and (common across items within a) testlet-specific components, which should help inform decisions concerning whether there is item and/or testlet bias and whether the differential functioning cancellation can be ignored.
Future Directions
As with any study, there are some limitations that emphasize the need for future research. First, only a small-scale simulation study was conducted to assess parameter and SE estimation of the MMMT-2 models’ DIF and DTLF coefficients and to demonstrate what SIBTEST’s DIF and DBF estimates measure. In particular, only one test length (50 items) with one testlet length (five items per testlet) was investigated here. Future research on the MMMT-2 could assess its performance with different test and testlets’ lengths as well as with tests that do not include only testlet-based items. Moreover, only a total sample size of 2,000 was investigated with balanced reference and focal group sample sizes. Both the test length and sample size were used to assess MMMT-2 model parameter estimation under good conditions. Future research could test the limits of parameter recovery for the MMMT-2 model using unbalanced group sample sizes and shorter test lengths. In addition, only 100 replications were generated per condition. Future research could assess the stability of the estimates’ and the pattern of results noted (especially for the SE estimates) for the MMMT-2 model parameters when more replications are used. Last, only uniform DIF and DTLF was modeled and assumed (with both the MMMT-2 and SIBTEST) when testing for differential functioning in the current study. The simplest Rasch-based MMM was assumed here including the assumption of unique item discriminations. Future research could extend the MMMT-2 model to permit item-specific discriminations as has been done with other MMMs (see, e.g., Rijmen, Tuerlinckx, De Boeck, & Kuppens, 2003). In addition, the MMMT-2 could be easily extended (see, e.g., Williams & Beretvas, 2006) to permit modeling of polytomous item responses’ functioning although this extension was not described nor evaluated here.
Second, estimation of the MMMT-2 model was performed only using SAS’s RSPL estimation procedure. Clearly, additional estimation procedures could be investigated including Laplace’s method and Markov chain Monte Carlo estimation. Given the underestimation of the DIF and DTLF noted for the MMMT-2 model’s estimates, it would be useful to see whether a different estimation procedure performs better than RSPL did here. Moreover, although the positive SE bias that was noted for some of the MMMT-2 model’s DIF and DTLF estimates would lead to the lesser evil of conservative (rather than liberal) statistical tests of DIF and DTLF, alternative estimation procedures might perform better in terms of SE estimation. It should also be noted that with most DIF and DTLF assessments, the practical rather than the statistical significance of the DIF (or DTLF) is emphasized when inferences are made about whether there is differential functioning. One benefit of the MMMT-2 is that the coefficients used to assess DIF and DTLF directly correspond to the item (and testlet) difficulty scale and thus interpretation of the magnitude of the DIF and DTLF is much easier than interpretation of, for example, statistics like SIBTEST’s βs.
Another important benefit of the MMMT-2 for DIF and DTLF assessment is that the source of the potential DIF and DTLF does not have to be known. The current study only investigated assessment of DIF and DTLF under the MMMT-2 in scenarios in which an observed factor (here interpreted as gender, for example) provided the source of the differential functioning. However, a version of the MMMT-2 model that includes random effects for each testlet (see Equation 7) and/or for each item’s difficulty can instead be estimated. If the variance components associated with each testlet’s (and/or item’s) random effects seems excessively large, then it should be inferred that there might be some latent factor causing items or testlets to function differentially that needs further exploration. This finding should then lead test developers to look in more depth at the relevant testlet (or item) to assess whether there is some observed variable that might explain this variability (leading to additional analyses that include an observed source of DIF or DTLF in the MMMT-2 model). Note also that, as with other MMMs, the source of DIF and DTLF does not have to entail a categorical grouping variable (or variables). Interval-scaled measures (e.g., examinees’ reading comprehension scores) can be included as sources of potential DIF and DTLF (see, e.g., Beretvas & Williams, 2004).
In addition, as noted previously, besides the utility of the MMMT-2 model’s parameterization, the model also offers a flexible framework common across the family of MMMs. An important extension to the MMMT-2 is the potential to include additional levels to model higher (or lower) levels and classifications of clusters. For example, a third level representing the clustering of students within schools could be added to the model to model this added source of dependence. Besides more appropriate variance decomposition, this permits addition of, for example, school descriptors to the model to explore their contribution to the explanation of impact, DIF, and DTLF. Alternatively, another level could be added at the lowest level if a researcher is interested in growth in item scores over time (see, e.g., Pastor & Beretvas, 2006).
As a final note, it should be emphasized that SIBTEST performed well in terms of its recovery of what its DIF and DBF statistics are intended to capture. Thus, as noted above, SIBTEST’s DBF statistic is designed to capture the total amount of differential functioning of items in a bundle (here, testlet). This sums together item-specific differential functioning and the differential functioning component that might be common across items (DTLF). This does mean that the resulting DBF estimate does not correspond with the current study’s conceptualization of DTLF (as differential functioning that is common in direction and magnitude across items in a testlet). Similarly, SIBTEST’s DIF statistic captures the total amount of differential functioning exhibited for an individual item. This again includes both item-specific differential functioning and any component common across all items in the testlet. Thus, again, the resulting SIBTEST estimates performed exactly as they are designed to do. The point of the comparison of the SIBTEST DIF and DBF estimates with those of the MMMT-2 model was to emphasize the distinctions between SIBTEST and the MMMT-2 model’s conceptualizations. Researchers are encouraged to recognize these differences and to use the conceptualization that best matches their research question’s purpose.
Footnotes
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
