Abstract
Cognitive diagnostic models (CDMs) are of growing interest in educational research because of the models’ ability to provide diagnostic information regarding examinees’ strengths and weaknesses suited to a variety of content areas. An important step to ensure appropriate uses and interpretations from CDMs is to understand the impact of differential item functioning (DIF). While methods of detecting DIF in CDMs have been identified, there is a limited understanding of the extent to which DIF affects classification accuracy. This simulation study provides a reference to practitioners to understand how different magnitudes and types of DIF interact with CDM item types and group distributions and sample sizes to influence attribute- and profile-level classification accuracy. The results suggest that attribute-level classification accuracy is robust to DIF of large magnitudes in most conditions, while profile-level classification accuracy is negatively influenced by the inclusion of DIF. Conditions of unequal group distributions and DIF located on simple structure items had the greatest effect in decreasing classification accuracy. The article closes by considering implications of the results and future directions.
Introduction
Cognitive diagnostic models (CDMs) have been developed with the purpose of providing examinees with finer grained diagnostic information on multiple dimensions (DiBello & Stout, 2007). CDMs assume that items on a measurement instrument can tap into multiple latent variables, interchangeably referred to as attributes, skills, or traits. A common example comes from fraction subtraction where solving an individual item requires multiple processes such as converting a whole number to a fraction, separating the whole number from a fraction, simplifying, and subtracting numerators (Tatsuoka, 1990). Each item in the fraction subtraction test is associated with one or more of the attributes assessed by the instrument. The CDM takes the item by attribute associations, formalized in the Q-matrix, and estimates the probability that a student has mastered the given attributes. Thus, a student would receive a score report classifying her performance on the test in terms of mastery or non-mastery on each of the attributes. 1 The statistical classifications from the CDM for each attribute enable students and teachers to plan how to improve learning outcomes. In this example, teachers can know which students struggle to convert a whole number to a fraction and develop instructional strategies for guiding those students. CDMs’ classification accuracy connects to subsequent decision making in learning.
The potential of cognitive diagnosis for learning and teaching is great, but the development of CDMs is more advanced than the CDM development processes necessary for establishing quality (Hou, de la Torre, & Nandakumar, 2014). In this study, we are particularly concerned about the examination of differential item functioning (DIF) in CDMs. DIF is an important indicator of test fairness. DIF measures potential test bias by examining the degree to which individuals from different groups with the same ability level have the same probabilities of responding correctly. In the context of CDMs, DIF analyses indicate the invariance of the prescribed attribute–item relationships across different groups (Hou et al., 2014). For example, a Q-matrix may unintentionally more closely match processes used by students in a particular curriculum as opposed to another. Another possible example is test accommodations resulting in different response processes among groups which would impact the parameter estimates across groups (Svetina, Dai, & Wang, 2017). In low-stakes assessments, differently motivated groups could have different item parameters.
While methods have been developed to estimate DIF in CDMs, other studies have not documented the degree to which DIF items impact classification accuracy. The purpose is to provide practitioners using CDMs an understanding of how different DIF-related scenarios affect classification accuracy. First, we provide the background for the study by discussing CDMs in general and the specific CDM used in the study. In addition, we situate this study in the existing CDM-DIF literature. The following section describes the design of the simulation study, data generation, and analysis plan. Results are then summarized, highlighting the main findings. Finally, we discuss the implications of this study for future research.
Background
CDMs
CDMs combine multidimensional modeling with criterion referenced latent attribute classifications (Rupp, Templin, & Henson, 2010). Similar to item response theory (IRT) models, CDMs model the probability of a successful response based on examinee and item characteristics (Almond, DiBello, Molder, & Zapata-Rivera, 2007). 2 In the place of a continuous latent variable associated with a student’s standing on a general attribute (e.g., general mathematical ability), CDMs estimate a profile of categorical, fine-grained latent attributes like those in the fraction subtraction example. This results in each individual being classified as a master/non-master of the attributes, allowing for personalized learning interventions.
CDMs require a confirmatory loading structure—the Q-matrix—that maps the multidimensional skills required by the individual items using a complex or simple structure (Rupp et al., 2010). The Q-matrix specifies which attributes are necessary to answer an item correctly. A typical approach to creating the Q-matrix is to consult subject matter experts as to which skills are necessary to respond to a particular item. A sample Q-matrix for four items and five attributes is presented in Table 1. In the Q-matrix, an entry of 1 indicates the attribute is required to correctly answer the item, while 0 indicates the attribute is not required. For example, Item 1 requires only Attribute 4, while Item 3 requires Attributes 2, 3, and 4. The probabilistic portion of the model then accounts for deviations from expected results (e.g., guessing correctly or incorrectly) by predicting the probability of attribute mastery.
Sample Q-Matrix for Five Attributes Across Four Items.
Deterministic Inputs, Noisy, “and” Gate (DINA) Model
We employ one of the most popular and parsimonious CDMs, the DINA model (Junker & Sijtsma, 2001). The DINA model is the most commonly used in empirical applications (Sessoms & Henson, 2018). The DINA model has also been shown to be flexible enough to fit in small sample contexts like classroom assessment (Sun & Suzuki, 2013) or large-scale assessments (H. Li, Hunter, & Lei, 2016). In addition, the ease of interpreting the model outcomes lends itself to broad implementation. Because of our goal to aid practitioners to assess DIF in practical uses of CDMs, we chose DINA as it is most likely to be employed in empirical applications.
Let
Parameter
where
DIF in CDMs
Traditionally, DIF is defined as occurring when the probability of correctly responding to an item differs by a grouping variable unrelated to the construct of interest, after conditioning on the latent variable or total score (Clauser & Mazor, 1998). As such, DIF items are a source of bias in assessment and represent a potential threat to construct validity. In CDMs, the definition of DIF is adapted to be “an effect where the probabilities of correctly answering an item are different for examinees with the same attribute mastery profile but who are from different observed groups” (Hou et al., 2014, p. 99). Specifically in the DINA model, DIF exists when the estimated item parameters differ for the focal and the reference groups after matching.
To this point, few studies have considered the impact of DIF in CDMs. The few studies that have investigated DIF in the context of CDMs (Hou et al., 2014; F. Li, 2008; X. Li & Wang, 2015; Milewski & Baron, 2002; Svetina et al., 2018; Zhang, 2006) have primarily focused their attention on comparing various methods to identifying DIF or DAF (differential attribute functioning). Milewski and Baron considered the possibility of differential attribute profiles between groups after controlling for overall score in comparing a variety of traditional DIF identification methods (Mantel–Haenszel (MH)), binary standardization, polytomous standardization, and analysis of covariance. Zhang (2006) also compared two traditional DIF methods (MH and Simultaneous Item Bias Test (SIBTEST)), at the item level as opposed to the attribute level, while also considering different conditioning variables (overall score and attribute profiles). Zhang found that matching on attribute profiles resulted in better Type I error and power rates than overall score. F. Li (2008) adapted a higher order (HO) DINA model to simultaneously account for DIF and DAF through the lower and higher level models, respectively. She found that the HO DINA approach produced better Type I error and power rates than the MH method. Hou et al. (2014) adjusted the Wald test for use in the CDM sphere, and found that it outperformed MH and SIBTEST in most contexts while producing inflated Type I error rates when items poorly discriminated. X. Li and Wang (2015) introduce an additional approach to CDM-DIF detection to effectively address multi-group variables (e.g., DIF across multiple races). Their log-linear cognitive diagnosis model–DIF method regresses item parameters on grouping variables and performed similarly to the Wald method in most conditions but outperformed it when slipping and guessing parameters were high. Svetina et al. (2018) compared logistic regression to MH and the Wald method under Q-matrix misspecification which could lead to structural DIF (e.g., reference and focal groups using different attributes to solve items). Results suggested that while MH and logistic regression generally yielded reasonable Type I error rates, the Wald method yielded higher Type I error rates in most conditions, and all methods performed produced higher Type I error rates when Q-matrix misspecification occurred in one rather than in both groups.
While various methods have been found to effectively detect DIF and its influence on item parameters, it is unclear the extent to which DIF affects CDMs’ primary purpose of accurately classifying respondents’ attribute mastery. Assessments with traditional measurement models assess the threat of DIF to score interpretation by comparing the estimated magnitude of DIF to established thresholds to determine whether the DIF is negligible or introducing significant bias (Zwick & Ercikan, 1989). Depending on the method used for estimating DIF, there are threshold rules established in traditional testing situations, for example, if using MH, large DIF is defined as items that have a delta value greater than the absolute value of 1.5 (Zwick & Ercikan, 1989). Applying traditional thresholds of what qualifies as consequential DIF from traditional testing procedures to CDMs is inappropriate given the differences in the measurement models being used. In addition, DIF thresholds in CDMs should be based on how DIF impacts classification accuracy as the classification results drive subsequent inferences and decisions. The findings will help practitioners understand how different conditions associated with the presence of DIF impact DINA classification accuracy.
Method
Design
We conducted a Monte Carlo simulation study based on a hypothetical cognitive diagnostic assessment. We assumed that DIF items are present in the assessment and calibrated with the non-DIF items, which is not an uncommon approach to DIF (Cho, Suh, & Lee, 2016). This approach illustrates the impact on classification when no action is taken to address DIF items. To examine this, we manipulated (a) DIF magnitude, (b) the percentage of items exhibiting DIF, (c) the type of DIF, (d) type of item exhibiting DIF, (e) the nature of the group distribution, and (f) the sample size for each group. The fully crossed study resulted in 196 separate conditions.
DIF magnitude
We examined a number of magnitudes of DIF, starting with those magnitudes (0.05, 0.10) that previous literature identified as small and large, in comparison with a baseline of no DIF (Hou et al., 2014; F. Li, 2008). We also examined additional levels (0.15, 0.20) to consider more extreme instances of DIF.
Type of DIF
We examined the relative impact of uniform and non-uniform DIF. Typically, uniform DIF occurs where items favor the reference group relative to the focal group. Table 2 describes how type and magnitude of DIF conditions are combined. We simulated uniform DIF more favorable to the reference group by increasing the slipping parameter for the focal group and increasing the guessing parameter for the reference group. This approach resulted in lower probabilities of correctly answering for focal group examinees who mastered all required attributes and increased probability of correctly responding for non-masters of required attributes among the reference group. Non-uniform DIF affects groups differentially, favoring one group’s probability of answering correctly in certain attribute profiles but impairing the same group’s probability with other attribute profiles. This was accomplished by increasing the slipping and guessing in the focal group, which increased the probability of a correct response for non-masters and decreased the probability for masters in the focal group (Hou et al., 2014).
Combination of Type and Magnitude of DIF.
Note.
Percentage of DIF items
Previous studies have suggested that actual tests can have a wide range of possible proportions of DIF items (F. Li, 2008). We used the levels considered in previous studies, 10% (three items) and 30% (nine items) of items (F. Li, 2008; X. Li & Wang, 2015; Zhang, 2006).
Items with DIF
We also manipulated the type of DIF items according to their Q-matrix item structure. The Q-matrix used in this study includes both simple and complex items. We considered the following possibilities for types of DIF items: (a) only simple items, (b) only complex items, and (c) a mix of simple and complex items. For the mixed item condition, when DIF was 10% of items, one item was simple and two were complex, and when DIF was on 30% of the items, four of the items were simple and five were complex.
Group distributions
A complicating factor in many empirical analyses of DIF is unequal reference and focal group distributions (F. Li, 2008; Mazor, Clauser, & Hambleton, 1992). Thus, we followed previous literature by considering the impact of equal and non-equal examinee distributions, where the non-equal distribution was simulated by setting the multivariate normal focal group distribution one standard deviation lower than the reference group.
Sample size per group
Another feature common to empirical DIF analyses is the unequal sample size of the reference and focal groups (Quesen & Lane, 2018). To improve the understanding of the impact of this, we considered both an equal sample size between groups (n = 1,000) and an unequal sample size (n = 1,300 for reference group, n = 700 for focal group).
Data Generation
We simulated the data using the DINA model (see Equation 2) based on a 30-item, five attribute Q-matrix. As noted above, the DINA model is the most commonly used CDM. Similarly, five attributes is the most common number of attributes in an assessment (Sessoms & Henson, 2018). The Q-matrix was taken from Hou et al. (2014). We randomly generated item discrimination parameters using a uniform distribution (.1, .3) to approximate a realistic, well-fitting model (de la Torre & Douglas, 2004; Hou et al., 2014). Previous literature indicated that correlation levels had relatively minor impact on DIF item parameter recovery (Svetina et al., 2018), so we set attribute correlation in the generating variance–covariance matrix to be equal to .7. This level of correlation is consistent with the correlations between test subscores in empirical tests (Sinharay, 2010). For each of the 192 conditions, 100 replications were generated, and results were analyzed across replications within each condition.
Analysis
We evaluated the results using two outcome variables: attribute-level classification accuracy (ACA) and profile classification accuracy (PCA). Classification accuracy evaluates to what degree the estimated mastery of attributes agrees with the true profile, both at the individual attribute and attribute profile level, defined as
and
where
Results
The classification accuracy results at the attribute level are reviewed first followed by the profile level. The results for the reference and focal groups are presented separately in figures showing the classification accuracy rate across conditions. Figures 1 and 2 show the ACA rates for the reference and focal groups, respectively, while Figures 3 and 4 do the same for the PCA rates. Each figure includes horizontal panels for each DIF magnitude level, which are split to compare the equal and unequal group distribution conditions. The vertical panels represent the item types (simple, complex, and mix), and the shape and color of the points represent the DIF type and percentage of DIF items, respectively. In addition, dashed and dotted lines are included in the graphs to represent baseline conditions with no DIF present and equal or unequal group distributions, respectively.

ACA rates across conditions: Reference group.

ACA rates across conditions: Focal group.

PCA rates across conditions: Reference group.

PCA rates across conditions: Focal group.
ACA Results
Figures 1 and 2 demonstrate the relative difference in impact between the reference and focal groups. As expected given our design, the simulation results favored the reference group. The results indicated that the lowest ACA for the reference group was around 0.85 (Figure 1), and the lowest ACA for the focal group was around 0.72. ACA rates for the focal group were less than 0.85 in more than a third of the 192 conditions (see Figure 2).
DIF magnitude averaged across other conditions had a relatively limited impact on ACA rates. The average ACA when DIF magnitude was at its highest fell only 0.02 or 0.03 (equal and unequal distributions, respectively) from the ACA rate when there was no DIF. DIF magnitude was more influential when 30% of the items exhibit DIF, the DIF items were simple, and the DIF was non-uniform. Holding these factors at the values specified, increasing DIF from 0 to 0.20 is associated with 0.10 drop (to 0.81) and 0.14 drop (to 0.72) for equal and unequal group distributions, respectively.
Unequal group ability distributions had the largest impact on ACA in conditions with and without DIF. For example, in no DIF conditions, ACA rates for the focal group were 0.06 lower than the reference group when group distributions were unequal. This difference remained fairly constant, but it increased (a) when DIF magnitude increased, (b) when 30% of the items were modeled with DIF, and (c) when DIF was present in simple items. The gap between the equal and unequal group distributions grew from 0.07 when DIF was small (0.05) to a gap of 0.09 when DIF magnitude was at its largest (0.20). This means that when DIF was as large as 0.10, non-uniform DIF on 30% of items, and the items were simple, the equal group condition achieved an ACA of 0.86 while the unequal group condition ACA was 0.79. The influence of unequal group distribution did not affect the ACA for the reference group in a meaningful way because the unequal group condition systematically lowered the mean for the focal group only.
The type of item where the DIF occurs had a meaningful influence on ACA for both the reference and focal groups. In most conditions, when DIF only occurred in simple items, ACA was always at least 0.01 lower than when DIF was present in complex items or in mixed items. This difference was exacerbated by DIF magnitude and percentage of items with DIF. As DIF magnitude and percentage of items with DIF increased, the reduction in ACA for simple items grew larger than it did when items were complex or mixed. For example, the ACA for the focal group when DIF was 0.20 and non-uniform on 30% of the items was 0.81 when on simple items, 0.88 when on mixed items, and 0.89 when on complex items. Uniform DIF ameliorated this impact somewhat, as the same conditions described in the previous sentence would result in ACA of 0.86 for simple items, and 0.90 for mixed and complex items. The results also showed that mixed items always had an equal or lower ACA than complex items.
The type of DIF had differential influences on the groups. As expected, uniform DIF had a more meaningful impact on the reference group ACAs while non-uniform DIF impacted the focal group more. For example, when 30% of items have 0.20 magnitude DIF, the average ACA for the focal group when the DIF was uniform was 0.89 and when DIF was non-uniform was 0.86. The rates for the reference group at those same condition values were 0.89 and 0.91, respectively. The influence of the DIF type was moderated by the type of DIF item and the DIF magnitude. Thus, when DIF magnitude was small, differences between non-uniform and uniform DIF in terms of ACA were minimal to none. When DIF magnitude was 0.20, the difference between uniform and non-uniform conditions increased to 0.05 and 0.10 in the focal group (equal and unequal distributions, respectively) and to 0.02 in the reference group.
The remaining conditions provided some general patterns. As expected, when the proportion of items was 30% as opposed to 10%, ACA rates were lower, especially as DIF magnitude increased. Also, ACA remained essentially the same when the sample size was equal or unequal. The figures’ slight overlaps of shapes represent the differences that came from equal and unequal sample size.
For individual attributes, ACA generally fell the lowest when DIF occurred in an unequal group distribution. When DIF occurred on complex structured or a mix of simple and complex structured items, none of the confluence of conditions resulted in ACA falling below 0.80. However, if DIF was located only on simple structured items, DIF magnitude as large as 0.10 resulted in ACA rates falling below 0.80 (assuming unequal distributions and 30% of items showed uniform DIF). DIF magnitude had a significant impact on the focal group ACA, but primarily on the reference group’s ACA when the DIF was uniform. Uniform and non-uniform types of DIF also influenced the ACA to some extent.
Profile-Level Classification Accuracy Results
Figures 3 and 4 show the profile-level results for the reference and focal groups, respectively. Similar to the ACA analysis, the simulation was designed to primarily affect the focal group. As the graphs indicated, and consistent with the CDM literature (Liu, Huggins-Manley, & Bradshaw, 2017; Templin & Bradshaw, 2014), the highest PCA (0.69) was much lower than the highest ACA (0.91), and PCA rates had a much wider range of values (approximately 0.69-0.15 compared with 0.91-0.67). The data’s random noise was compounded when classifying a vector of attributes. PCA rates were superior to the probability of randomly selecting the appropriate attribute profile (i.e., 1/32 or ~3%) but are too low to rely upon.
The factors that influenced ACA rates exerted much of the same impact on PCA rates. The magnitude of DIF had limited influence in some conditions but significant influence in other conditions. Thus, when 10% of items exhibited DIF, and items were complex, increasing the DIF magnitude by 0.30 resulted in only a 0.01 decrease in PCA in the focal group. In most other conditions, changing the DIF magnitude resulted in decreasing PCA rates. The decrease in PCA rates was particularly dramatic when DIF was located on simple and mixed items as seen in Figure 4, where gaps grew larger going from left to right. For example, when non-uniform DIF occurred on 30% of mixed items in an equal group distribution setting, focal group PCA steadily fell from 0.69 in no DIF conditions to around 0.60 when DIF magnitude was 0.20.
Unequal group distributions resulted in an even greater difference in the focal group’s PCA rates than in comparable ACA rates. Whereas ACA rates differed by 0.06 between the group distribution conditions with no DIF, PCA rates differed by 0.20. Similar to ACA rates, this gap grew larger as DIF magnitude increased, particularly when 30% of items were DIF and DIF was located in simple structure items. Thus, when DIF magnitude was 0.20, PCA in the equal group condition was around 0.50 while PCA in the unequal group condition was 0.23.
The type of item structure had a meaningful influence on PCA rates. When DIF was located in simple items, the PCA rate was always lower than when it was located on complex or a mix of items. The difference in PCA rates between the reference group and focal group increased from complex to mix to simple structure items. Non-uniform DIF on simple structure items resulted in even greater decreases in PCA for the focal group, particularly when group distributions were equal. When DIF was located on complex items, the impact of the percentage of items with DIF was much larger than when DIF was located on a mix of items.
Other general patterns were also observed. First, PCA rates were always lower in uniform than non-uniform conditions. Outside of mix structured DIF items, PCA was always lower when 30% of the items showed DIF compared with 10% of DIF items. Similar to ACA rates, unequal sample sizes had a minimal influence on the PCA results.
Discussion
Research in CDMs has increased over the past several years as demand for formative fine-grained assessment has grown. Theoretically, CDMs can provide classifications that lead to this kind of feedback in a way that other models cannot. However, there has been limited research examining how known challenges to measurement impact CDM outcomes. Previous examinations of DIF in CDMs have assessed impacts on parameter precision and have not reported on the impact of DIF on CDM classification as they found limited variation (L. Hou and X. Li, personal communication, February 17 and 19, 2019). The findings are consistent with these studies excepting the meaningful impact we observed with the novel consideration of unequal distributions and the DIF items’ structure.
At the individual attribute level, diagnostic classifications were mostly robust to DIF conditions. Assume an 80% classification accuracy rate was acceptable in a low-stakes assessment context. Based on this hypothetical criterion, then only six out of 192 conditions would result in an unacceptably low ACA. ACA fell below this criterion primarily when DIF occurred exclusively on simple structured items and when group distributions were unequal. In such situations, DIF of 0.10 could result in inadequate attribute classification, also assuming 30% of items showed uniform DIF. However, if a group’s ability distributions were equal, then DIF magnitude would need to be as great as 0.30 to result in inadequate ACA.
The influential impact of the item structure and group’s ability distribution on classification accuracy was interesting and prompted reflection. As the literature on the identifiability of the DINA model has demonstrated, without simple structure items for each attribute, the model cannot be identified, resulting in inaccurate and inconsistent estimation of parameters and classifications (Gu & Xu, 2018; Xu & Zhang, 2016). Given the relative importance of such items for model identifiability, it is reasonable that increasing the random noise on simple structure items would result in less effective measurement of the given attributes and poor classification accuracy. DIF on complex items was less influential because identifiability only requires that each attribute is measured by three items, and in this case, the given Q-matrix well exceeded that number for each attribute.
The unequal group distribution condition estimated the DINA model with the combination of two different multivariate normal distributions. This created a bimodal distribution as opposed to a customary uniform or multivariate normal distribution of attribute profiles. Previous research has indicated that under non-normal distributional conditions, invariance of DINA model parameters does not hold (de la Torre & Lee, 2010). In de la Torre and Lee’s (2010) empirical example, estimated guessing and slipping parameters markedly varied across different clusters within the dataset. In this study, the unequal distribution conditions set the focal group one standard deviation below the reference group, creating what de la Torre and Lee described as “extreme underlying distributions” (p. 125). Because of the sample’s nature, the item parameters estimated for the entire sample meaningfully differed from the guessing and slipping behaviors that occurred in the portions of the population that deviated the most from the bulk of the sample, primarily those in the lower half of the focal group. The examinees in the focal group were farther from the norm of the distribution and were subject to varying item parameter estimates, resulting in lower focal group classification accuracy. The propensity for differing distributions in populations makes this a relevant concern (Penfield & Lam, 2000).
The classification of entire attribute profiles was not robust. Without introducing any DIF, this study randomly generated item parameters between 0.10 and 0.30, and found PCA under this model was at best 0.69. If the model was estimated with data from unequal group distributions, the focal group PCA was much lower. Empirical applications of the DINA model that make inferences based on attribute profiles should ensure that items are highly discriminating, well fitting to the Q-matrix.
We recommend that analysts and scholars do not extrapolate the results of this simulation study beyond the confines of the conditions we employed. This initial study of DIF influence on diagnostic classification is useful for its focus on factors commonly investigated in CDM studies, that is, DINA model, five attributes, complex and simple item–attribute structures, moderately discriminating item parameters, typical proportions of DIF items, and a variety of DIF magnitude levels. Further research should explore the robustness of CDM classifications to DIF by considering different numbers of attributes and attribute relationships and different levels of item discrimination. While we anticipate that fewer attributes and higher item discrimination would result in more robust classification accuracy rates both at the attribute and profile level, how these factors relate to item type and group distributions would be of great interest. In addition, future research should consider how robust more flexible, general CDMs are to DIF. Finally, this study considers the approach of not modeling DIF separately, which has been shown to be a common strategy (Cho et al., 2016). Future research should consider approaches to dealing with DIF. Multiple-group and DIF modeling approaches show particular promise.
Overall, this study suggests a number of recommendations for analysts concerned about DIF occurring in the DINA model. First, the magnitude of DIF and number of DIF items must be large to meaningfully diminish classification accuracy at the individual attribute level. On the contrary, even moderate levels of DIF can notably reduce profile-level classification accuracy. Second, disparate ability distributions between groups will reduce classification accuracy, even in the absence of DIF. Analysts should assess group distributions to understand how this might impact their classifications. Third, analysts should be aware of the attribute structure of the DIF items. In this study, we found that simple structure items with DIF had more impact on classification accuracy than if DIF is associated with complex structure items.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
