Abstract
When diagnostic assessments are administered to examinees, the mastery status of each examinee on a set of specified cognitive skills or attributes can be directly evaluated using cognitive diagnosis models (CDMs). Under certain circumstances, allowing the examinees to have at least one opportunity to correctly answer the questions and assessments, with repeated attempts on the items, provides many potential benefits. A sequential process model can be extended to model repeated attempts in diagnostic assessments. Two formulations of the sequential generalized deterministic-input noisy-“and”-gate (G-DINA) model were developed in this study. The first extension uses the latent transition analysis (LTA) approach to model changes in the attributes over attempts, and the second extension constructs a higher order structure of latent continuous variables and latent attributes to account for the dependences of the attributes over attempts. Accurate model parameter estimation and correct classifications of attributes were observed in a series of simulations using Bayesian estimation. The effectiveness of the developed sequential G-DINA model was demonstrated by fitting real data from a longitudinal mathematical test to the developed model and the longitudinal G-DINA model using the LTA approach. Finally, this article closes by discussing several important issues associated with the developed models and providing suggestions for future directions.
Keywords
Over the past few decades, interest in the development of cognitive diagnostic assessments to provide diagnostic information about students’ strengths and weaknesses according to a set of cognitive skills or attributes has been increasing; this approach provides the advantages of facilitating learning and instruction in educational environments. In contrast to the conventional approaches to scaling and ordering students along a proficiency continuum (i.e., item response theory [IRT] models), cognitive diagnosis models (CDMs) hold great promise for extracting relevant diagnostic information, and they serve as an important means of constructing diagnostic assessments for identifying the mastery status of individuals on a set of multiple fine-grained skills (Leighton & Gierl, 2007; Rupp, Templin, & Henson, 2010). Diagnostic assessments based on certain standards can be used for different purposes. In state-mandated testing programs, for example, school performance can be evaluated using large-scale diagnostic assessments to satisfy state and federal accountability requirements; in practical instructional settings, diagnostic assessments can be used to aid instructors in understanding and addressing the specific needs of students and to guide teachers in designing effective instructional interventions for those in need (Huff & Goodman, 2007).
In response to the demand for more fine-grained information and personalized feedback on students’ strengths and weaknesses, tests that are cognitively diagnostic in nature can be administered to students multiple times in a sequential process by allowing examinees multiple attempts on test items until they submit a correct response to those items. Such an approach, that is, giving examinees repeated attempts at problem solving, was defined terminologically as answer-until-correct test administration (Pressey, 1926), and the literature has documented that the sequential assessment approach has a considerable number of potential benefits that are associated with student learning and development; these benefits include, for example, a reduction in test anxiety because of multiple attempts being allowed and the subsequent feedback when the examinees complete each attempt (Attali & Powers, 2010; DiBattista & Gosse, 2006; Epstein, Epstein, & Brosvic, 2001; Kluger & DeNisi, 1996).
The large body of research on the effectiveness of administering an answer-until-correct test has been supported by experimental evidence with the use of partial feedback and other materials to supplement students’ knowledge content levels (for an intensive discussion, see Ye, Fellouris, Culpepper, & Douglas, 2016). In an introductory psychology course, for example, Epstein et al. (2001) found that students who received immediate response feedback using an answer-until-correct testing procedure on multiple-choice examinations performed better than those using the traditional examination form. Another example was demonstrated with the use of constructed-response items on an answer-until-correct examination by Culpepper (2014). Examinees were allowed to provide a maximum number of five opportunities to submit a correct response to test items in a calculus-based probability theory course, and in this study, the sequential process of cognitive operations over repeated attempts was observed and the variation of content knowledge supplement over attempts was significantly large such that the benefits of retrieval-based activity were acknowledged.
At least two principal psychometric methods for analyzing sequential data collected from repeated trials deserve further discussion. The first method counts the number of attempts that an examinee requires to correctly answer an item, treats the required attempt number that an examinee has as a reciprocal function of an observed score, and then employs polytomous IRT models to analyze the observed scores (e.g., Attali, 2011; Attali & Powers, 2010; Muñiz & Menéndez, 2011). Under this scoring rubric framework, let T denote the maximum number of attempts allowed; when an examinee correctly responds to an item on attempt k (
When repeated attempts on test items are allowed, a succession of cognitive operations over repeated trials should be considered and can be decomposed into different sequences that are reflective of different success rates on each attempt. Sequential IRT models have been developed for analyzing data that are collected from multiple attempts on items in cognitive tests, survival trials in experiments, or practices in psychomotor research (Akkermans, 2000; Albert & Chib, 2001; Bechger & Akkermans, 2001; Spray, 1990; Tutz, 1990). These proposed sequential IRT models can account for the sequential process of cognitive or psychomotor operation when examinees are allowed to respond to specific tasks via several attempts; however, certain methodological limitations should be noted. For cognitive assessments, it is unlikely to assume that the success rates on each attempt are independent because examinees could receive feedback or instructional intervention subsequently to supplement their content knowledge; consequently, their ability levels are expected to change over attempts. Recently, Culpepper (2014) proposed a series of sequential IRT models for repeated attempts on test items in which the probability of success on an attempt was conditional on that of success on the previous attempt by adding attempt-specific threshold parameters or person-specific growth parameters to model the dependences of multiple attempts. A testing program that allows examinees opportunities to correctly answer questions provides diagnostic information; however, neither Culpepper’s models nor the previous sequential IRT models can provide sufficient diagnostic information, in contrast to CDMs, because the proficiency levels of examinees are scored on a proficiency continuum using conventional IRT models, which offer limited diagnostic information.
Allowing examinees multiple opportunities to answer test items correctly implies that the changes in latent traits or attributes would occur over attempts. In sequential IRT models for repeated attempts, Culpepper (2014) incorporated a linear growth model for assessing the changes in examinees’ latent continuous traits. Note that examinees’ latent traits are not necessarily assumed to increase over attempts in the latent growth model because the growth rate may be negative as they practice different trials. In the sequential CDMs, latent attributes that examinees possess on each attempt are expected to change across multiple measurements, as in sequential IRT models. For the assessment of attribute changes over time, latent transition analysis (LTA) for CDMs (Kaya & Leite, 2017; Li, Cohen, & Bottge, 2016) and multilevel higher order CDMs (Huang, 2017; Huang & Hung, 2014) have been developed for longitudinal surveys in the literature. However, the two approaches do not consider the sequential process of cognitive operations and do not serve as ideal models for analyzing answer-until-correct test data.
In this study, sequential CDMs were developed by the authors to account for the attribute mastery statuses of examinees on multiple successive diagnostic assessments. Variations between the different measurement occasions (i.e., multiple attempts) in the mastery statuses of the individuals with respect to a set of defined attributes can be modeled using the LTA approach or a higher order CDM approach (Huang, 2017; Huang & Hung, 2014; Kaya & Leite, 2017; Li et al., 2016). This article is organized as follows. The most general CDM of the generalized deterministic-input noisy-“and”-gate (G-DINA) model (de la Torre, 2011) is adopted and extended in this study, which is introduced in Online Appendix A in the online supplemental material. The next section describes the sequential CDM framework and discusses the development and extension of the model by originally combining the sequential process model and the G-DINA model, followed by a description of two approaches to modeling changes in the latent attributes. A series of simulations are conducted to assess the parameter recovery and classification accuracy of the examinees, using Bayesian estimation to calibrate the model parameters. An empirical analysis is presented to demonstrate the applications and implications of the proposed models. The final section draws conclusions about the new models and provides suggestions for future research.
Sequential CDM Framework
When allowing examinees multiple attempts on a diagnostic assessment, given a maximum number of attempts T, the number of attempts that examinee i requires to submit a correct response to item j can be recorded and it is denoted as
and the conditional probability of an incorrect response for attempt m is
where
Therefore, the probability of making
if
if
Because examinees are allowed multiple attempts on test items, the probability of a correct response would change as the number of attempts increases. In such cases, examinees can practice multiple times on test items and may have the opportunity to make use of additional materials through partial feedback in educational interventions to supplement their content knowledge or to refine the response strategies that are related to problem solving. Therefore, attribute mastery statuses of examinees are treated as dynamic over attempts, which is similar to the approach adopted in sequential IRT models for answer-until-correct tests (Culpepper, 2014). In addition to the changes in attribute mastery statuses, items may be easier in later attempts than in early attempts, and item parameter drift may occur. It is thus possible to add an additional item parameter to capture the variation of the item parameter over time for each attempt except for the first attempt. Item parameter invariance over time or attempts is a common assumption in longitudinal surveys (e.g., Huang, 2015, 2017) and item parameter drift may be caused by several factors (Park, Lee, & Xing, 2016). Following reasonable assumptions and to avoid reaching beyond the focus of this study, it can be assumed that the item parameters are invariant over time in this study.
Figure B1 of Online Appendix B illustrates an example in which the maximum number of attempts is three for examinee i in response to item j with a reduced attribute vector
LTA Approach
Stage sequential changes in multiple attributes can be captured by the transition probabilities between two adjacent measurement occasions through LTA (Kaya & Leite, 2017; Li et al., 2016). For notational convenience, the subscript i that indicates a specific person is omitted in the following formulation. Let the entry in
where
Higher Order Modeling Approach
Because the attributes are seldom independent and their correlations should be considered, a higher order latent variable can be assumed to govern the mastery status of each attribute, and higher order CDMs have been developed to describe the dependences of the multiple attributes (de la Torre & Douglas, 2004). Multilevel extension of the higher order CDMs provides an alternative to assessing the changes in the latent attributes and to simultaneously estimating the continuous latent trait, which can serve as an overall evaluation of the examinees’ performances (Huang, 2017; Huang & Hung, 2014). Let
where λ1k is the discrimination parameter and λ0k is the location parameter for attribute k. The function described by Equation 8 is equivalent to the response function of the two-parameter logistic model (2PLM), but the outcome variables in Equation 8 are latent rather than observed, as in the 2PLM. As with the G-DINA model parameters described above, the λ1k and λ0k parameters are assumed to be time-homogeneous and are estimated to be equal for each attempt. Combining Equations A1 and 8 leads to a higher order G-DINA model when translating the full attribute vector
In the higher order CDMs, such as the higher order G-DINA model, both
Therefore, the likelihood function of
where
To model the latent growth for the
where
Note that although the higher-order-structured parameters (i.e., the
The research hypotheses and expectations formulated in this study are as follows: (a) the model and person parameters can be recovered satisfactorily in the sequential G-DINA model using the LTA and higher order modeling approaches with the use of Bayesian estimation; (b) a large sample size can enable more precise estimations of the model parameters; (c) the LTA approach can provide higher accuracy classifications than the higher order modeling approach because the transition probabilities between two attempts are directly modeled rather than through a higher order structure; and (d) the use of the sequential G-DINA model to fit the empirical longitudinal data would be more efficient than the use of the non-sequential G-DINA model because fewer items are administered to examinees if they correctly answer the items on early attempts.
Method
Simulation Design
Two simulation studies were conducted to assess the quality of parameter estimation in the sequential G-DINA model with respect to two different approaches of modeling the changes in examinees’ mastery statuses over attempts. The first simulation used the LTA approach, and the simulation design was adapted from an empirical analysis performed by Li et al. (2016). Because the LTA approach was extended in the sequential G-DINA model, the simulation design and generated values were set to be similar to those reported in a previous study rather than simulating the conditions arbitrarily. The test length was fixed to 23 items that measured four attributes, and the construction of the real-data
Transition Probabilities for Attempt Changes.
Note. The superscript indicates four transition conditions, where the values of 1, 2, 3, and 4 denote the following conditions: remaining in nonmastery status, transitioning from nonmastery to mastery status, transitioning from mastery to nonmastery status, and remaining in mastery status, respectively.
It is likely that certain misconceptions arise when examinees receive more instructional materials on the following measurements (National Research Council, 1997). Therefore, the probability of transitioning from mastery to nonmastery status may be greater than that of remaining in mastery status. In addition, difficult cognitive attributes, for example, problem-solving skill, could result in a higher transition probability from mastery to nonmastery status (Li et al., 2016). The generated transition probability parameters (see Table 1) indicate that a few conditions involved a rather larger probability for the transition from mastery to nonmastery status. For example, the transition probability for the abovementioned situation was sufficiently large from Attempt 1 to 2 but decreased dramatically from Attempt 2 to 3. Acknowledging the possible and occasional backwardness in attitudes, although such a situation is rare in answer-until-correct test administrations, any increasing monotonicity for attribute mastery profile in this study is not assumed.
To generate the mastery proportions of the four attributes similarly to the study of Li et al. (2016), the tetrachoric correlations among the attributes on the first attempt were fixed to .5, and the mastery thresholds for each attribute were set to 0.08, –0.03, 1.28, and 0.58; consequently, the proportions of the examinees who mastered each of the four respective attributes on the first attempt were 0.47, 0.51, 0.10, and 0.28. After generating the attribute mastery status on the first attempt for each examinee, the attribute mastery status of the following attempts could be determined according to the transition probabilities listed in Table 1. Because the saturated G-DINA model was used to generate the item responses and the stabilities of the parameter estimates could be influenced by the number of observations (e.g., de la Torre, 2011), samples of 500 and 1,000 examinees were employed to assess the effects of the sample size on the parameter estimation in the developed sequential G-DINA model.
The parameters in the G-DINA model were generated as follows. When one attribute was measured (i.e., when
In the second simulation study, a higher order model was used to generate examinees’ mastery statuses of the four attributes; the other settings were set the same as in the first simulation study. The responses of a large sample of 3,000 examinees were generated using the LTA approach and were calibrated using the higher order modeling approach to obtain stable parameter estimates. When the G-DINA model parameter estimates were set to the generated values, the parameter estimates of the higher order model (i.e., the discrimination, location, mean, variance, and covariance parameters) were used as the generated values for the higher order sequential G-DINA model. Thus, the results of the second simulation study could be compared with those of the first simulation study.
Therefore, the generated values for the four respective attributes were set to 1, 0.741, 0.710, and 1.927 for the discrimination parameters and 0.182, 0.274, 3.709, and 1.602 for the location parameters. The regression residual was randomly generated from a normal distribution with a mean of 0 and variances of 0.507, 0.033, 0.183, and 0.036 for the four attempts, and the
For both simulation studies, each condition was replicated 30 times. The authors attempted to increase the number of replications to 100 for several conditions and found that the resulting estimates were very close to those obtained from 30 replications, which implied that the sampling variation was very small and the differences among them could be neglected. Therefore, 30 replications were deemed sufficient to provide reliable results, as in Bayesian CDMs involving fewer replications than those used in this study (e.g., de la Torre & Douglas, 2004, 2008; Huang & Wang, 2014).
Analysis
The computer program WinBUGS 1.4 (Spiegelhalter, Thomas, & Best, 2003), which implements Bayesian estimation with Markov Chain Monte Carlo (MCMC) methods, was used to calibrate the parameters in the developed sequential CDMs. The Metropolis-Within-Gibbs algorithm was used to construct a Markov chain to generate samples for parameter estimation in this study. For the two types of sequential G-DINA models, in addition to the G-DINA model parameters (i.e., the
Prior distributions should be specified for each model parameter before simulating the Markov chain to obtain the target density of the parameters. The settings of the prior distributions described below were consistent with or similar to those reported in previous studies involving Bayesian estimation in CDMs (e.g., Huang, 2017; Huang & Hung, 2014; Huang & Wang, 2014; Li et al., 2016). A beta distribution with both hyperparameters equal to 1 was used for the G-DINA model probability and transition probability parameters; a normal prior with a mean of 0 and a variance of 4 was used for the threshold, location, and random-slope mean parameters; a uniform distribution between −1 and 1 was used for the attribute correlations; a lognormal prior with a mean of 0 and a variance of 1 was used for the discrimination parameters; a gamma prior with both hyperparameters equal to 0.1 was used for the inverse residual variances; and finally, a Wishart distribution with a diagonal matrix whose nonzero entries were all equal to 0.1 and with two degrees of freedom was used as the prior for the inverse variance–covariance matrix.
The multivariate potential scale reduction factor (Brooks & Gelman, 1998) was used as the criterion for verifying the convergence of the parameter estimates using three parallel chains for five randomly selected simulated data points in each condition. Consequently, the authors performed 15,000 iterations to calculate the posterior mean for the parameter estimates over the iterations after the first 5,000 iterations were used as burn-in. Parameter estimates were then computed by averaging the samples from the remaining 10,000 iterations with a thinning interval of five. In addition, label switching is one major challenge associated with mixture models when Bayesian estimation is implemented for parameter calibration. In the simulation and empirical analyses, no multiple nodes in the marginal posterior distributions were observed for the specified parameters whether within a single MCMC chain or when using different initial values, implying that label switching did not occur. The quality of the model parameter estimation was evaluated by computing the bias and root mean square error (RMSE) for each estimator, and the recovery of the examinees’ attribute mastery statuses was assessed by computing the correct classification rate (CCR) for each individual attribute.
Results
The mean and standard deviation of the bias and RMSE are provided for different types of parameters instead of for individual parameters, owing to space constraints. Tables 2 and 3 summarize the results of the parameter recovery for the sequential G-DINA model using the LTA and higher order modeling approaches, respectively. The two approaches estimated a common set of G-DINA model parameters but used different schemes to account for the dependences of the attributes and to model the stage sequential change. The results show that the two approaches yielded satisfactory parameter recovery for the G-DINA model parameters, and the difference between the two approaches was very small. For the LTA approach, attribute correlations, mastery thresholds, and transition probabilities were estimated acceptably, and the large sample size provided more precise parameter estimation. For the higher order modeling approach, relatively larger bias and RMSE values were observed for the discrimination, location, residual variance, and covariance–variance parameters. This result was obtained because the quality of the latent continuous trait parameter estimation in the higher order CDMs depended strongly on the number of measured attributes (Hsu & Wang, 2015) and the latent growth trajectory would be influenced by the imprecise latent continuous trait estimation (Huang, 2017; Huang & Hung, 2014). Nonetheless, as shown in Table 2, from sample sizes of 500 to 1,000, satisfactory parameter recovery could be obtained when the large sample size was used.
Parameter Recovery for the LTA-Approach Sequential G-DINA Model.
Note. LTA = latent transition analysis; G-DINA = generalized deterministic-input noisy-“and”-gate; RMSE = root mean square error.
Parameter Recovery for the Higher Order Approach Sequential G-DINA Model.
Note. G-DINA = generalized deterministic-input noisy-“and”-gate; RMSE = root mean square error.
Figure 1 shows the mean CCR for the two approaches across replications for each individual attribute in the four attempts. Although the LTA approach had better attribute recoveries than the higher order modeling approach in most conditions, the mastery status of the attributes appeared to be classified accurately for both approaches because of the higher CCR values. Because Attributes 3 and 4 were measured by a moderate number of items, 8 and 9 items, respectively, in contrast to Attributes 1 and 2, which were measured by a large number of items, 19 and 15 items, respectively, it was not surprising that Attributes 3 and 4 had lower CCR values than those of Attributes 1 and 2. In summary, with a sufficiently large sample, the parameters in the sequential G-DINA model using the two different approaches to account for the dependences of the attributes over attempts were recovered accurately using Bayesian estimation.

CCR for the sequential G-DINA model for sample sizes of 500 (top panel) and 1,000 (bottom panel).
Empirical Demonstration
The data used for application demonstration were obtained from a longitudinal basic-ability-assessment survey in which a mathematical test was administered to 3,527 Taiwanese junior high school students at three points in time (Chang, 2007). Four attributes, namely, analysis, numbers, geometry, and algebra, were identified by content experts for a subset of the mathematical test; therefore, seven multiple-choice items were analyzed. The application of the sequential G-DINA model to the multiple-choice items was justified because the mathematical test was treated as a low-stakes examination and because multiple-choice tests have been widely used in the answer-until-correct scoring procedure (e.g., Ben-Simon, Budescu, & Nevo, 1997; Epstein et al., 2001).
The
Table 4 presents the frequency and proportion of the sequential response patterns that satisfy the assumption of the sequential G-DINA model for each item. For each item, more than 60% of the examinees showed consistent response patterns, as specified by the sequential CDMs, and the responses of the 597 examinees who had consistent response patterns for all of the items were included for the fit of the developed model. Note that because a relatively small number of examinees and items was used, the data were originally collected in a longitudinal survey, and a possible misspecification of the
Frequencies and Percentages for the Specified Attempt Patterns in the Longitudinal Mathematical Test.
Note. The total number of examinees was 3,257.
The LTA-approach sequential G-DINA model was used to fit the data because it provided more precise parameter estimation and more accurate attribute classification in the simulation study. The same dataset was fit to the LTA-approach G-DINA model developed for assessing the changes in the attributes (Li et al., 2016). A comparison between the two models provides evidence to support the efficacy of the LTA-approach sequential G-DINA model when the attribute classifications of the two models are comparable. For the fit of the LTA-approach G-DINA model, the examinees’ responses to seven items across three measurement times were included for data analysis, whereas each item was scored polytomously with four ordered response categories, where scores of 1, 2, 3, and 4 were recorded for the attempt patterns of (1, 1, 1), (0, 1, 1), (0, 0, 1), and (0, 0, 0), respectively, when the LTA-approach sequential G-DINA model was used.
Table 5 presents the estimated successful probability of the seven items for each of the reduced attribute profiles in the LTA-approach sequential G-DINA model. Note that different items can require different sets of attributes even though they have the same reduced attribute profile. It can be observed that the majority of items functioned in the manner that the disjunctive (or compensatory) CDMs assumed; therefore, highly constrained CDMs, such as the DINA model, were not the best choice for the fitting model. Result patterns similar to those of the LTA-approach sequential G-DINA model can be observed in the LTA-approach G-DINA model, although the results are not discussed in detail, and the correlation of the estimated probabilities between the two models was as high as .88. Because item exposure control should be considered in a large-scale assessment, further scrutiny of test contents in terms of different probabilities for different attribute patterns was impractical in this empirical analysis demonstration.
Probabilities of Success on the Seven Items for Attribute Mastery Patterns.
Finally, the attribute classification consistency between the two models was examined, and a higher consistency rate indicated that the LTA-approach sequential G-DINA model was comparable with the LTA-approach G-DINA model and was thus more efficient because the test item would be terminated thereafter if an examinee could correctly answer it on an attempt. The results indicate that the mean consistency rates between the two models over the four attributes were .88, .90, and .91 for the three respective attempts, which suggests that the LTA-approach sequential G-DINA model can serve as an efficient and alternative method when allowing examines to answer test items over repeated attempts.
Summary and Discussion
Allowing examinees multiple attempts on test items until they correctly answer questions has received increasing attention and offers many advantages. Sequential IRT models have been introduced for such occasions when latent continuous traits are intended to be measured (Culpepper, 2014). Cognitive diagnostic assessments can be administered to examinees multiple times and allow the examinees multiple opportunities to submit a correct response. Therefore, a new class of sequential CDMs for identifying examinees’ mastery statuses of the attributes over attempts was developed in this study, and two formulations in the sequential G-DINA model were introduced to model the dependences of the attributes on repeated attempts.
The results of two simulation studies demonstrated that the LTA-approach sequential G-DINA model performed better than the higher order approach sequential G-DINA model in both model parameter estimation and attribute classification. A larger sample size increased the accurate item and person parameter estimation, as documented in previous studies (e.g., de la Torre, Hong, & Deng, 2010; Huang & Wang, 2014). Because the simulation experiment should be as reasonable as practical CDM assessments, the generated values were adopted from a real data analysis study (Li et al., 2016), and several factors were not manipulated in this study. For example, different attribute structures could influence the model parameter estimations, and examinees using multiple strategies could interfere with the sequential process; these issues are left for future study.
A longitudinal mathematical test was implemented in the application of the sequential G-DINA to real data analysis because there is little empirical research regarding answer-until-correct diagnostic assessments to date. The responses to each item that were correct on an early attempt but incorrect on a later attempt were removed from the data analysis because the sequential CDMs assume that once examinees succeed on a question in a trial, the following attempts on that question are treated as successful. The LTA approach to modeling the stage sequential change for the attributes was used in the empirical analysis. Comparing the results obtained by fitting the LTA-approach sequential G-DINA model with those obtained by fitting the LTA-approach G-DINA model provided supporting evidence regarding whether the developed model produced comparable estimation efficiency and served as an alternative when repeated attempts on test items were allowed in cognitive diagnostic assessments. The results of the empirical analysis indicated that the two models produced similar estimated probabilities for different attribute patterns and had high consistency in attribute classifications; therefore, the use of the developed model to fit the repeated-attempt data was justified. Because of item exposure control in the longitudinal ability assessment, as in other international large-scale assessments, the
It should be noted that the data used for demonstration in the empirical example pertained to a low-stakes assessment; therefore, a relatively high proportion of examinees who recorded a correct response to items on a previous attempt but incorrect responses on the following attempts was observed. Whether high-stakes assessments could increase the number of the response patterns that the sequential CDMs assume deserves further investigation. Nonetheless, test developers were encouraged to devote themselves to designing multiple-choice or constructed-response items with diagnostic function and to provide retrieval-based practice for students with the use of partial feedback that can improve learning as students proceed in a sequence of attempts (Grimaldi & Karpicke, 2012; Karpicke & Blunt, 2011; Karpicke & Grimaldi, 2012). In addition, with ongoing advances in computer technology, a test administration and partial feedback system can be constructed in an e-learning environment, and the effectiveness of diagnostic assessments can be improved when using e-learning tools.
Recently, Ma and de la Torre (2016) proposed a sequential CDM for polytomous responses, which could provide more information than the traditional CDMs. In contrast to the models developed in this study, which involved repeated attempts, Ma and de la Torre’s models focused on a sequence of ordered categorical responses within an item and assumed that an item consisted of a finite number of sequential steps and that each of the ordered categories represented a specific step that can be specified by a different set of attributes. Although the two models were formulated from different methodological perspectives, extending the proposed sequential G-DINA model to fit graded-response data that involve different sequential steps would be an interesting topic for future study.
Supplemental Material
Online_appendix – Supplemental material for A Sequential Process Model for Cognitive Diagnostic Assessment With Repeated Attempts
Supplemental material, Online_appendix for A Sequential Process Model for Cognitive Diagnostic Assessment With Repeated Attempts by Su-Pin Hung and Hung-Yu Huang in Applied Psychological Measurement
Footnotes
Acknowledgements
The authors thank the editor and two anonymous reviewers for their constructive comments on earlier drafts of this article. Finally, the authors would like to express their deepest appreciation to Professor Wen-Chung Wang for his persistent encouragement and substantial feedback on their research career over the past years.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The first author was supported by the Ministry of Science and Technology, Taiwan (Grant NO. 106-2410-H-006-057-MY2). The second author was supported by the Ministry of Science and Technology, Taiwan (Grant NO. 106-2628-H-845-001-MY2).
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
