Abstract
A well-known stopping rule in adaptive mastery testing is to terminate the assessment once the examinee’s ability confidence interval lies entirely above or below the cut-off score. This article proposes new procedures that seek to improve such a variable-length stopping rule by coupling it with curtailment and stochastic curtailment. Under the new procedures, test termination can occur earlier if the probability is high enough that the current classification decision remains the same should the test continue. Computation of this probability utilizes normality of an asymptotically equivalent version of the maximum likelihood ability estimate. In two simulation sets, the new procedures showed a substantial reduction in average test length while maintaining similar classification accuracy to the original method.
Keywords
In the past several decades, computerized adaptive tests (CATs) have received much attention in educational and psychological research due to their efficiency in achieving the goal of assessment, whether it is to estimate the latent trait of test takers with high precision or to accurately classify them into one of several latent classes. In the latter case, the adaptive nature of CATs is used in educational testing to make inferences about the location of examinees’ latent ability relative to one or more prespecified cutoff points along the ability continuum. When there is only one cutoff point and two proficiency groups, this type of CAT is commonly referred to as adaptive mastery testing (AMT; Spray & Reckase, 1996; Weiss & Kingsbury, 1984). In this setting, an examinee will pass the test and be declared a “master” if evidence collected based on test responses indicates that the latent ability is above the threshold for mastery.
Among other advantages of AMT over conventional paper-and-pencil tests is the fact that different examinees might receive different sets of test items that best suit their ability levels. When test items are selected adaptively to match examinees’ true ability, high-performing examinees will not be given easy items unnecessarily. Similarly, low-performing examinees will not be under undue pressure when trying to solve very difficult items. In any implementation of AMT, the choice of classification rule plays a vital role. Several methods that have been proposed include the sequential probability ratio test (SPRT; Spray, 1993; Thompson, 2011; Weissman, 2007), the sequential Bayes (SB) approach (Spray & Reckase, 1996), the generalized likelihood ratio (GLR) approach (Bartroff, Finkelman, & Lai, 2008; Thompson, 2011), and the confidence interval (CI) approach (Thompson, 2011; Weiss & Kingsbury, 1984). All of these approaches stem from the framework of sequential analysis (Wald, 1947), whereby inferences are continually made after each additional observation is obtained. Ideally, inferential procedures based on sequential analysis should be allowed to continue indefinitely until the desired confidence level is reached. In practice, however, there is usually a fixed upper bound n on the number of observations that the practitioners are willing to obtain. In the context of AMT, n is the maximum number of test items that any examinee can receive before a classification decision is made. When termination of the test is forced to occur after n observations are obtained, it is said that a truncated sequential procedure is being used.
Further improvements to sequential procedures for classification purposes have been proposed in the literature. Eisenberg and Ghosh (1980) discussed a modification rule called curtailment that allows for a sequential procedure to be stopped early if no subsequent observations can alter the final classification decision. By this specification, the curtailed version of a sequential procedure always yields the same classification results as the original procedure. Therefore, the two methods always have the same proportion of correct decisions (PCDs). However, the curtailed version has lower average test length (ATL) due to the possibility of early stopping. As a result, curtailed truncated sequential procedures could be useful in many different applications, including in clinical trials and educational testing. In educational testing, administering the fewest possible number of items while trying to make accurate classification decisions is important because a shorter test reduces the operational cost of assessment as well as alleviating the problem of overexposing items in high-stakes testing programs.
As a further improvement to the aforementioned idea of curtailment, a sequential procedure can also be halted during interim analyses when subsequent observations have only a small probability of altering the final results of analysis. Within this framework of stochastic curtailment that was first introduced by Lan, Simon, and Halperin (1982) in the context of clinical trials, several formulations are available depending on how the probability of whether the final results of analysis will change should the sequential procedure continue is calculated. Regardless of what approach is followed, many applications of stochastic curtailment have arisen in various fields. In the particular context of educational measurement, Finkelman (2008, 2010) used stochastic curtailment as a stopping rule for AMT in addition to the original truncated SPRT (TSPRT) and its curtailed version. It was shown that the stochastically curtailed TSPRT, albeit having slightly lower PCD than that of the original version, greatly improves the ATL whether or not the test is constrained from content balancing or exposure control perspectives.
Although curtailment and stochastic curtailment have both been applied in the AMT setting, their applications have only been studied when the TSPRT or the GLR approach is used as the stopping criterion. This fact limits their usefulness because the TSPRT or the GLR approach might not always be the stopping method of choice. Previous research has found that both the TSPRT and the GLR stopping criteria work best when test items are selected based on maximum Fisher information at the cut-off point (Reckase & Spray, 1994; Thompson, 2011). In terms of item bank structure, this requires that many items have high information at the cut-off score. If the item bank is relatively flat in terms of its information function, selecting test items based on maximum Fisher information at the interim ability estimate is more appropriate. In that case, the CI approach works best as a stopping rule due to the decreasing conditional standard error of measurement (CSEM; Thompson, 2007). In addition, because the TSPRT and the GLR stopping criteria are typically used in conjunction with the item selection method that maximizes Fisher information at the cut-off point, the same set of items will be selected for every examinee. Both stopping criteria are therefore prone to the problem of overexposure of test items. On the contrary, different sets of items will be selected for different examinees under the CI approach as it is usually coupled with an adaptive item selection method. Therefore, the distribution of item exposure rates might benefit from using a CI-based method, although a rather uneven distribution is still likely to be observed if item selection is based on Fisher information. Preference toward using the CI approach as the stopping rule over the TSPRT could also be attributed to the arbitrariness inherent in the latter. The TSPRT requires that the likelihood function be compared at 2 points close to the cut-off score, with the region along the latent trait continuum falling between those 2 points referred to as the “indifference region” (Thompson, 2007; Wald, 1947; Weissman, 2007). The TSPRT thus assumes that the practitioners are indifferent to the fact that the probability of misclassifying examinees is typically greater than the nominal Type I error rate within this region (Y.-C. Chang, 2005), an assumption that some practitioners might not be willing to make. It is also not clear how wide the indifference region should be, a factor that further affects how early the AMT can terminate. This dependence of the TSPRT on the width of the indifference region is explained by Thompson (2011).
The purpose of this article is to examine whether the CI stopping rule for AMT can benefit from the use of curtailment and stochastic curtailment. In particular, it is investigated whether curtailment and stochastic curtailment can substantially reduce the ATL of the CI stopping rule, without compromising classification accuracy. If so, the curtailed and stochastically curtailed versions of the CI stopping rule would be attractive options under the circumstances (outlined earlier) where the CI approach is preferable to the TSPRT and the GLR.
In constructing the CI for the person ability parameter, an interim ability estimate needs to be available. While the maximum likelihood (ML) estimator is commonly used for this purpose, it is sometimes not uniquely defined when the three-parameter logistic (3PL) model (explained below) is used, as it is in this article. Therefore, for the purposes of this study, the modified version of the ML estimator introduced by Chang and Ying (2009) will be used. This modified estimator is unique whenever it exists and is asymptotically equivalent to the original ML estimator. It is this modified ML estimator that will guide adaptation of curtailment and stochastic curtailment to the framework of the CI stopping rule.
Item Response Theory (IRT) and the Truncated CI Stopping Rule in AMT
IRT
In AMT, the classification decision for examinees is built upon the relationship between their latent ability and item responses. This relationship is commonly characterized via IRT models. One of the most common models used when test items are scored dichotomously (i.e., correct or incorrect) is the 3PL model. With
where
Using the model introduced earlier, the likelihood function upon observing an examinee’s responses to the ith test item is given by
Truncated CI Stopping Rule in AMT
The idea of using the CI stopping rule for classification of examinees in AMT originated from Weiss (1983). With
where
The CI for θ can be constructed following Equation 2 even when
When the examinee’s true ability is very close to the cutoff score, the desired early stopping scenario might not occur. The CI for θ constructed around
Curtailed and Stochastically Curtailed CI Stopping Rules in AMT
Curtailed CI Stopping Rule
As explained in the previous section, early stopping using the truncated CI stopping rule in AMT can occur any time at or after n0, but before n items are administered if either the lower bound of the interim CI exceeds the cut-off score
The above formulation of the truncated CI stopping rule in AMT is simple and thus can be implemented very easily in practice. To see why further improvement using curtailment is beneficial, consider, for example, the following hypothetical situation: On a test with n = 6 and cut-off point
The choice of item parameters in the preceding example was based on the assumption that an ideal item pool was available; that is, for any updated ability estimate, there exists an item with difficulty parameter close to the interim ability estimate. Although this setting might not be realistic in practice, the possibility of shortening AMT when the final classification decision is certain—or nearly certain—should be clear. It is in this regard that the notion of curtailment and stochastic curtailment will be used. Both methods strive to make the same final classification decision as the truncated CI stopping rule. With curtailment, the truncated CI stopping rule in AMT can be modified as follows: Let DC and KC be the classification decision and stopping time, respectively, of the curtailed CI stopping rule. Then, for any
As seen from the preceding formulation of the curtailed CI stopping rule, the authors only stop early to declare mastery if the current ability estimate is at or above the cutoff point and the examinee has a probability of 1 to be declared a master at the end of the test. This can be checked in practice by calculating the examinee’s final ability estimate if all future responses were to be incorrect. If that final ability estimate is still at or above the cutoff score, the test is stopped, and the current mastery classification decision is preserved. Similarly, the authors only stop early to declare nonmastery if the current ability estimate is below the cutoff point and the examinee has a probability of 1 to be declared a nonmaster at the end of the test. To check this, the final ability estimate if all future responses were to be correct is calculated. If the resulting final ability estimate is still below the cutoff score, then the test is stopped, and the current nonmastery classification decision is preserved.
As explained earlier, when the CI stopping rule is used in AMT, test items administered to an examinee are selected to maximize information at his or her interim ability estimate. In the context of curtailment, the identity of all future items to which the all-correct or all-incorrect responses (alluded to earlier) will be assigned needs to be known. The identity of the next test item can easily be obtained by utilizing the test responses given so far. The identity of all other future test items can be figured out iteratively by first calculating the examinee’s ability estimate after the next item is answered correctly (if all-correct future responses are assumed) or incorrectly (if all-incorrect future responses are assumed), then choosing another item that maximizes information at the new ability estimate, and repeating the process. As another alternative, a “representative set” of items can be used as surrogates for those unknown future items. When performing curtailment and stochastic curtailment of the TSPRT, Finkelman (2008) suggested that the representative set consist of the
Stochastically Curtailed CI Stopping Rule
Standard Formulation of Stochastic Curtailment
A more aggressive modification to the truncated CI stopping rule in AMT is obtained by using stochastic curtailment. Here, instead of requiring that the probability be equal to one that the same final classification decision is obtained should the test continue, the only requirement is that the probability be at or above a specified threshold. For an examinee currently considered a master (i.e., an examinee whose current maximum likelihood estimate [MLE] is greater than or equal to the cutoff point), his or her probability of remaining a master is required to be at least
set KS = k and DS = N (i.e., stop testing and declare nonmastery) if
continue testing to stage k+ 1 otherwise. If k = n, the test is terminated and DS = M is set if and only if
The previous formulation of the stochastically curtailed CI stopping rule is very similar to that of the curtailed CI stopping rule presented earlier. In particular, when
Probability Calculations of the Stochastically Curtailed CI Stopping Rule
As a first alternative to evaluate the probability calculations in Equations 3 and 4, the authors consider using
As a second alternative, the lower or upper end point of the interim CI can be used for
The authors note that in each of the stopping rules introduced above, the ML ability estimate has been used to construct the end points of all CIs. In particular, in the case of the stochastically curtailed CI stopping rule, the ML estimate of θ is used not only in constructing the CI that serves as a criterion for early stopping but also in evaluating the probabilities in Equations 3 and 4. The rest of this section is devoted to explaining how those probabilities can be evaluated using an alternative to the ML estimate due to the latter’s minor drawback explained next.
An Alternative to the MLE: The Modified MLE
At any stage k during the test, the ML estimate
When the 3PL model is used, the likelihood equation simplifies to
The solution to Equation 5, which is called the modified ML estimate of θ, is unique whenever it exists because the left-hand side of 5 is a monotone decreasing function of θ. A necessary and sufficient condition for existence is as follows:
For large n, Equation 6 holds almost certainly, and the corresponding solution
Select the first item with parameters
For each
With
In AMT with the curtailed and stochastically curtailed CI stopping rules, the probabilities in Equations 3 and 4 need to be evaluated, conditioning on all item responses that have been observed. Therefore, to use the consistency and asymptotic normality results of H.-H. Chang and Ying (2009) in AMT, a slight generalization is needed for conditional distributions. For this purpose, Theorem 1 will be used, which is available in Appendix A as an online supplement to this article. Using the modified ML estimate of θ proposed by Chang and Ying and Theorem 1, stochastic curtailment for the CI stopping rule in AMT can now proceed by making use of asymptotic theory. Using the result of Theorem 1, the probability calculations in Equations 3 and 4 can now be approximated by a normal probability.
In practice, the asymptotic result based on Theorem 1 is used to decide whether to terminate early only when k is far less than n. If k is close to n (say, with five remaining items on the test), the probability in Equation 3 can be evaluated using exact calculations. This is done by considering all possible response patterns that the examinee can generate for the hypothetical future items from the representative set. With five remaining items on the test before reaching the maximum test length, an exact calculation can be evaluated by considering all 25 = 32 response patterns that the examinee can generate when answering those five items selected from the representative set. A hypothetical final ML ability estimate is then computed in each case, and the proportion of all estimates that meet or exceed the cutoff point is then evaluated, weighted by the respective probability of each response pattern. The probability calculation in Equation 4 is obtained analogously via the response patterns for which the final ML estimate does not reach the cutoff point.
Simulation
The previous section has described motivation for using curtailment as well as stochastic curtailment in conjunction with the truncated CI stopping rule in AMT. In this section, results of a simulation study are presented to provide comparisons between the stopping rules in terms of their ATL as well as PCD. The four stopping rules studied are the original truncated CI stopping rule, the curtailed CI stopping rule, and the two formulations of the stochastically curtailed CI stopping rule described in “Stochastically Curtailed CI Stopping Rule ” section. In the simulation set presented herein, the stopping rules were compared when test items were selected purely based on a psychometric criterion. Another simulation set (available in Appendix B as an online supplement to this article) compared the stopping rules when other test constraints such as content balancing and item exposure control were considered.
Simulation Design
The 3PL model was used with an item pool consisting of 500 items with IRT
To avoid the multiple roots problem of the likelihood equation of the 3PL model as discussed in the “An Alternative to the MLE: The Modified MLE” section, the modified ML estimate of H.-H. Chang and Ying (2009) was used for all four stopping rules. In particular, the truncated CI stopping rule was based on the CI of θ that was constructed at each stage following 2, only with the modified ML estimate of
In all stopping rules and both simulation sets, test items were selected following the algorithm proposed by H.-H. Chang and Ying (2009) as described in “Stochastically Curtailed CI Stopping Rule” section. In particular, Equation 6 was checked at each stage during the test. If the condition was met,
Similar to Finkelman (2008, 2010), each simulation set was conducted to allow for a matched comparison between the different stopping rules. For every simulated examinee, the four stopping rules were run simultaneously, using the same response pattern to make a classification decision. If one stopping rule ended before the others, the remaining methods were allowed to continue until they reached their respective stopping time. To evaluate accuracy of classification decisions, a correct decision was defined as mastery if the true ability of the examinee was greater than or equal to the test cutoff point and nonmastery otherwise.
Simulation Results
Table 1 presents the PCD, ATL, as well as average losses (explained in the following) for the four stopping rules being compared.
Simulation Set 1: Proportion of Correct Decisions, Average Test Lengths, and Average Losses With Cw = 100.
Note. CI = confidence interval; SC = stochastic curtailment; ML = maximum likelihood; PCD = proportion of correct decisions; ATL = average test length.
Comparing the original (truncated) CI stopping rule with its curtailed version, the difference in PCD was never higher than 0.009 = 0.9% at any ability value. In the opening section, it was explained that in theory, the PCD of the curtailed stopping rule would always be the same as that of the original method. The differences observed here between the PCDs of the two methods can be understood by revisiting the notion of “representative set” used in the simulation. Due to the adaptive nature of the item selection algorithm, the identity of future items that would be administered should the test continue was unknown. A set of approximate future items was therefore needed for the curtailed stopping rule to compute the final modified ML ability estimate of each examinee under hypothetical all-incorrect (or all-correct) future responses. In reality, the original CI stopping rule might proceed with a different set of future items chosen adaptively based on the examinee’s subsequent responses. If the two sets of items are exactly the same, then the PCDs of the two methods will also be the same. This is the case, for example, if test items are chosen based on maximizing Fisher information at the cutoff point as is commonly the case when the TSPRT is used as a stopping rule in AMT. Having seen that the curtailed CI stopping rule did not sacrifice much classification accuracy compared with the original method, it is of interest to see how many test items the former can save. Comparing the ATL of the original CI stopping rule and that of its curtailed version, the latter was able to save an average of 2.62 items across all 13
Turning to the more aggressive stochastically curtailed CI stopping rule (i.e., the rule whereby the probability calculations in Equations 3 and 4 are evaluated at the modified ML estimate
The more conservative stochastically curtailed CI stopping rule (i.e., the rule whereby the probability calculations in Equations 3 and 4 are evaluated at the lower and upper end point of the CI of
To further evaluate whether the shorter test lengths of the curtailed and stochastically curtailed stopping rules are worth considering despite their slightly lower PCDs, each of these three methods was compared with the original CI stopping rule using a classification efficiency index proposed by Vos (2000). With
This index, which ranks different stopping rules by taking both PCD and ATL into account, has also been used by Finkelman (2008, 2010) in the context of AMT with variants of the TSPRT stopping rule.
Table 1 presents the average loss at each
Summary and Discussion
The purpose of this article was to implement curtailment and stochastic curtailment in the context of AMT with the CI stopping rule. The goal of these methods is to provide savings of test items without unduly compromising classification accuracy. Previous research (Finkelman, 2008, 2010) had demonstrated the success of both methods when the TSPRT stopping rule is used. More recently, stochastic curtailment was also applied in conjunction with the GLR stopping rule and was shown to yield increased efficiency while maintaining similar accuracy (Huebner & Fina, 2014). However, as explained in the Introduction, no previous research had applied curtailment or stochastic curtailment to the CI stopping rule, which is preferable to the TSPRT and the GLR in some testing contexts.
The results of the simulation study confirmed the usefulness of the methods in terms of their relative ATLs and PCDs compared with the original CI stopping rule. On a test with a maximum test length of 50 items and no constraints, the more aggressive stochastically curtailed stopping rule reduced the ATL of the original CI stopping rule by an average of more than 11 items without reducing the PCD by more than 3.3%. Based on the loss function in Equation 7 with
The results presented in this article were obtained using the modified ML ability estimate of H.-H. Chang and Ying (2009). This ability estimate can be computed fairly easily by conducting a grid search of the solution to Equation 5 when the inequality in Equation 6 holds. Otherwise, the interim ability estimate will be assigned a value from a decreasing sequence
In this article, the authors focused on evaluating whether curtailment and stochastic curtailment improve the CI stopping rule. With evidence favoring the curtailed and stochastically curtailed CI stopping rules, a possible direction for future research is to compare them with the curtailed and stochastically curtailed versions of the TSPRT and the GLR. In addition, performance of the methods also needs to be further investigated under different structures of item pools, different minimum and maximum test lengths, and other test constraints in addition to content balancing and item exposure control. Extending the methods to classification tests with more than 1 cutoff point also provides fruitful research opportunities, as many statewide tests under No Child Left Behind typically classify students to multiple proficiency groups (Finkelman, 2010).
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
