Abstract
Multidimensional computerized adaptive testing (MCAT) provides a mechanism by which the simultaneous goals of accurate prediction and minimal testing time for a screening test could both be met. This article demonstrates the use of MCAT to administer a screening test for the Computerized Adaptive Testing–Armed Services Vocational Aptitude Battery (CAT-ASVAB) under a variety of manipulated conditions. CAT-ASVAB is a test battery administered via unidimensional CAT (UCAT) that is used to qualify applicants for entry into the U.S. military and assign them to jobs. The primary research question being evaluated is whether the use of MCAT to administer a screening test can lead to significant reductions in testing time from the full-length selection test, without significant losses in score precision. Different stopping rules, item selection methods, content constraints, time constraints, and population distributions for the MCAT administration are evaluated through simulation, and compared with results from a regular full-length UCAT administration.
Keywords
Introduction
In a selection setting, significant savings could be realized if potential applicants could be administered a screening test beforehand that would yield an accurate prediction/estimation of their true score with a minimal amount of testing time. The selection test could then be administered only to examinees whose predicted/estimated performance from the screening test suggests they would meet the qualification requirements on the full-length test. Multidimensional computerized adaptive testing (MCAT) provides a mechanism by which the simultaneous goals of accurate prediction and minimal testing time for the screening test could both be met. This is similar to previous research for paper-and-pencil (PP) tests that showed using multidimensional item response theory (MIRT) can improve the precision of ability estimates by borrowing information from each content domain (Yao, 2010; Yao & Boughton, 2007), while also considering the response time for each of the content domains in a CAT mode.
This article demonstrates the use of MCAT to administer a screening test for the Computerized Adaptive Testing–Armed Services Vocational Aptitude Battery (CAT-ASVAB) under a variety of manipulated conditions. CAT-ASVAB is a test battery administered via unidimensional CAT (UCAT) that is used to qualify applicants for entry into the U.S. military and assign them to jobs. The primary research question being evaluated is whether the use of MCAT to administer a screening test can lead to significant reductions in testing time from the full-length selection test, without significant losses in score precision.
To answer the primary research question, a simulation study is conducted using ASVAB item parameters to administer a 20-item screening test under a variety of manipulated conditions: (a) fixed-length test versus variable-length test; (b) content constraints versus no content constraints; (c) three conditions regarding response time management: no response time constraint, Methods 1 and 2 with response time constraints; (d) four levels of correlations between domains for the population distribution; and (e) three sets of CAT pools. The manipulated conditions were chosen to help evaluate how best to implement a CAT-ASVAB screening test under MCAT administration. The results for all varying conditions on the screening test are compared over 20 replications by examining average bias in ability estimates, average absolute bias, correlations between estimated abilities and their true values, average response time used, and average test length. Results from the screening test are also compared with results from a full-length unidimensional CAT-ASVAB administration using the same criteria.
Multidimensional Item Response Theory (MIRT) Models
Following the notation of the MIRT model in Yao and Schwarz (2006), for a dichotomously scored item
where
The item and test information function (indicated by
In this study, Maximum a posteriori (MAP) is used by finding the mode that maximizes the posterior distribution. Yao (in press) has found that MAP yields better precision than maximum likelihood (MLE). Strong priors are applied, which may yield better precision.
Applications
Data
The purpose of the CAT-ASVAB screening test is to predict/estimate a score on the Armed Forces Qualification Test (AFQT). AFQT scores are computed from four ASVAB subtests, Arithmetic Reasoning (AR), Word Knowledge (WK), Paragraph Comprehension (PC), and Mathematical Knowledge (MK), and are used for selection into the military.
Simulation studies using three sets of item pools are used to evaluate the MCAT approach to estimating AFQT scores. The item parameters used to generate responses come from real item pools for the AR, WK, PC, and MK tests. The first item pool contains a large number of item parameters sampled proportionally from operational CAT pools, and represents an idealized case for a screening test, where there are many items available to select from in the MCAT administration. The pool contains 2,845 items in total, with 807, 839, 446, and 753 items for AR, WK, PC, and MK, respectively.
For the second item pool, the item parameters come from four operational PP forms; they are combined and put together to form the item pool. There are 420 available items in total, with 120, 140, 60, and 100 items for AR, WK, PC, and MK, respectively.
The third item pool contains a small number of item parameters sampled proportionally from operational CAT pools, and represents a more likely scenario for a screening test, where there are relatively few items available to select from in the MCAT administration. The pool contains 253 items in total, with 69, 74, 43, and 67 items for AR, WK, PC, and MK, respectively.
The summary statistics for the third item pool (small number of CAT items) are very similar to the first item pool (large number of CAT items). The second item pool (moderate number of PP items) has lower information than the other item pools (CAT items). The summary statistics for the item pools are displayed in Table 1.
Item Statistics for the Three Item Pools.
Note. The response time is in unit second. AR = Arithmetic Reasoning; WK = Word Knowledge; PC = Paragraph Comprehension; MK = Mathematical Knowledge; PP = Paper and Pencil; CAT = computerized adaptive testing.
The AFQT composite score for the item pools derived from operational CAT pools (Pools 1 and 3) is obtained based on the four content scores, expressed as
The average item response times for the four content areas AR, WK, PC, and MK are 79, 14, 69, and 42, in unit seconds, respectively; they are derived by computing the means of the response times for all the items in each content area for all 4,839,372 examinees taking the CAT-ASVAB from 1998 to 2012. Using separate response times for each item is possible, however, using the average response time for each content area makes the computation simple and should not affect the results of this study. 1
MCAT Simulation Design
Item selection procedures
For the AFQT composite score, the pre-determined weight
For each item
For each item m, compute its information and add to the current posterior information,
Select item
has a minimum value (Method 1, labeled “D1” in the results), or
has a maximum value (Method 2, labeled “D2” in the results).
Update ability
Compute the current error variance
Compute the current subscale estimates
Item exposure rate is not considered in the item selection. 2
Population distributions
Three samples of 3,000 examinees (labeled P1-P3 in the results) are simulated from a multivariate normal distribution with a mean of (0,0,0,0) where correlations between the domains are
Response time constraints
There are two variations for
Content constraints
Constraints are imposed here that require the minimum number of items administered to be 0, 1, 2, or 3 for each of the four content areas under MCAT administration. The conditions are labeled M0, M1, M2, or M3, respectively, in the results.
Fixed versus variable-length tests
A CAT administration process is a cyclical procedure that utilizes a stopping rule (Reckase, 2009; Wainer, 2000). The stopping rule can specify a fixed number of test items be administered (a fixed-length test), or a desired level of precision be reached (a variable-length test). Previous studies on MCAT item selection procedures have been conducted under the fixed-length condition (Huang, Chen, & Wang, 2012; Li & Schafer, 2005; Luecht, 1996; Mulder & van der Linden, 2009; Veldkamp & van der Linden, 2002; W.-C. Wang & Chen, 2004; Yao, 2012, 2014a). However, MCAT selection procedures with varying test lengths using different stopping rules deserve more attention (C. Wang, Chang, & Boughton, 2013; Yao, 2013). In Yao (2013), variable-length stopping rules were used for five MCAT item selection procedures, but the precision requirement and the selection criteria were imposed on the subscores, not the composite scores. This study focuses on the composite score in both the selection criteria and the precision requirement. To achieve the same level of precision, test length may need to vary across examinees, depending on the examinees’ ability, the quality of the item pool, and the CAT selection procedures.
For each variation, two procedures are applied; one is a fixed-length test of 20 items (labeled F in the results), and another is a variable-length test (labeled V in the results) that is stopped when a maximum test length of 20 is reached or when the precision or SEE for the AFQT score reaches a level of 0.25. The choice of the precision level depends on the quality of the item pool (i.e., information); a precision level of 0.25 is appropriate for the item pools in this study, and it is comparable with the precision level for a test of fixed length 20. 3 The simulation conditions are summarized in Table 2.
Simulation Conditions Under MCAT Administration.
Note. MCAT = multidimensional computerized adaptive testing; F = fixed length; V = variable length; D1 = Method 1; D2 = Method 2.
Evaluation Criteria
A number of criteria are used to compare results across the different conditions. They include (a) the bias and absolute value of the bias, averaged over replications and examinees (labeled Average Bias [BIAS] and Average Absolute Value of the Bias [ABSBIAS], respectively); (b) the correlation between estimated and true ability values, averaged over replications; (c) the reliability of the ability estimates; (d) false positive and negative rates; (e) response time, averaged over examinees and replications; and (f) test length, averaged over examinees and replications. The reliability is the average of the reliabilities over n replications and the reliability is the square of the correlations between the estimates and the true values. False positive and negative rates at two cut score points are examined: AFQT scores of 31 and 50.
4
False negative is defined as the percentage where the true value is classified as pass but the estimate is classified as fail; false positive is defined as the percentage where the true value is classified as fail but the estimate is classified as pass. AFQT cut scores of 31 and 50 correspond to
Results
Comparison of MCAT screening test and UCAT full-length test
First, results are compared between the MCAT screening test and the UCAT full-length test for population P4 using Pool 3; the results are presented in Table 3. The bottom line in Table 3 presents the results for a UCAT simulation under the current operational CAT-ASVAB setting and is considered as the baseline from which to compare performance of the MCAT administrations. As expected, the MCAT results all show much shorter average testing times than the UCAT administration (7-17 min vs. 45 min). The reduced number of items administered (20 or fewer items for the MCAT case vs. 55 items for the UCAT case) is a contributing factor to that finding. As might be expected, reliability is highest for the UCAT case, again corresponding to the longer test length. However, reliability of some MCAT conditions approaches that of the UCAT case, for fixed-length test conditions. On all evaluation criteria, the MCAT results show some loss in precision when compared with the UCAT case; however, the loss in precision does not appear substantial, especially for the fixed-length conditions. For a screening test, a smaller false negative rate (the percentage where the true value is classified as pass but the estimate is classified as fail) is desirable. There are some 20-item MCAT conditions that have a smaller false negative rate than the full-length UCAT test, which means that fewer qualified examinees will be screened out by the screening test. However, a non-qualified examinee (AFQT
BIAS, ABSBIAS, Correlations, Reliability, False Negative and False Positive Rates at Cut Score AFQT = 31 and 50, Average Response Time (in Minutes), and Average Test Length for Population P4 for Pool 3 for MCAT With Test Length 20 and a Full-Length UCAT.
Note. BIAS = Average Bias; ABSBIAS = Average Absolute Value of the Bias; AFQT = Armed Forces Qualification Test; MCAT = multidimensional computerized adaptive testing; UCAT = unidimensional computerized adaptive testing; F = fixed length; V = variable length; D1 = Method 1; D2 = Method 2; M0 = minimum item number is 0; M1 = minimum item number is 1; W1 =
Comparison of MCAT conditions
To further evaluate how MCAT might be implemented operationally, results are compared between item selection Method 1 (D1) and item selection Method 2 (D2), between content constraints (M1, M2, M3) and no content constraints (M0), between weights for the time constraints (W0 vs. W1), between variable-length (V) tests and fixed-length (F) tests, across the three item pools, and across the four sets of samples (P1-P4). For populations P1 to P3 under MCAT administration, content constraints M0 and M1 were imposed. For population P4 under MCAT administration, content constraints M0, M1, M2, and M3 were imposed. In general, it is observed that for item selection Method 1, if there are no time constraints and no content constraints (condition D1M0W1), the procedure selects items from all four content areas, where WK has the most items selected, followed by AR, MK, and PC. The number of items selected are similar in the condition where there are no time constraints and there are content constraints (D1M1W1). For Method 1, when there are time constraints, the procedure selects most of the items from WK, with some from MK. AR and PC items are selected for only some examinees but the numbers are very small. Compared with Method 1, Method 2 selects more items from AR, PC, and MK, and the response time is higher. Tables and figures are used to demonstrate the results. Both summary statistics and individual-level information are examined.
Table 4 summarizes the BIAS, ABSBIAS, correlation, reliability, percentage of misclassification, average response time, and average test length for population P2 (r = .5) for the three pools and for both the fixed-length and variable-length tests. For Method 2, two weights for the times
BIAS, ABSBIAS, Correlations, Reliability, Percentage of Misclassification, Average Response Time (in Minutes), and Average Test Length for Population P2 for the Three Pools With Test Length = 20.
Note. BIAS = Average Bias; ABSBIAS = Average Absolute Value of the Bias; AFQT = Armed Forces Qualification Test; F = fixed length; V = variable length; D1 = Method 1; D2 = Method 2; M0 = minimum item number is 0; M1 = minimum item number is 1; W1 =
To examine the effect of population correlations, Figures 1 and 2 show the correlations, response time, and test length for the fixed-length and variable-length tests for Pool 3 for populations P1 to P3 for content constraints M0 and M1, respectively. As the population correlation increases, response time decreases and the correlation between the estimates and the true values for the four content areas and AFQT increases, especially for AR and MK. As strong priors are applied in the MAP estimate, the subscore estimates for AR and WK are improved as the correlation between the content areas increases. Method 1 with time constraints of

Correlations and response time for a fixed-length test for the three populations P1-P3, two item selection methods Dl and D2, and time constraints and content constraints for Pool 3 with a 20-item test and 20 replications.

Correlations, response time, and test length for a variable-length test for the three populations P1-P3, item selection methods Dl and D2, time constraints and content constraints for Pool 3 with a 20-item test and 20 replications.
Compared with no content constraints (i.e., no minimum required item number, indicated by M0), a slightly larger time is observed when content constraints are imposed (indicated by M1). If time is not considered (indicated by conditions D1M0W1 and D1M1W1) in selecting items, the precision is the best and the test length is the shortest, however, the response time is the largest and, on average, it can be as high as 16.5 min for the fixed-length test and 8.27 min for the variable-length test. Figure 3 shows the response time and test length against their true ability values for all 3,000 examinees for population P4 using Pool 3 for MCAT administration with a 20-item test. Differences between Methods 1 and 2, between fixed-length and variable-length tests, and between time constraints and no time constraints can be visually observed.

Response time and test length against their true ability values for the 3,000 examinees for P4, applied to item Pool 3, with a 20-item test using MCAT.
To examine the differences between fixed-length and variable-length tests and how the stopping rules perform, SEEs are examined for some conditions. Figure 4 displays the SEE against the true ability values for all 3,000 examinees for population P2 with Pool 1 with a 20-item test for both Methods 1 and 2. It is observed that when there are no time constraints and content constraints (conditions FP2D1M0W1, VP2D1M0W1), the SEEs are smaller than 0.25 for most of the examinees. The fixed-length test has a smaller SEE than the variable-length test (e.g., FP2D1M0W1 vs. VP2D1M0W1), as the test length for the variable length is at most 20 items (the length for the fixed-length test). When there are time constraints, the SEEs are higher than those without time constraints (e.g., FP2D1M0W2 vs. FP2D1M0W1). Method 2 has smaller SEEs than Method 1 (e.g., FP2D2M0W2 vs. FP2D1M0W2 and VP2D2M0W2 vs. VP2D1M0W2).

SEEs against their true ability for the 3,000 examinees for Methods 1 and 2 for population P2 with Pool 1 with a 20-item test.
A SAS reg procedure was conducted for the estimated AFQT and the estimated four content domain area scores and the true relations were recovered for all the conditions with the default error. Table 5 lists the number of selected items for each of the four content areas for the fixed-length and variable-length tests for population P2 for content constraints M0 and M1; it is clear that many more items from WK were selected for Method 1 than Method 2.
Number of Selected Items by Content Area for a 20-Item Test With Population P2 and Pool 1 Averaged Over Replications.
Note. AR = Arithmetic Reasoning; WK = Word Knowledge; PC = Paragraph Comprehension; MK = Mathematical Knowledge; F = fixed length; V = variable length; D1 = Method 1; D2 = Method 2; M0 = minimum item number is 0; M1 = minimum item number is 1; W1 =
Discussion
The simulation results suggest that using MCAT to administer a 20-item screening test can shorten testing time substantially over the unidimensional full-length CAT-ASVAB, while still yielding a precise estimate of actual overall ability. The approach demonstrated here could be useful in any selection setting where a short, yet precise screening test is desired to reduce the rate of unqualified examinees taking the full-length selection test.
Two methods in selecting items were studied. Compared with Method 2, Method 1 selects more WK items (smallest response time) and fewer items from the other three content areas (longer response times), has a longer test length, and has lower precision; the precision differences between the two methods are smaller when the correlations between subscores are higher than 0.5—such correlations can be met in practice. For MCAT ability estimates, strong priors are used, and as the correlation between the abilities for the population distribution increases, the correlation between the estimates and the true values for the four content areas and the overall increases.
Variable-length tests may be considered; certain examinees can reach the precision level with much fewer items and much lower response times. However, fixed-length tests may be easier for users to understand and/or justify.
Three different item pools were examined. For pools with more high-quality items (i.e., more informative), the required precision level can be reached with lower response times if a variable-length test is used. As this is a screening test, no item exposure control is considered. Therefore, a larger item pool does not necessarily yield better results than a smaller item pool, even if the item statistics for these pools are similar (i.e., Pool 1 vs. Pool 3).
A key element to a shorter testing time in the MCAT screening test is the modeling of response time in the item selection method. For example, for item selection Method 1, average testing time is much higher when there is no response time constraint (i.e., the weight assigned to response time in the model is set to 0) than when a response time constraint is used (16-17 min vs. 7-8 min, respectively, for a fixed-length test). When response time is accounted for in the item selection model, the result is a much shorter testing time, but the tradeoff is an imbalance in content, as the test with the shortest item response times (WK) dominates the selected items. The addition of content constraints has a minimal effect on reducing the number of WK items administered, mainly because only a small number of items per content area are constrained to be administered. If content constraints are included that result in a more balanced distribution of content, then the likely tradeoff would be an increase in testing time and possible reduction in reliability.
If high reliability is the primary desired outcome, increasing the test length may not necessarily improve the precision level when response time constraints are used, because many items are selected from WK due to its shorter response time, and even more WK items would be expected to be administered, which may not increase the precision level. For condition D1M0W2, additional simulations were conducted for a test length of 30 and a precision level of 0.2, and the correlation/reliability for AFQT score was no better than that for a test length of 20 and a precision level of 0.25. Therefore, a screening test with item selection Method 1, variable-length test, response time constraints, and a precision level of 0.25 could be implemented operationally with no anticipated loss in score precision when compared with a 30-item test length.
A secondary research question regarding how the MCAT administration compares with a UCAT administration of the screening test was answered through additional simulation, namely, would it be worth it to implement MCAT instead of UCAT for an operational CAT-ASVAB screening test? Readers can find the research design and the results in the online version of the article at the appendix.
Compared with UCAT, the precision increased when MCAT was applied. To facilitate the comparison of UCAT and MCAT, a 55-item MCAT was also simulated. MCAT outperformed UCAT at both the 55- and 20-item test lengths with regard to reliability and test time, but there was a greater gain in performance with the shorter test length. However, there are some UCAT 20-item screening test conditions that look like they could be acceptable in practice, in terms of their relative reliability and testing times. A good UCAT screening test condition can be realized but is best chosen based on the success of the MCAT screening test. A requirement for the minimum number of items in each of the content areas increases the reliability for the MCAT administration, but also increases the average response time. If minimizing testing time is the primary interest in a screening test, the minimum required item number should be less than or equal to 3; under these conditions in MCAT, the rest of the items are selected from WK for most of the examinees, especially for Method 1. More items from the test with the shortest response times (relative to the other three content areas) will decrease both the reliability and the average response time.
In conclusion, MCAT certainly looks to be a promising approach to a screening test; MCAT can incorporate content constraints and response time simultaneously for better prediction and short testing time. In practice, the balance between the precision and the response time (the weight for the two methods) needs to be adjusted based on their values in the application data. The tradeoffs between reliability, testing time, content representation, and the labor in implementing MCAT methods are all factors that need to be considered in choosing an approach for operational implementation of a screening test.
Footnotes
Acknowledgements
We thank the reviewers and the editor for their valuable input on the earlier version of this manuscript.
Authors’ Note
The views expressed are those of the authors and not necessarily those of the Department of Defense or the U.S. government.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
