Abstract
The Air Force Officer Qualifying Test (AFOQT) is the primary selection tool for officer applicants in the U.S. Air Force (USAF) for nearly seven decades. The AFOQT is revised and modified periodically, with rigorous equating and linking effort to ensure comparability and connectivity across forms. The most recent version of AFOQT is Form T that includes 10 cognitive ability and knowledge subtests. Despite the continuing validation effort of the AFOQT across forms, it was mostly directed to the general population of officer applicants, but not to any specific subpopulation. The current investigation reported three studies in an attempt to provide evidence for factor structure and criterion-related validity of AFOQT Form T for pilot applicants via four analytical approaches: meta-analysis, exploratory factor analysis (EFA), confirmatory factor analysis (CFA), and structural equation modeling (SEM). The results suggested that AFOQT Form T data are best represented by a bifactor model with a general ability and four specific abilities, and that each latent construct has a distinct predictive utility for pilot performance criteria.
The Air Force Officer Qualifying Test (AFOQT) is a multiple-aptitude test battery composed of 10 ability subtests in the most recent version, Form T. The battery is designed to assess a variety of cognitive abilities (e.g., verbal, quantitative, spatial) and aviation-related knowledge constructs. The AFOQT is used for officership qualification and initial job placement for officers selected in the U.S. Air Force (USAF). The subtest configuration of AFOQT has changed from 16 ability subtests (Forms O, P, and Q), to 11 subtests (Form S), and, most recently, to 10 subtests (Form T). Its reliability has been studied extensively (e.g., Berger et al., 1990; Carretta et al., 2016; Glomb & Earles, 1997; Skinner & Ree, 1987) and was found based on Cronbach’s alpha at a range from .71 to .91 (Barron et al., 2016; Carretta et al., 2016). The AFOQT has been validated for performance criteria of officer training (Roberts & Skinner, 1996) and several Air Force occupations (e.g., Arth, 1986; Carretta, 2010; Finegold & Rogers, 1985). However, most of its criterion-related validation efforts were directed to aviation jobs such as flying, navigation, and air battle management (e.g., Barron et al., 2016; Carretta, 2008; Carretta & Ree, 1995; J. F. Johnson et al., 2017; Olea & Ree, 1994). For job decision purposes, operational composites are typically derived from the AFOQT using different sets of subtests. Currently, sex overlapping composite scores are computed from the entire AFOQT Form T battery: Pilot, Combat Systems Officer, Air Battle Manager, Academic Aptitude, Verbal, and Quantitative (Aguilar, 2017). In addition to cognitive ability tests, Forms S and T also included the Self-Description Inventory+, a personality inventory measuring the Big Five factors (Neuroticism, Extraversion, Openness, Agreeableness, and Conscientiousness) and Machiavellianism trait (Aguilar, 2017). Moreover, a Situational Judgment Test was introduced in Form T as an experimental assessment, which reflects a tendency to make use of different measurement methods for enhancing the dependability of selection assessment (Carretta et al., 2016).
Table 1 shows the AFOQT subtests across the last five forms. A configuration of 16-subtest AFOQT appeared first in Form O (Skinner & Ree, 1987) and was maintained across two succeeding forms: Form P (Berger et al., 1990) and Form Q (Glomb & Earles, 1997). The subsequent versions of AFOQT excluded some of the subtests; five subtests were omitted from Form S (Drasgow et al., 2010), and six subtests were omitted from the current Form T (Carretta et al., 2016). Most subtests in the 16-subtest AFOQT were carried forward from earlier versions (Skinner & Ree, 1987), and the retained subtests in later versions were all taken forward from the 16-subtest version. Because of this, the accumulation of data (e.g., subtests’ correlation matrices) from different AFOQT versions for potential meta-analysis study is well grounded, especially with the rigorous equating and alignment procedure regularly conducted when replacing an old version with a newer version (e.g., Glomb & Earles, 1997; Steuck et al., 1988).
The AFOQT Configuration Across Five Different Forms and CHC-Based Classification for Its Subtests.
Note. Description of AFOQT subtests was obtained from Carretta and Ree (1995); the CHC-based classification of subtests was obtained from Stanek and Ones (2018). AFOQT = Air Force Officer Qualifying Test; CHC = Cattell–Horn–Carroll theory.
Introducing a new AFOQT form usually takes a long process, to ensure validity and reliability of the test and to ascertain comparability with previous forms. The AFOQT developers make efforts to construct a form equivalent to its successor in terms of test content and length, item difficulty and discrimination, and stylistic features (Glomb & Earles, 1997). Equating analysis is carried out to empirically derive linkage between composite scores (e.g., Pilot, Academic Aptitude, Verbal, Quantitative) of the new and former AFOQT forms. Different equating methods had been used in this effort, such as linear equating (based on z-score values) and equipercentile equating (based on percentile ranks). The extent of similarity between the distribution of the equated new test scores and the distribution of the reference test scores impacts the decision as which equating method to use; the method that yields the utmost similarity in the two distributions is the one generally preferred. For example, Steuck et al. (1988) applied equipercentile equating with various polynomial smoothing methods (linear, quadratic, and cubic) for linking AFOQT composites of Forms P1 and P2 with those of Form O. Similarly, Glomb and Earles (1997) found that equipercentile equating method with cubic smoothing was a better alternative to linking AFOQT Forms Q1 and Q2 with Form P. The outcome of equating analyses is conversion tables through which scores on the newly developed two parallel test forms (e.g., T1 and T2) are linked with scores on the previous form.
Despite the fact that the construction of AFOQT was influenced by Thurstone’s (1938) multiple-aptitude theory (Ree & Carretta, 1996), it has not been attached to any specific intelligence theory. Drasgow et al. (2010) found evidence indicating that the psychometric g tapped by the AFOQT is in line with Carroll’s (1993) three-stratum theory of intelligence, and suggested that its underlying latent structure is similar to that found in many non-military contexts. It is still important, nonetheless, to draw a connection between AFOQT factor structure and the common intelligence models founding ability test batteries. Human cognitive structure is often described via statistical models specifying its posited organization (Reeve & Bonaccio, 2011). Some models are more indicative of particular theories than others (Brunner et al., 2012). For example, a single-factor model is a good representation of Spearman’s theory (Spearman, 1904), while a higher order factor model containing three levels of abstraction is more indicative of Carroll’s (1993) three-stratum theory. Attempting different modeling techniques when assessing factor structure of cognitive data can be useful and, in fact, is a recommended strategy (Jöreskog & Sörbom, 1993). Although no agreed-upon model has yet been reached for the best conceptualization of cognitive ability data, theoretical considerations usually guide researchers’ choices on how the data are best modeled. As observed by Molenaar (2016), for example, Horn’s (1968) arguments were in favor of the correlated-factor model, A. R. Jensen’s (1998) and W. Johnson and Bouchard’s (2005) arguments were in support of the higher order factor model, and Gignac’s (2008) and Beaujean’s (2015) arguments were in favor of the bifactor model. In fact, Beaujean (2015) was conveying a different perspective for Carroll’s (1993) view on intelligence suggesting that Carroll’s published views are better represented by a bifactor model rather than a higher order model. In a comparison between three statistical models—correlated-factor, higher order, and bifactor models—Morgan et al. (2015) recognized a wide overlap of fit values across the models irrespective of the true structure from which they were obtained, and they suggested to judge the models on substantive and conceptual grounds. In this study, we factor analyzed AFOQT data to explore its latent structure and to test which cognitive model represents these data more fittingly (i.e., single-factor, correlated-factor, higher order factor, and bifactor models).
Factor-analytic studies of AFOQT are relatively small in number. There exist four factor-analytic investigations for the AFOQT, of which only one study assessed the data by exploratory factor analysis (EFA), while the rest applied a confirmatory approach through confirmatory factor analysis (CFA) procedures. This group of studies assessed the three AFOQT forms described above: the 16-subtest version, Form S, and Form T. Skinner and Ree’s (1987) study of AFOQT Form O was the only study that determined its factor structure through an exploratory technique and, very often, is the reference most cited to justify a specific factor structure choice. One to six factor solutions were examined in this study for AFOQT data, collected from 3,000 officer applicants. Through principal factor analysis and oblique rotation, the results suggested a five-factor structure to be the best representation of the data. The five ability factors were named Verbal, Quantitative, Space Perception (spatial and mechanical), Perceptual Speed, and Aircrew Interest/Aptitude. Intercorrelations among the five factors ranged from .22 (Verbal and Perceptual Speed) to .50 (Space Perception and Perceptual Speed), with an average of .36, a pattern that was interpreted as a possible existence of two to three higher order factors if it was taken further for a second-order factor analysis (see also Warne & Burningham, 2019).
Carretta and Ree (1996) further examined the same data using CFA procedures. Seven CFA models were assessed in the study, including a single factor, correlated four-factor (based on four operational composites), correlated five-factor (based on previous EFA study), bifactor (three different models), and higher order factor. The best-fitting model among the seven was a bifactor model containing a general factor (i.e., g, general intelligence, general mental ability, or psychometric g) and the five group factors suggested by Skinner and Ree (1987), with CFI (comparative fit index) of .96, RMSEA (root mean square error of approximation) of .07, and SRMR (standardized root mean square residual) of .03. However, the other two bifactor models, as well as the higher order factor model (Vernon-like model), also showed acceptable fit for the data (CFI > .95 and RMSEA < .08).
Drasgow et al. (2010) examined the factor structure of AFOQT Form S data (11 subtests) that was collected from 12,511 officer applicants from the USAF, with additional goals of establishing measurement invariance across sex (male/female) and race (White/African American/Hispanic/Other groups). A total of nine models were tested in the study, seven of which were defined using item-parceling procedures. Similar to Carretta and Ree (1996), results indicated that the data were best represented by a bifactor model containing a general intelligence factor and five content-specific factors representing Verbal, Quantitative, Spatial, Perceptual Speed, and Aircrew Aptitude/Interest (CFI = .98, RMSEA = .05, and SRMR = .06). Measurement invariance of the AFOQT across sex and racial/ethnic groups was determined in the study.
On the validation of AFOQT Form T, Carretta et al. (2016) assessed five different factor structures: a single-factor model, four correlated-factor model, five correlated-factor model, higher order model with four factors, and higher order model with five factors. Differing from the other two studies, no bifactor models were tested in the study, possibly due to the small number of subtest variables in Form T (10 subtests). Akin to Drasgow et al.’s (2010) study, this study also applied item-parceling technique to provide multiple composites for each subtest to enable factors’ specifications for CFA. Across two independent samples (N1 = 5,681; N2 = 5,199), the best-fitting model was a higher order model with the general intelligence at the top and five lower order factors at the bottom, indicating Verbal, Math, Spatial, Aircrew Aptitude, and Perceptual Speed (Sample 1: CFI = .87, RMSEA = .05, and SRMR = .07; Sample 2: CFI = .91, RMSEA = .04, and SRMR = .05).
Criterion-related validity studies have become an increasingly popular approach for supporting a test’s psychometric properties (e.g., American Educational Research Association et al., 2014). Numerous validation studies have been published and reported for the AFOQT in the effort of validating its subtests and composites for different criteria of job performance (e.g., Carretta, 2010; J. F. Johnson et al., 2017), including flight performance. The recent meta-analysis of ALMamari and Traynor (2020) has synthesized AFOQT literature of pilot performance that spanned nearly five decades of research and provided a quantitative summary for its predictive validity. Based on 32 independent samples from 26 studies, most AFOQT scores (16 subtests and 1 composite) related significantly with pilot performance measures, but with varying degree of predictive power and varying evidence of generalizability across settings. Overall results showed that subtests indicating aviation-related acquired knowledge and aptitude (Aviation Information [AI] and Instrument Comprehension [IC]) and perceptual speed (Scale Reading and Table Reading [TR]) are the best predictors of pilot performance among the 16 subtests, with evidence supporting their generalizable validity. The next best predictors are two subtests from quantitative ability domain (Arithmetic Reasoning [AR] and Data Interpretation [DI]) that showed stronger validities than the remainder of subtests in the battery. Subtests demonstrating verbal ability (VA, Reading Comprehension [RC], and Word Knowledge [WK]) had the weakest relations with flight performance and showed the lowest generalizability evidence across settings (ALMamari & Traynor, 2020). Nonetheless, extending these validation efforts to a higher order construct level can provide deeper understanding of the role of cognitive abilities on pilot performance. Connecting latent ability factors underpinning AFOQT Form T with pilot performance helps to derive more validity evidence for the battery. On this account, the strength of associations between AFOQT Form T ability factors and pilot performance measures was examined in a criterion-related validity design in the present research.
In light of the preceding, it is clear that the effort to examine the AFOQT psychometric properties has continued across the forms. However, there are several limitations to note about previous factor-analytic studies. First, EFA investigations of AFOQT data are rare during the long history of AFOQT use and the regular revisions implemented to the test. Furthermore, a long time has passed since reporting on the only available EFA study (more than three decades ago), although it is often used as a theoretical basis for suggesting a five-factor model for AFOQT data. EFA studies are vital, particularly when an instrument undergoes continued revision, as in the case of AFOQT. Second, two of the three CFA studies used multi-item composite (i.e., parcels) method to deal with the issue of small number of indicators available for the AFOQT short versions (Forms S and T). Although item parcels in CFA have been common in research, the use of this procedure is not always recommended with multidimensional data (see Bandalos, 2002; Matsunaga, 2008; Sass & Smith, 2006), such as those obtained from cognitive test batteries. It can be used, however, when data are non-normally distributed, are coarsely categorized, have a small variable to sample-size ratio, or are believed to produce better model fit over solutions at the original item level (Bandalos, 2002). Third, as highlighted above, previous CFA studies utilized item-level data for assessing AFOQT factor structure, which was facilitated by parceling. Replicating some of these models with subtest-level scores (obtained from subtests’ correlation matrices) instead of item-level data (obtained from the raw scores) can provide further support for the AFOQT factor structure. The use of test batteries’ subtest scores to assess cognition schema is common in psychometric studies (e.g., Canivez et al., 2017; Dombrowski et al., 2018, 2019; McGill, 2020). Fourth, AFOQT investigations of criterion-related validity focused primarily on subtest level, rather than construct level, which may not be sufficient to draw a firm conclusion about the cognitive ability–pilot performance relationships. Despite the usefulness of this effort, our inference about the role of cognitive abilities in shaping pilot performance is only possible to a certain degree, and there remains a need to capture higher order constructs with more precision. This can be achieved relatively well by using structural equation modeling (SEM) procedures that necessitate conceptualizing each construct in a model with multiple indicators, so its operationalization is more accurate (i.e., free of measurement error) and more representative (i.e., broader construct measurement) (e.g., Kline, 2015).
The AFOQT Form T is still in need of an in-depth assessment for pilots’ samples, including (a) EFA investigation; (b) CFA investigation not relying on item-parceling procedures, following the commonly proposed cognitive models; and (c) criterion-related validation of the chosen model(s) for pilot performance. These three goals are addressed in the current research via three separate studies.
General Overview of the Three Studies
The first study in the present work was designed to provide evidence for the internal construct validity of AFOQT Form T for pilot applicants. Because EFA is an efficient means for establishing a test battery’s construct validity (Thompson & Daniel, 1996), this method was preferred as an initial assessment for analyzing AFOQT data. Furthermore, instead of relying on a single data set, this study aggregated and meta-analyzed the correlation matrices of AFOQT subtests reported in prior studies. Cheung and Chan’s (2005) meta-analytic SEM (MASEM) approach was chosen as a methodology for this investigation. However, only the first stage of the two-stage method was applied in the study. The resulting pooled correlation matrix of AFOQT Form T subtests was then examined by means of EFA and, further, by second-order EFA with the Schmid–Leiman orthogonalization procedure (Schmid & Leiman, 1957). The scarcity of AFOQT EFA studies makes this investigation of particular interest and would be a useful addition to the AFOQT factor-analytic literature.
To provide further validation for the solution suggested by the meta-analysis and EFA study, three AFOQT data sets of pilot samples were selected from the pool of correlation matrices that were collected for the meta-analysis study so they can be examined comprehensively by means of CFA. Each selected correlation matrix met three criteria: (a) It included the 10 subtests forming AFOQT Form T, (b) it included correlations between AFOQT subtests and pilot performance criteria, and (c) it had not been analyzed previously with SEM procedures. The selected data sets were separately tested for four model representations: single-factor, correlated-factor, higher order factor, and bifactor models. This is an advantageous investigation for AFOQT Form T factor structure because it provides evidence supporting its construct validity and enhances the understanding of latent variables measured by the battery. Moreover, most factor-analytic investigations of cognitive models stem from intelligence test batteries that are designed for the mere assessment of intelligence construct(s). Conversely, the AFOQT is a selection test battery that has a specific focus on officer recruitment and applicants’ selection who possess aptitude to be successful candidates for the USAF officer jobs (e.g., pilot, air battle manager), which make this investigation a unique and contributory to the literature of cognitive models.
After validating the most plausible model for AFOQT via both factor-analytic procedures, EFA and CFA, an additional examination was then sought focusing on the criterion-related validity (i.e., predictive validity) of ability factors underlying AFOQT Form T versus several criteria of pilot performance. The same three data sets selected for Study 2 were also used in this study. Due to the practical goals usually foreseen when designing a selection test battery, as AFOQT, it would be extremely important for any organization to ensure that the battery is capable of predicting future outcomes for the intended job performance. Hence, the current investigation had practical-oriented goals that complement the theoretically oriented goals sought by Studies 1 and 2.
Study 1: Meta-Analysis and EFA
The goal of this study was to examine the factor structure of the AFOQT by means of EFA procedures. To provide more dependable results, an MASEM was applied on the AFOQT data that were collected from prior investigations. Subtest intercorrelations were meta-analyzed to produce a weighted pooled correlation matrix for further EFA assessment. Due to the space limitation, the meta-analysis part of this study is reported in the appendix, and only the EFA part of the study is reported here. As presented in the appendix, tests of heterogeneity suggest that there is large between-study heterogeneity for most coefficients in the pooled correlation matrix and that a large part of the variance accounted for is at the study level. Therefore, subsequent EFA results should be interpreted with caution. The small number of correlations included in the meta-analysis may be a potential cause of such inflation in the between-study variance (e.g., Sidik & Jonkman, 2007).
The resulting pooled correlation matrix for the 10 subtests comprising AFOQT Form T was then extracted and used as an input for the subsequent factor-analytic examinations. An EFA with Principal Axis Factoring (Fabrigar et al., 1999) and Promax rotation (Gorsuch, 1983) was first applied. For a better judgment of the number of factors to be retained, multiple criteria were used, including the Kaiser’s (1960) MinEigen greater than one criterion, Cattell’s (1966) scree test, Horn’s (1965) parallel analysis (HPA), Velicer’s (1976) minimum average partial (MAP) method, and Sequential Chi Square Model Tests (SMT; Auerswald & Moshagen, 2019). The suggested factor solution was then assessed for interpretability and theoretical conceivability (Fabrigar et al., 1999). To be considered a plausible solution, each factor had to be marked by two or more salient loadings (Gorsuch, 1983). Salient factor pattern coefficients were defined as those loadings equal to or greater than .30 (Child, 2006).
The model suggested by EFA was further analyzed using an exploratory bifactor modeling approach. Specifically, a second-order EFA with the Schmid–Leiman orthogonalization procedure (Schmid & Leiman, 1957) was performed. This analysis helps to understand a test battery presumed to measure higher order factor and correlated traits (Dombrowski, 2015). The explained common variance (ECV; O’Connor Quinn, 2014) and Omega reliability (Omega [ω], Omega subscale [ωs], Omega Hierarchical [ωh], Omega Hierarchical subscale [ωhs]; Reise et al., 2013) were computed to assess whether the AFOQT Form T should be interpreted as primarily unidimensional or multidimensional. The combination of the two exploratory approaches provides evidences supporting AFOQT construct validity and also enhances modeling strategies of subsequent studies. All analyses in this work were performed using R package (R Core Team. 2020).
Results
The pooled correlation matrix was used to perform EFA to assess the AFOQT’s internal structure. The MAP suggested one factor, eigenvalue greater than 1 suggested three factors, and HPA, SMT, and scree plots (Figure 1) suggested four factors. More support is apparent for the four-factor solution of the AFOQT meta-analytic data, and hence, it was suggested as the most plausible underlying factor structure. Overextraction of factors in EFA is arguably more favored than underextraction because loading estimates of true factors tend to include less error in the case of overextraction (Wood et al., 1996). Table 2 shows the EFA results for the four-factor AFOQT model. The extracted factors can be interpreted as a verbal ability (Verbal Analogies [VA], RC, WK, and General Science [GS]), a quantitative ability (AR and Math Knowledge [MK]), an aviation-related acquired knowledge (IC and AI), and a spatial-perceptual ability (Block Counting [BC] and TR). The assessment of factors’ intercorrelations of this solution revealed stronger relations for quantitative ability with verbal ability (.67) and spatial-perceptual ability (.55), but weaker relations for aviation-related acquired knowledge with spatial-perceptual ability (.31) and quantitative ability (.34) factors.

Scree Plot of Eigenvalues Derived From the AFOQT Form T Meta-Analytic Data With Means and 95th Percentile Eigenvalue Estimates From the Parallel Analysis.
Factor Pattern (Structure) Coefficients of AFOQT Form T Meta-Analytic Data of Pilot Applicants From Principal Axis Factor Extraction With an Oblique (Promax) Rotation
Note. Bold font indicates significant loading (>.30). AFOQT = Air Force Officer Qualifying Test.
This solution showed some similarities to the theorized five-factor model suggested for the AFOQT. However, two notable differences are particularly worth mentioning. First, the GS test had its strongest loading on the verbal ability factor (0.38) but loaded weakly on aviation acquired knowledge factor (0.22), although the latter had been suggested as a primary factor for the test (Carretta & Ree, 1996). Second, due to the omission of six subtests from the 16-subtest version of the AFOQT, majority of which were indicative of spatial ability construct, BC and TR subtests formed one broad factor representing spatial-perceptual ability.
Given the plausibility of the four-factor model as revealed by EFA results, this solution was transformed to a second-order model with the Schmid–Leiman orthogonalization procedure. The results are presented in Table 3. After transformation, the four factors of verbal, quantitative, acquired knowledge, and spatial-perceptual ability remained as distinct factors, with the same subtests’ composition. Despite such a convergence with the structure posited in the first-order EFA, however, a notable reduction in factor loadings was noticed, and three subtest loadings were less salient (≅0.25).
Sources of Variance in AFOQT Form T Meta-Analytic Data of Pilot Applicants According to an SL Orthogonalization With Four First-Order Factors.
Note. AFOQT = Air Force Officer Qualifying Test; SL = Schmid–Leiman; IECV = item explained common variance; ECV = explained common variance; ω = Omega; ωs = Omega subscale; ωH = Omega Hierarchical; ωHS = Omega Hierarchical subscale.
The hierarchical g-factor accounted for 48% of the common variance. The general factor also accounted for between 9% (AI) and 83% (AR) of individual subtest variance, with a median of 45.5%. The model-based estimate of internal reliability was .86 for the general factor (ω) and between .60 (spatial-perceptual factor) and .82 (verbal factor) for the specific factors (ωs). The Omega Hierarchical coefficient was estimated to be .68 for the general factor (ωh) and between .18 (quantitative factor) and .53 (acquired knowledge factor) for the specific factors (ωhs). The ωh coefficient suggests that general factor may be sufficient for interpretation, whereas the ωhs coefficients suggest that specific constructs may be inappropriate for interpretation beyond the general factor as little variance exists beyond the general factor (Reise, 2012).
On the whole, this study was useful in understanding the underlying constructs of the AFOQT subtests of Form T. Traditionally, AFOQT researchers had advocated a five-factor model, even when the subtests were reduced in number from 16 to 11 (Form S), and then to 10 (Form T). The four-factor model of AFOQT Form T suggested by EFA in the current study appeared a competitive structure for the five-factor model frequently suggested for the AFOQT, with even more favored characteristics. It met the requirement of simple structure, with each factor marked by two to four salient loadings, and none of the subtests had salient cross-loadings on other factors (≥.30). The hierarchical EFA indicates that a four first-order and one higher order factor model can also be a viable presentation of the AFOQT Form T, even with the weak multidimensional interpretation determined by dimensionality assessment indices.
Study 2: CFA
The goal of this study was to further demonstrate the internal construct validity of the four-factor AFOQT model (verbal, quantitative, aviation-related acquired knowledge, and spatial-perceptual ability) on three separate data sets for pilot applicants. Three different representations of the four-factor model (i.e., correlated-factor, higher order factor, and bifactor models), as well as a single-factor model, were tested in the current study in an attempt to assess how representative they are for the AFOQT data and whether they are replicable across pilot data. This study will provide additional evidence for the construct validity of AFOQT Form T.
Data of this study were obtained from three prior AFOQT investigations, all of which reported intercorrelation matrices between the 16 AFOQT subtests for pilot students, including the 10 subtests remained in the AFOQT Form T. The Online Supplemental Material A includes all three correlation matrices. Of note, these three data sets were among the set of correlation matrices that were meta-analyzed in Study 1. Due to the substantial heterogeneity detected by the meta-analysis for AFOQT subtests’ intercorrelations, the resulting matrix was not used in the present CFA or subsequent SEM analyses. In addition, the located pilot performance measures in the collected data, that would also be required for Study 3’s investigation, varied widely, which reduce the number of cases (k) available for each specific criterion, thus minimizing the opportunity of producing a reliable meta-analysis for each unique ability–performance relationship. Hence, we found that examining separate AFOQT data sets is more appropriate for the sought-after assessment of construct- and criterion-related validity.
The correlation matrix of the first sample was reproduced from Duke and Ree’s (1996) study for 1,082 Undergraduate Pilot Training (UPT) students. The second data set was sourced from Olea and Ree’s (1994) study, which reported two data sets for pilots and navigators. The pilot sample consisted of 1,867 UPT students. The correlation matrix of the third data was reproduced from Arth et al.’s (1990) technical report for 695 pilots, which also included data for 632 navigators. Subjects of all three samples were pilot trainees in their UPT program in the USAF. This program typically consists of three main phases: ground school phase, primary training phase, and advanced training phase. The AFOQT was the primary selection tool used in qualifying the subjects for officer training programs. Subjects in the samples had completed at least a 4-year baccalaureate degree before training. In addition to their qualification on the basis of AFOQT scores, the selected applicants had to meet other selection standards such as academic achievement, medical, moral, physical fitness, personal recommendations, and prior flying experience.
The three data sets were tested for four AFOQT factor structures: single-factor model, correlated-factor model containing four constructs, higher order factor model with general factor at the top and four lower order constructs at the bottom, and bifactor model containing general factor and four specific constructs that were all modeled as lower order factors. This extensive examination of AFOQT Form T factor structure was necessary to uncover the best representation of its data and to facilitate comparison across AFOQT forms.
CFA models were estimated using maximum likelihood (ML). For identification purposes, the loadings on the factors with only two indicators (three of the four factors) were constrained to be equal. Model fit was assessed according to several goodness-of-fit indices, including the CFI (Bentler, 1990), RMSEA (Browne & Cudeck, 1992), and SRMR (Hu & Bentler, 1999). Due to the large sample used in all primary studies, chi-square values (χ2) were not considered for judging model fit (e.g., Schermelleh-Engel et al., 2003; Vandenberg, 2006), although it was reported for completeness. A good fit of the hypothesized model to the observed data requires CFI ≥ 95, RMSEA ≤ .05, and SRMR ≤ .08 (Hu & Bentler, 1999). Given models’ complexity, CFI ≥ 90 and RMSEA ≤ .08 were considered consistent with an acceptable model (Bentler, 1990; Browne & Cudeck, 1992; Schermelleh-Engel et al., 2003; Stevens, 1996). In addition, the relative goodness of fit of the competing models were compared using the Akaike information criterion (AIC; Akaike, 1974) and the sample-size adjusted Bayesian Information criterion (BIC; Schwarz, 1978), where smaller values indicate a better fit. Specifically for bifactor models, several indices were also obtained to determine whether the AFOQT Form T should be interpreted as primarily unidimensional or multidimensional, including the ECV (O’Connor Quinn, 2014), Omega reliability (ω, ωs, ωh, ωhs; Reise et al., 2013), H index of construct replicability (Hancock & Mueller, 2000), Factor Determinacy (FD; Gorsuch, 1983), and the percentage of uncontaminated correlations (PUC; Rodriguez et al., 2016).
Results
Figure 2A to D displays the four CFA models tested in this study. Fit statistics of the four models for each of the three data sets are presented in Table 4. Across the three data sets, bifactor models were the most likely representation of AFOQT data, albeit not excellent, χ2(28) = 340.39, p ≤ .001, CFI = .90, RMSEA = .10, SRMR = .06 for Sample 1; χ2(28) = 358.41, p ≤ .001, CFI = .94, RMSEA = .08, SRMR = .05 for Sample 2; and χ2(28) = 150.47, p ≤ .001, CFI = .94, RMSEA = .08, SRMR = .05 for Sample 3, while single-factor models were the least likely representation. Correlated-factor and higher order factor models had fairly comparable fit statistics but were not as good as bifactor models neither as weak as single-factor models. AIC and BIC values took the same pattern across the three samples, starting from the lowest to the highest: bifactor, correlated-factor, higher order factor, and single-factor models. Such a pattern indicates that the bifactor models are better representation of AFOQT Form T data among the competing models. Table 5 shows the factor loadings for the bifactor models across the three data sets. Subtests’ loadings on both general factor and their corresponding specific factors did not differ notably, with the exception of those marking quantitative ability factor (AR and MK), where they had higher loadings on the general factor, and those marking aviation-related knowledge factor (IC and AI), where they had higher loadings on their specific factors. This may suggest that quantitative ability is a focal element in the construction of general ability, while aviation acquired knowledge is not so important for its construct. The AI subtest, the subtest with the weakest loading on the general factor, loaded negatively on the general factor in Sample 3, which indicates a deviation of its construct from the intelligence-specific constructs, perhaps due to its saturation with domain-specific knowledge related to flying (for discussion, see Schneider & McGrew, 2018, p. 117).

(A) Single-Factor Model for AFOQT, (B) Correlated-Factor Model for AFOQT, (C) Higher Order Factor Model for AFOQT, and (D) Bifactor Model for AFOQT.
Model Fit Indices for Four Competing Confirmatory Factor-Analytic Models Across Three AFOQT Data.
Note. AFOQT = Air Force Officer Qualifying Test; CFI = comparative fit index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; AIC = Akaike information criterion; BIC = Bayesian information criterion.
Subtest Loadings on Their Corresponding Specific Factors and on General Factor Across Three Pilot Samples.
Note. All factor loadings were significant at p < .001, with exception of those indicated in the table. VA = Verbal Analogies; RC = Reading Comprehension; WK = Word Knowledge; GS = General Science; AR = Arithmetic Reasoning; MK = Math Knowledge; IC = Instrument Comprehension; AI = Aviation Information; BC = Block Counting; TR = Table Reading.
p > 05. **p < .01.
Table 6 demonstrates several indices derived from bifactor models to evaluate the quality of unit-weighted total and subscale score composites or factor score estimates. Across the three bifactor models, the hierarchical g-factor accounted for between 45% and 48% of the common variance. The specific factors accounted for between 2% (quantitative factor) and 95% (acquired knowledge) of common variance of the items in each specific factor. Model-based estimate of internal reliability ranged from .83 to .85 for the general factor (ω) and from .55 to .84 for specific factors (ωs). The Omega Hierarchical coefficient was estimated to be between .58 and .62 for the general factor (ωh) and between .02 (quantitative factor) and .67 (acquired knowledge factor) for specific factors (ωhs). The ωh and ωhs coefficients suggest that general factor and some specific factors may be sufficient for interpretation. Many specific constructs, particularly quantitative factor, seemed inappropriate for interpretation beyond the general factor as little variance exists beyond the general factor (Reise, 2012).
Dimensionality Assessment of AFOQT Using Several Indices From Bifactor Models.
Note. AFOQT = Air Force Officer Qualifying Test; OmegaH = Hierarchical Omega; ECV = explained common variance; H = H index of construct replicability; FD = Factor Determinacy; PUC = percentage of uncontaminated correlations.
The H index of construct replicability indicates that the g-factor can be replicable (>.80), but none of the specific factors seemed replicable latent constructs (Hancock & Mueller, 2000). The Factor Determinacy Index indicates that factor score estimates of all constructs, including g, should not be used (FD < .90; Gorsuch, 1983, p. 260). The joint evaluation of ECV, PUC, and ωh suggests the presence of some multidimensionality, but it might not be strong enough to disqualify the interpretation of the AFOQT Form T as unidimensional (Reise et al., 2013; Rodriguez et al., 2016). The results demonstrate a manifestation of a higher order factor, with signs of multidimensionality.
Study 3: Criterion-Related Validity
This study attempted to assess the relations of the four ability latent factors (verbal, quantitative, aviation-related acquired knowledge, and spatial-perceptual ability) with pilot performance criteria after its plausibility had been established by the meta-analytic EFA and the CFA on three pilots’ data sets. Given the support previously established for the bifactor model as a favored expression of the four-factor solution of AFOQT data, the intended SEM predictive models relating latent ability factors with latent/observed performance criteria were also specified on this basis. This should also lead to the ongoing debate on whether the general ability or specific abilities that has a greater effect in the prediction of job performance. Although bifactor models have seen a growing interest (Reise, 2012; Rodriguez et al., 2016) and have successfully been utilized in predictive studies (e.g., Gustafsson & Balke, 1993), their heavy use has been in factor-analytic studies as a measurement model for test batteries or other instruments. Their application as part of SEM models, particularly as a framework for ability–performance relationships, is not as frequent as might be expected. A unique feature of this model is that it allows for disentangling the general effect from specific effects. It partitions domain-general from domain-specific variance and allows the unique effects of specific abilities to emerge and manifest (e.g., Zhang et al., 2020). Accordingly, the latent factors underlying AFOQT Form T were linked to pilot performance measures using this modeling approach.
The same three data sets utilized in the CFA study were used in this study because they contained correlations associating AFOQT subtests with pilot performance measures. The primary studies from which these data were extracted, however, did not attempt an assessment of ability–performance relationships at latent construct level, as the goal pursued in the current study. Hence, the criterion-related validity design and the research method applied are unique to this study. The measures used to indicate pilot performance varied across the three studies, which can be useful to draw a more reliable conclusion about the role of abilities in predicting various types of pilot performances. The decision whether to model performance scores as latent or observed in the SEM models was informed by the number of performance measures available in the data and the phases of training in which they were collected (primary, advanced). There was a preference for modeling performance as latent whenever a sufficient number of indicators were available.
Duke and Ree’s (1996) study reported three performance measures; each represented a distinct type of performance. These were as follows: (a) the exceedance of average flying hours in the primary phase of training, (b) the exceedance of average flying hours in the advanced phase of training, and (c) the students’ class rank upon completion of the flight training course. The first two measures were selected for this study to construct latent performance for flying skills. Olea and Ree’s (1994) study reported six performance measures, two of which were used in this study to represent a latent pilot performance. The two measures were check flight ratings during the primary phase of training and check flight ratings during the advanced phase of training. Arth et al.’s (1990) technical report offered only one performance measure representing the final outcome of training (pass or fail), and hence, it was modeled as observed. The scales of measurement for the measures in Samples 1 and 2 were interval and that for Sample 3 was nominal (dichotomized pass/fail).
A similar analytic procedure was planned for all three data sets to investigate the predictive relations between ability factors and performance measures. The ability factors were linked to performance measures (latent variable or observed variable) in a bifactor SEM model. Fitting a bifactor model can be a suitable approach for the current investigations because the relations of both general ability and specific abilities with performance measures are simultaneously estimated. Every ability factor, including the general ability factor, will have a path (i.e., regression) coefficient estimating its effect on performance criteria, controlling for other abilities in the model. Thus, the unique contribution of each ability can be estimated. The loadings on the factors with two indicators were constrained to be equal. All SEM models were estimated using ML, and the traditional model fit indices (CFI, RMSEA, and SRMR) were used for models’ assessment. Regarding the interpretation of the resulting effect sizes, the normative correlation guidelines suggested by Gignac and Szodorai (2016) were considered: .10, .20, and .30 indicate relatively small, typical, and relatively large, respectively.
Results
Figure 3A and B displays the structural models defined to estimate the predictive relations between cognitive ability factors underpinning AFOQT Form T and pilot performance. In these figures, the effects of the four ability factors (or their residuals) on each latent (Samples 1 and 2) or observed (Sample 3) performance measure were separately estimated. Similar to multiple regression, each effect represents the degree to which an ability factor uniquely predicts performance criterion after accounting for the effects of the other ability factors in the model (and after partialling out performance variance that is domain-general).

Bifactor SEM Model on (A) Latent Pilot Performance and (B) Observed Pilot Performance.
All three ability–performance models fit the data well, with indices within the acceptable ranges, χ2(42) = 357.57, p < .001, CFI = .90, RMSEA = .08, SRMR = .05 for Sample 1; χ2(42) = 419.50, p < .001, CFI = .94, RMSEA = .07, SRMR = .05 for Sample 2; and χ2(61) = 158.24, p < .001, CFI = .94, RMSEA = .07, SRMR = .05 for Sample 3. Table 7 displays the standardized path coefficients for the relations. For Sample 1, aviation-related acquired knowledge (β = .43), quantitative ability (β = .25), and spatial-perceptual ability (β = .18) were the three positive and significant predictors in the model (p < .001). A differing results about the relative roles of abilities in flying performance were found in Sample 2, with aviation knowledge (β = .41) and the general factor (β = .27) being the two positive and significant predictors in the model (p < .001). For Sample 3, the aviation acquired knowledge (β = .42) was the only factor contributing substantially in the prediction of pilot performance. Verbal ability did not contribute positively in any of the three predictive models; in fact, its unique effect was consistently negative. Hence, three different patterns of predictive relations emerged across the three samples; however, all supported a critical role for the aviation-related acquired knowledge factor as a powerful predictor for future pilot performance. The predictive role of general ability, quantitative ability, and spatial-perceptual ability for pilot performance appeared inconsistent across data, whereas that of verbal ability was consistently unimportant.
Prediction of Pilot Performance by General Ability and Specific Abilities via Bifactor Models.
Note. Bold font indicates significant structural regression coefficient.
p > .05. *p < .05. **p < .01. ***p < .001.
Sensitivity Analysis
To assess the robustness of current findings, we performed EFAs and CFAs similar to that applied in Studies 1 and 2 using two more AFOQT data sets from the general population of USAF officer applicants. This is meant to be a sensitivity analysis to detect any variation in results due to the source of data (AFOQT Form T/other AFOQT forms) or sample characteristics (non-pilots’ applicants/pilots’ applicants). Also, such an analysis provides further evidence on the extent that the results revealed by our study are replicable across samples. Correlation matrices were reproduced from Carretta et al.’s (2016) technical report, which was conducted to assess the psychometric properties of AFOQT Form T at initial item-, test-, factor-, and composite-level. Details of all results can be found in the Online Supplemental Material B.
EFA results on AFOQT Form T data of officer applicants were fairly similar to those found for AFOQT Form T meta-analytic data of pilots. The MAP suggested one factor, eigenvalue greater than 1 and scree plots suggested two factors, and HPA and SMT suggested four factors. A four-factor solution for AFOQT data remained more suggestive solutions according to factor retention criteria. There is one notable difference, however, concerning the physical science (PS) subtest, the modified version of GS subtest in Form T. Our result from the AFOQT meta-analytic data for pilot samples corresponded well to the theorized model tested by Carretta et al. (2016), where the GS/PS was found to load on verbal ability factor. Differently, the sensitivity analysis on samples from general officer applicants indicated that the PS subtest loaded on quantitative ability factor (≅0.52), rather than verbal ability factor, and also had salient cross-loading on aviation knowledge factor (≅0.36). Hence, more investigations are required to determine whether this variation in results represents a true change in the latent construct underpinning PS subtest or possibly a difference related to sample characteristics (i.e., general officer applicants, pilot applicants). The mathematical requirement of physics problems could be one reason made the PS to lean toward quantitative factor rather than verbal factor. Similarly, the fact that the PS subtest does not contribute to any of the AFOQT Form T composites may have impacted applicants’ motivation and performance leading to construct contamination.
Forcing a four-factor solution in a second-order EFA with the Schmid–Leiman orthogonalization procedure resulted in holding the same four factors suggested by EFA (verbal, quantitative, acquired knowledge, and spatial-perceptual ability), with salient factor loading. Albeit not salient (<.30), the PS and IC subtests had sizable cross-loading on acquired knowledge and spatial-perceptual factors, respectively (≅0.23). Fit statistics of the four CFA models for both data sets were also comparable with those derived from AFOQT Form T meta-analytic data for pilots. Across the two data sets, bifactor models described the AFOQT better than any other models, while single-factor models proven to be the least fitting models. Correlated-factor and higher order factor models had fairly comparable fit statistics. Overall, internal construct validity of AFOQT data seems reasonably stable, regardless of specific officer applicants’ subpopulation from which the data are collected or specific AFOQT forms are administered. A four-factor model expressed in a bifactor model represents AFOQT factor structure appropriately, whether the data collected from USAF officers’ general applicants or from a specific subpopulation such as pilots’ applicants.
General Discussion
The present investigation attempted to validate the AFOQT Form T that has been in operational use in USAF since 2015. Via the application of meta-analysis, EFA, CFA, and SEM procedures on prior AFOQT investigations on pilots’ samples, the goals intended in this research were achieved to a good extent. Studies reported in the current investigation provided a useful assessment for the AFOQT factor structures. In addition to the exploratory approach employed in Study 1 (i.e., EFA and second-order EFA) to evaluate the factor structure of meta-analytic AFOQT data, the subsequent CFA study provided further evidence for the most plausible factor structure of AFOQT subtest data based on three pilots’ sample. This investigation can be informative for the emerging line in psychometric research that compares several competing models of intelligence models to understand the best representation of human cognitive structure (e.g., Beaujean, 2015; Cucina & Byle, 2017; Morgan et al., 2015; Murray & Johnson, 2013).
The factor structure assessed in Study 1 via EFA and second-order EFA yielded a useful understanding of the ability factors underlying AFOQT Form T. What makes this examination of particular interest is that the EFA procedure was applied to a pooled correlation matrix of pilot cognitive data, spanning decades of AFOQT research. The evidence resulting from this study thus may be stronger and more reliable than that derived from any single set of data. We concluded that the four-factor model is the factor structure most representing AFOQT Form T data. This conclusion was supported by model assessment criteria (i.e., HPA, SMT, and scree plots) and was readily interpretable as four distinct ability factors. Due to the existence of only one EFA investigation for an earlier AFOQT 16-subtest version (i.e., Skinner & Ree, 1987), the present results can be beneficial and add new findings for the AFOQT accumulated research. The evaluation of the four constructs underlying AFOQT Form T indicates that they resemble the AFOQT’s theorized five-factor structure, but with one common factor grouping the two single subtests remained in Form T for spatial ability and perceptual speed factors (BC and TR, respectively). The four-factor structure (i.e., verbal ability, quantitative ability, aviation-related acquired knowledge, and spatial-perceptual ability) can be an alternate population model to that of the five-factor structure.
The second-order EFA with the Schmid–Leiman orthogonalization procedure was a useful complementary analysis for the first-order EFA to evaluate the appropriateness of making uni- or multidimensional interpretations of the AFOQT Form T data. The results indicated that the lower order four factors are unique constructs in the AFOQT structure and represent distinct factors, even with the presence of general factor. However, the small amount of variance accounted for by lower order factors and low Omega Hierarchical estimates suggests that the interpretation of AFOQT Form T is more demonstrable at the level of general factor than at the level of specific construct. Hence, caution should be heeded when interpreting the lower order factors on the AFOQT Form T. This conclusion was further supported by the results derived from bifactor indices (e.g., ECV, OmegaH [Hierarchical Omega], PUC). Nonetheless, according to a recent simulation study (Dueber, 2020), low OmegaHS estimates can be sufficient for separately interpreting subdomains when they are evaluated jointly with OmegaS. When OmegaS are high, such as those found in the current study, OmegaHS become less relevant to determine whether a subscore has added value. Similarly, Dueber (2020) showed that the ECVSS value (i.e., the proportion of common variance of the items in specific factor S explained by specific factor S) is best interpreted in conjunction with subscore’s reliability (OmegaS). For example, the ECVSS value as low as .30 can be sufficient to warrant interpretation of a subscore when the subscore’s reliability (OmegaS) is .80. It has been revealed that the importance of ECVSS diminishes as OmegaS increases further.
The CFA study was a useful investigation for AFOQT factor structure to verify whether the suggested four-factor latent factor is a viable structure for pilot applicants’ data. The study has also a theoretical implication for intelligence theory because the tested CFA models were specified similarly to those frequently discussed as plausible structure underlying cognitive data. Given that the AFOQT is a selection test battery, the findings may be of a special interest and supplement those concluded from standard intelligence test batteries. By assessing three pilots’ samples, a single-factor model (Spearman, 1904) appeared to be the weakest among the four competing models (i.e., single-factor, correlated-factor, higher order factor, bifactor), and fit indices did not support its viability. Evidence from other test batteries also shows that this model rarely yields acceptable goodness-of-fit statistics for cognitive data (Schneider & Newman, 2015), which indicates that the variance in cognitive tests cannot be explained by only one single ability. By comparison, a four correlated-factor model showed an acceptable fit for AFOQT data. Studies consistently show that this model possesses strengths over single-factor model and tend to fit the cognitive data adequately (e.g., Morgan et al., 2015), despite that Gignac and Kretzschmar (2017) caution against taking the results from correlated-factor models, on their own, as evidence for separability, or uniqueness, of each of the hypothesized specific factors. Instead, they recommended to further analyze the data using a higher order model.
When taking into account the general ability factor, our results demonstrated that a higher order factor model can also be a likely representation for AFOQT data and was fairly comparable with that of correlated-factor model. The appropriateness of hierarchical models to describe human cognitive ability structure has been determined in many studies (e.g., Cucina & Howardson, 2017; Gustafsson, 2001; Watkins, 2010). Reeve and Blacksmith (2009) found that almost 57.6% of the articles applying CFA methods have used a higher order modeling strategy. Contrarily, recent findings suggest that bifactor structure represents cognitive data more fittingly than higher order structure. The less constrained, bifactor models were found to be more conceivable than that of strict, higher order factor models (e.g., Cucina & Byle, 2017). Even when samples are obtained from a true higher order structure data, Morgan et al.’s (2015) simulation study showed that a bifactor model tends to produce a better fit than higher order model. Concurrently, when there is unmodeled complexity (e.g., correlated residuals and cross-loadings), bifactor models are still found to be more plausible than higher order models (Murray & Johnson, 2013). Our study for AFOQT Form T, same as previous AFOQT studies of Forms Q and S (Carretta & Ree, 1996; Drasgow et al., 2010), provides support for this general conclusion about the eminence of bifactor models for presenting cognitive data over other CFA models.
Study 3 assessed the predictive validity of AFOQT latent factors for pilot performance using bifactor predictive SEM models. Isolating domain-general from domain-specific effects showed negligible domain-specific effects for verbal ability (all three samples), quantitative ability (Samples 2 and 3), and spatial-perceptual ability (Samples 2 and 3) on pilot training performance criteria. Aviation acquired knowledge was the factor most contributing to performance in the three samples with an average effect of .42. Surprisingly, g-factor related to pilot training performance criteria significantly in Sample 2 only (β = .24), but trivially in Samples 1 and 3. The few significant relations between specific abilities and performance criteria after removing the general-domain effect indicate that the contribution of some specific cognitive abilities in flight performance can be more than only g. This contributes to the ongoing discussion about the relative role of general and specific abilities for the prediction of job performance (J. E. Hunter, 1986; Kell & Lang, 2017, 2018; Lang et al., 2010; Ree et al., 1994). Overall, in the presence of general ability, one to three AFOQT Form T latent abilities demonstrate meaningful predictive validity for pilot performance. Aviation-related acquired knowledge, in particular, related uniquely with all performance criteria across the three samples, suggesting strong predictive value for pilot performance that can be even stronger than that of general ability. A similar result was also manifested with different pilot performance criteria (ALMamari & Traynor, 2021). The construct of job/technical knowledge in aviation is consistently found to be an influential predictor for subsequent pilot performance (e.g., ALMamari & Traynor, 2019, 2020; Ree et al., 1995; Zierke, 2014). It should be noted, however, that the current findings were based on uncorrected correlation coefficients, which may preclude uncovering the true effect of general ability construct and its components on performance criteria (e.g., A. R. Jensen, 1998; A. R. Jensen & Weng, 1994).
Limitations and Future Research
Through the application of several analytic procedures, the current study achieved, to a good extent, the intended validation goals. The construct validity and criterion-related validity of AFOQT Form T for pilots’ samples were adequately established. However, there are some limitations to be noted when interpreting the results of the present work. First, all studies included in the meta-analysis investigation of Study 1 were conducted by the same group of researchers at a somewhat close period of time, which may raise a question about the dependency of the synthesized correlation coefficients. One critical assumption in meta-analysis is that the effect sizes being synthesized are independent, and failure to meet this assumption could cause misleading or even wrong results (Cheung, 2019). Hence, this observation about studies included in the meta-analysis needs to be borne in mind when interpreting the results of Study 1. Second, due to the practical organizational goals of AFOQT use, its coverage does not correspond to any particular intelligence framework. Future research may apply methods that allow joint assessment of constructs and hence facilitate a better understanding of those constructs that may be matching/mismatching the current measurement scope of AFOQT. The cross-battery assessment approach (e.g., Flanagan & McGrew, 1997) is one promising procedure for this research direction. Joint CFA between AFOQT and other test batteries can be useful for validating the current construct underlying AFOQT and for suggesting further improvement.
Third, in this study, we used uncorrected data, not the corrected data, in the attempt to provide results that are more correlative to the expected outcomes from cognitive testing. Due to the range restriction typically associating job selection and academic admission, the range of observed scores is limited and the correlation coefficients among variables are restricted in range, particularly for those central in the selection process. The results obtained, therefore, are likely underestimates of the true relations between cognitive abilities (and pilot performance). This limitation is well known in the literature and is often treated with some forms of correction (e.g., Raju & Brand, 2003; Sackett & Yang, 2000). Correction for range restriction (and measurement unreliability) usually yields stronger relationships between variables (e.g., ability vs. performance). As seen in Study 3, the significant predictors of pilot performance yielded generally small structural regression coefficients. Similarly, the effects of general ability on pilot performance criteria were found to be of small magnitude or even insignificant, which diverge from the overwhelming results emphasizing the dominant role of the general ability factor in workplaces based, most often, on corrected data (J. E. Hunter, 1986; Ree et al., 1994; Schmidt & Hunter, 2004). The range restriction in Samples 1 and 3 may conceivably be responsible for the unexpected low effect of g on pilot performance criteria. Hence, utilizing corrected data from these samples may show different conclusion about the relative contribution of cognitive abilities on pilot performance and may establish a stronger true effect for the g-factor.
Fourth, as a result of shortening the AFOQT in Form T, three of the four ability factors were represented by only two subtest indicators, which may limit capturing the construct under investigation. Three or more indicators are often recommended (Kenny et al., 1998; Marsh et al., 1998). However, other researchers found that two indicators or even one may be sufficient (Hayduk & Littvay, 2012). A large sample size, like those used in this study, can also help to compensate for a few indicators per factor (Koran, 2020). Furthermore, the several factor-analytic assessments of AFOQT data attempted in this study, via EFA and CFA, provided evidence supporting the construct validity of the ability factors, even with the relatively fewer indicators measuring their constructs. Aviation-related acquired knowledge, the strongest predictor of pilot performance revealed in this study, has only two indicators (IC and AI); therefore, the influence of the number of indicators seems minor. Fifth, in spite of several data sets utilized in Study 3, the modeling techniques were based primarily on a bifactor model, which has some inherent limitations (Reise, 2012; Reise et al., 2010). For example, the factors in a bifactor model, either general factor or grouping factors, are restricted to be uncorrelated. Also, each indicator in a bifactor model is allowed to load onto the general factor and to only one grouping factor. Due to the known intercorrelation between cognitive data, these assumptions may seem unrealistic, where group factors are conceptually related, or an indicator can mark more than one construct. It would be useful to attempt different approaches with other analytic procedures to give the results further credibility. Example of such approaches that have recently gained popularity includes relative weight analysis (J. W. Johnson, 2000), dominance analysis (Azen & Budescu, 2003), and non-g residuals with higher order models (Coyle, 2018). Replicating the results in Study 3 using some of these methods can give further confidence in the results.
Sixth, AFOQT Form T has one notable difference from earlier AFOQT versions at the subtest level. The GS test that was introduced and maintained throughout several AFOQT forms has been revised to the PS test in Form T, with content more focused on physical topics. Such changes may impact the narrow construct underlying this subtest, which may, in turn, influence its loading on the latent factors. Due to the reliance in this study on data from earlier AFOQT forms, it was not possible to examine the nature of this change and what impact it may have on the loading of these subtests. Although Carretta et al.’s (2016) examination of AFOQT Form T on the general population of USAF officer applicants indicated reasonable fit statistics for models that specified this subtest under verbal ability and aviation knowledge factors, our EFA and CFA within the context of sensitivity analysis on the same data indicated that the PS subtest is more saturated with quantitative ability than verbal ability.
Last, data of the current study were obtained from previous AFOQT forms, not from Form T, which may raise concern about potential variation in selectees’ performance related to the specific AFOQT forms administered, such as subtest administration order or test length. While the ideal strategy in the psychometric evaluation is to examine the most up-to-date version of a test, not an earlier version, we found our research strategy to provide reliable evidence and to satisfyies the validation requirement reasonably well. Researchers of this work are independent with no contact whatsoever with AFOQT development team, which hinders their access to newly collected row data for Form T. In addition, AFOQT data for pilot applicants, the targeted population in this research, may even be more difficult to obtain, which make the reliance on already published data sets a sensible option. With respect to the possible influence of subtest administration order and test length on selectees’ performance, studies examined these factors have shown inconsistent findings. For example, the assumption that longer tests lead to more cognitive fatigue, which has a negative effect on test takers’ performance, was contrasted by results from several studies (e.g., Ackerman & Kanfer, 2009; J. L. Jensen et al., 2013; Ryan et al., 2010; Tulsky & Zhu, 2000). Other factors seem more influential in the performance than the length of the test, such as differences in student personalities, interests, motivation, learning styles, or test-taking strategies (e.g., Ackerman & Kanfer, 2009). Similarly, evidences indicated that item order (or test order) has no or only slight effects on the test takers’ performance (e.g., Cole et al., 2017; Hohensinn et al., 2011; Weitensfelder, 2017; Tulsky & Zhu, 2000; Yousfi & Böhme, 2012). Some studies concluded that the test-order effect is not a potential threat to the internal validity of a test battery (Tulsky & Zhu, 2000; Yousfi & Böhme, 2012). Therefore, performance of test takers is not necessarily influenced by reducing the number of subtests in AFOQT Form T or changing their administration order.
Conclusion
The primary aim of the present study is to address the lack of research evidence on the factor structure and criterion-related validity of AFOQT Form T for pilot samples. We have done so by reanalyzing several data sets of pilot applicants using meta-analysis, EFA, CFA, and SEM procedures. The results are informative and have useful implications. The AFOQT Form T factor structure is best explained by four latent cognitive abilities marking verbal ability, quantitative ability, spatial-perceptual ability, and specific-domain acquired knowledge. A bifactor CFA model incorporating these four specific abilities, along with the general factor of ability, is found to be the best-fitting model for AFOQT Form T data. The predictive validity of AFOQT’s latent cognitive abilities for pilot performance was demonstrated suggesting that aviation-related acquired knowledge is the strongest predictor and verbal ability is the weakest predictor, while g-factor, quantitative ability, and spatial-perceptual ability have inconsistent prediction trend across samples. This study contributes to the AFOQT and ability test batteries literature and provides finding of practical and theoretical implications. In sum, the validity evidences collected for AFOQT Form T in this study support its conceptual groundwork and endorse its appropriateness for use in pilot selection.
Supplemental Material
sj-docx-1-asm-10.1177_10731911211058691 – Supplemental material for Factor Structure and Criterion-Related Validity of the Air Force Officer Qualifying Test Form T for Pilot Applicants
Supplemental material, sj-docx-1-asm-10.1177_10731911211058691 for Factor Structure and Criterion-Related Validity of the Air Force Officer Qualifying Test Form T for Pilot Applicants by Khalid ALMamari in Assessment
Footnotes
Appendix
Acknowledgements
The author acknowledges the insightful comments and help of Dr. Anne Traynor and Dr. Yukiko Maeda who commented on an earlier version of this work. Thanks go out as well to the editor and three anonymous reviewers who helped to improve the manuscript and made many useful suggestions.
Author’s Note
Part of this manuscript is adapted from the author’s academic thesis.
Khalid ALMamari is now affiliated to The Military Technological College, Muscat, P.O. Box 262, P.C. 111, Oman.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The author was supported during this work by the Royal Air Force of Oman and the Ministry of Higher Education in Oman.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
