Abstract
This study examined key assumptions underlying the interpretation of one of the most widely used multidimensional nonverbal tests of intelligence, the Universal Nonverbal Intelligence Test–Second Edition (UNIT2). Specifically, we examined the dimensionality of the UNIT2 and the interpretive relevance of its factors. We also examined the invariance of constructs measured by the UNIT2 across age groups, gender, race, and ethnicity. Structural analyses were conducted using data from 1,802 individuals aged 5 to 21 years who participated in the norming of the UNIT2. Results indicate that the UNIT2 is primarily a measure of psychometric g. Tests of invariance indicate that the factors measured by the UNIT2 are calibrated differently across age, gender, and racial groups. The Memory, Quantitative, and Reasoning factors represent psychometric g quite well. However, there is insufficient unique, reliable variance for the interpretation of the index scores reflecting the Memory, Quantitative, and Reasoning factors. Based on the results of this study, we question whether the administration of multidimensional nonverbal tests of intelligence is worth the time and effort when unidimensional tests may provide the same information.
Of the multidimensional nonverbal tests of intelligence published during the past two decades, none has attracted as much attention in the literature as the Universal Nonverbal Intelligence Test (UNIT; Bracken & McCallum, 2001) and the Universal Nonverbal Intelligence Test-Second Edition (UNIT2; Bracken & McCallum, 2016). The UNIT and its successor, the UNIT2, are an individually administered, norm-referenced tests for individuals aged 5 to 21 years. They are the only multidimensional nonverbal tests of intelligence administered entirely nonverbally (i.e., via pantomime). Thus, receptive language is not required as part of the input demands for completing items, and expressive language is not required as part of the output demands, as the examinee responds by manipulating chips or blocks or by pointing rather than by providing verbal responses.
According to McCallum, Bracken, and Wasserman (2001), “a fundamental strength of the UNIT is its strong theoretical base, consistent with the models of Carroll (1993) and Jensen (1980), both of whom consider intelligence to be hierarchically structured and multifaceted” (p. 111). As shown in Table 1, the UNIT2 consists of six subtests intended to assess three cognitive abilities (reasoning, memory, and quantitative reasoning) with two organizational strategies (symbolic and nonsymbolic). As Table 1 shows, within each of the organizational strategies, the problem solution requires one of the three cognitive abilities. Thus, similar to the original UNIT (Bracken & McCallum, 2001), the UNIT2 continues to emphasize the assessment of global intelligence but expands the number of cognitive abilities measured to include quantitative reasoning in addition to memory and reasoning.
Conceptual Model for the Universal Nonverbal Intelligence Test–Second Edition.
McCallum et al. (2001) stated that sound interpretation of UNIT scores is based on key assumptions. These assumptions also apply to scores from its successor, the UNIT2. They asserted that the “the first assumption is that intelligence is best characterized by a hierarchical model that has general intellectual ability, or g, at the apex” (p. 102). Accordingly, the UNIT2 subtests and scales were designed to be good measures of general intelligence. Another assumption underlying interpretation of the UNIT2 is that the test has acceptable psychometric properties. McCallum et al. asserted that, in addition to adequate internal consistency, scores should have adequate unique variance to warrant interpretation. This is an important assumption given that Bracken and McCallum (2016) encourage ipsative (i.e., intraindividual) analysis of an individual’s profile of index and subtest scores.
How well does the UNIT2 meet these assumptions? Bracken and McCallum (2016) reported the results of confirmatory factor analysis (CFA) to support the UNIT2 scoring structure. Specifically, they presented results for a one-factor model (i.e., a model with a single, general factor), a two-factor model (i.e., a model with correlated Memory and Reasoning factors and a model with correlated Quantitative and Reasoning factors), and a three-factor model (i.e., a model with correlated Memory, Quantitative, and Reasoning factors). They reported model fit indices that generally supported all the models tested. However, they did not conduct model comparisons across these models to determine the best fitting model. Although Bracken and McCallum concluded that results from the series of models they tested support a multidimensional scoring structure with factors at different levels of generality and abstraction, they did not actually examine the fit of this model. Therefore, they did not report empirical evidence to support the claim that that the UNIT2 Full Scale Battery and Abbreviated Battery composites represent a higher order general factor (i.e., psychometric g, which commonly is referred to simply as g) and lower order Memory, Quantitative, and Reasoning factors.
Bracken and McCallum (2016) also did not examine plausible alternative models, including (a) models consistent with the design of the UNIT2 (e.g., addressing the organizational strategies measured) and (b) models closely aligned with the Cattell–Horn–Carroll (CHC) theory (e.g., Schneider & McGrew, 2012). Furthermore, they did not evaluate an alternative to the proposed higher-order model in testing bifactor models with a general factor and three group factors (i.e., Memory, Quantitative, and Reasoning). Although both higher order and bifactor models are hierarchical, they differ in important ways (Beaujean, 2015). Higher order models presume a higher-order structure wherein the general factor has indirect effects on all mental tasks and lower order group factors have direct effects on subsets of mental tasks. In other words, lower order factors mediate the effects of the general factor on all mental tasks (Gignac, 2008). In contrast to higher order models, bifactor models, first developed by Holzinger and Swineford (1937), include a general factor and group factors that differ with respect to the breadth of effects, as the general factor has direct effects on all mental tasks and group factors have direct effects on subsets of mental tasks that share features (e.g., similar stimuli and similar response demands). In the bifactor model, the general factor and all group factors are orthogonal to each other.
As the scoring structure of the UNIT2 has not been thoroughly examined, it is impossible to determine if there is adequate correspondence between the proposed theoretical domains and the empirical domain as operationalized by the subtests administered to examinees. In other words, additional research is needed to determine if the proposed constructs are adequately defined and provide meaningful representations of targeted constructs that are distinguishable from conceptually similar constructs (e.g., are the Memory, Quantitative, and Reasoning factors distinguishable from the general factor as well as from each other?). Furthermore, when marshalling evidence to support score interpretations, it is imperative to test plausible alternative models, as it is possible that an alternative model may demonstrate greater correspondence between the theoretical and empirical domains than the proposed model does (Cronbach, 1990). Rejection of plausible alternatives strengthens arguments for proposed interpretations of test scores.
When marshalling evidence to support score interpretations, it also is essential to evaluate the interpretive relevance of test scores. A review of the test manual and the published literature indicate a need for additional evidence to establish the interpretive relevance of scores derived from the UNIT2. A score with interpretative relevance meets the following criteria:
provides a good representation of the construct targeted for measurement,
is distinct from conceptually similar constructs,
is likely to be replicable across data sets and methods, and
has adequate unique, reliable variance such that it is statistically distinguishable from test-derived scores reflecting conceptually similar constructs. (Benson, Beaujean, McGill, & Dombrowski, 2018, p. 3)
Purpose of the Present Study
This study was designed to reexamine evidence supporting the structural validity of the UNIT2. We addressed the following three research questions:
What is the latent structure of the UNIT2?
To what extent are the constructs measured by the UNIT2 invariant across age, gender, racial, and ethnic groups?
Do the index scaled scores have adequate unique variance to warrant interpretation in ipsative analyses?
To further examine its latent structure, we conducted both exploratory factor analyses (EFA) and CFA, as findings from these different analyses are known to differ in important ways and are held to be complementary when used to examine dimensionality (e.g., Carroll, 1995). Additional examination of dimensionality is warranted given that plausible alternative structural models for the UNIT2 have not been examined. Furthermore, our CFAs expand on results of those presented by Bracken and McCallum (2016) by: (a) examining higher order and bifactor models and (b) conducting model comparison to evaluate plausible alternative models that combine elements of distinct models (e.g., a model addressing reasoning, memory, and quantitative reasoning as well as the symbolic and nonsymbolic organizational strategies) or that are consistent with CHC theory (Schneider & McGrew, 2012). In addition, using the best fitting and most viable models determined when answering the prior research questions, we examined the invariance of the constructs measured by the UNIT2 across age, gender, racial, and ethnic groups. Last, we evaluated the interpretive relevance of proposed test scores using model-based indices to examine the dimensionality and replicability of factors as well as the properties of test-derived scores used to measure these factors (Rodriguez, Reise, & Haviland, 2016a).
Method
Participants and Norming Data
This study employed data from 1,802 individuals aged 5 to 21 years who participated in the norming of the UNIT2 (Bracken & McCallum, 2016). The UNIT2 norming sample was split relatively evenly across binary gender classifications (50.7% males and 49.3% females). Approximately 76% of participants were White, 15.1% were African American, 4.5% were Asian or Pacific Islanders, 1.6% were American Indian, Eskimo, or Aleut, and the remaining participants were from another category or did not specify. About 19.9% of participants were Hispanic.
Participants in the UNIT2 norming sample were recruited across 33 states and selected to reflect the population of the United States based the following characteristics: gender, race, ethnicity, parent education level, household income, and regional representation (ProQuest LLC, 2014). Examiners participating in norming evidenced appropriate certifications or advanced training and received additional training in administering the UNIT2. Their protocols were subjected to quality control procedures described by Bracken and McCallum (2016).
Instrument
The UNIT2 (Bracken & McCallum, 2016) is an individually administered intelligence test designed for ages 5 to 21 years. The UNIT2 includes six subtests (see Table 2). Subtest scores are reported as scaled scores (M = 10, SD = 3) and are the primary measures employed in this study. The UNIT2 yields three specific ability composite scores: Memory, Quantitative, and Reasoning. In addition, it yields an abbreviated global intelligence score (Abbreviated Battery) stemming from two subtests and three alternative multidimensional global intelligence scores, including (a) the Full-Scale Battery stemming from all six subtests, (b) the Standard Battery with Memory, and (c) the Standard Battery without Memory, which both stem from four subtests.
Descriptions of Universal Nonverbal Intelligence Test–Second Edition (UNIT2) Subtests.
Note. Subtest descriptions and UNIT2 classifications were obtained from Bracken and McCallum (2016).
All UNIT2 subtests have yielded coefficient alpha values of .84 or higher across 1-year age groups of the norming sample; average alpha values across age groups ranged from .89 (Spatial Memory) to .96 (Analogic Reasoning, Nonsymbolic Quantity, and Numerical Series). Test–retest reliability coefficients, obtained from a sample of 199 participants tested across an average of approximately 17 days and corrected for range restriction, ranged from .75 (Spatial Memory) to .94 (Cube Memory). Correlations between subtest scores and CFA results support the structure of the UNIT2 as measuring global intelligence and some specific abilities (as described in the Introduction). Criterion-related validity evidence in the form of correlations with other intelligence test scores, correlations with achievement test scores, and clinical-group comparisons also support the construct validity of the UNIT2 subtest, composite, and global intelligence scores (Bracken & McCallum, 2016).
Data Analysis
Exploratory Factor Analysis
The initial stage of data analysis involved EFA of the UNIT2 total sample using a maximum likelihood estimator. When multiple factors were extracted, a geomin rotation was employed. In contrast to classic mechanisms such as Promax that adjust factor loadings to improve interpretability without consideration of fit, Geomin is a modern mechanism that identifies a rotated solution with the best fit relative to a target function while minimizing discrepancies between factor loadings in the original and rotated solutions (Browne, 2001). We used multiple criteria for selecting the number of factors to extract (Gorsuch, 1983), including (a) eigenvalues > 1 (Kaiser, 1960), (b) scree plot coordinates (Cattell, 1966), and (c) observed values in parallel analysis > corresponding value for the 95th percentile random data eigenvalue (Glorfeld, 1995; Horn, 1965). In addition to extraction criteria, solution viability was evaluated by considering salient factor pattern coefficients (i.e., >.3) and the extent to which the factor pattern achieves simple structure (i.e., each factor defines a distinct cluster of observed variables).
Confirmatory Factor Analysis
The second stage of data analysis involved CFA of the UNIT2 total sample using a maximum likelihood estimator. The following alternative fit measures and criteria were used to determine acceptable model fit: the comparative fit index (CFI; >.95), the root mean square error of approximation (RMSEA; <.08), and the standardized root mean square residual (SRMR; <.05). As Bayesian methods have been shown to be useful for evaluating the relative merits of competing models (Kass & Raftery, 1995), and the Bayesian information criterion (BIC) provides a useful approximation to the log Bayes factor, the BIC was used for model comparisons. BIC values were transformed into BIC (Schwarz) weights using the formula developed by Wagenmakers and Farrell (2004). These weights reflect the probability that a candidate model is the best model among the set of candidate models examined.
Although the aforementioned strategies can detect multidimensionality, the presence of multidimensionality in a dataset does not mean that scores reflecting each dimension should be derived and interpreted. Rather, it is important to evaluate the extent to which each score reflects common variance reflecting a general factor or residual variance attributable to a group factor after controlling for the effect of a general factor (Reise, 2012). In other words, evidence is needed to show whether or not a given score can be interpreted as a reliable measure of a general or specific factor. To this end, we used model-based indices to examine the dimensionality and replicability of factors as well as the properties of test-derived scores used to measure these factors (Rodriguez et al., 2016a).
Invariance Testing
After identifying the best fitting model using the UNIT2 total sample, measurement invariance analysis was conducted using the following age-based groups: (a) 5 to 7 years, (b) 8 to 10 years, (c) 11 to 13 years, (d) 14 to 17 years, and (e) 18 to 21 years. Age invariance ensures that UNIT2 subtest scores are on the same measurement scale across age groups. Measurement invariance also was examined across gender groups (female and male), racial groups (Asian or Pacific Islanders, Black, and White), and ethnic groups (i.e., Hispanic or non-Hispanic). The extent to which measurement invariance exists is determined by testing the measurement model (i.e., specified relations between factors and UNIT2 subtest scores). Equivalence determinations are made by testing increasingly stringent levels of invariance.
First, we examined configural invariance by determining whether groups have the same number of factors and pattern of factor loadings. If configural invariance is supported, metric invariance can be tested by constraining factor loadings to be equal across groups. Invariant factor loadings suggest that a factor is calibrated in a similar way across groups. If metric invariance is supported, scalar invariance can be tested by constraining the subtests’ intercepts to be equal. Invariant intercepts indicate an absence of systematic group differences in additive biases as estimated by discrepancies between observed means and means that are expected based on factor loadings (Steinmetz, 2013). Differences in the CFI (ΔCFI) were used to test increasingly restrictive invariance models, with ΔCFI of .002 or more used as the criterion for identifying noninvariance (Meade, Johnson, & Braddy, 2008). Likelihood ratio tests (Δχ2) also were conducted, although this was not used as the primary criterion for evaluating model fit because the χ2 statistic is very sensitive to sample size.
Interpretive Relevance
Dimensionality was evaluated using three indices: explained common variance (ECV; Reise, Moore, & Haviland, 2010), percentage of uncontaminated correlations (PUC; Bonifay, Reise, Scheines, & Meijer, 2015), and average relative parameter bias (ARPB; Rodriguez, Reise, & Haviland, 2016b). ECV is viewed both as an index of general factor strength and as an index of closeness to unidimensionality. Higher values indicate general factor strength. As a frame of reference, ECV values of .80 and .70 indicate a low level of bias from ignoring group factors (i.e., less than 5% and 10%, respectively). PUC provides an index of the extent to which covariances among observed variables reflect a general factor and are uncontaminated by variance from group factors. PUC values of .80 or greater suggest unidimensionality (Reise, Scheines, Widaman, & Haviland, 2013). ARPB provides an index of the difference between factor loadings on a general factor in a unidimensional model and factor loadings on a general factor in a bifactor model. In this study, differences in factor loadings on a general factor reflect the extent to which group factors in the bifactor model account for variance in observed variables (i.e., subtests). ARPB values thus reflect bias that would result from not accounting for the effects of group factors by including them in the CFA model. ARPB values less than .10 to .15 indicate minimal bias when interpreting a dataset as unidimensional in nature (Muthén, Kaplan, & Hollis, 1987).
A factor determinacy index (FD; Beauducel, 2011) and a construct replicability index (H; Hancock & Mueller, 2001) were used to evaluate the meaningfulness and replicability of factors in our final model. Factor determinacy values provide evidence that test-derived scores reflect individual differences in performance and may be of value for subsequent analyses or applied assessment purposes. H evaluates how well scores represent constructs of interest and provides information regarding the extent to which factors believed to represent these constructs are likely to replicate across studies. H values greater than .80 are considered indicative of a well-defined factor (Hancock & Mueller, 2001).
We calculated omega coefficients to gain information regarding the properties of the UNIT2 specific ability composite scores (Memory, Quantitative, and Reasoning) and the global intelligence scores. Omega coefficients provide model-based reliability estimates for unit-weighted scores. We calculated coefficient omega total (Lucke, 2005), which accounts for all variance shared by indicators of a construct. When appropriate, we also calculated the related statistics omega hierarchical and omega hierarchical subscale, which reflect the reliability of a single common (general) factor and the reliability of group factors with the effects of the general factor removed, respectively (Reise, 2012; Rodriguez et al., 2016b).
Results
Exploratory Factor Analysis
EFA results support a single factor solution for the UNIT2. The eigenvalue for the first factor was 3.23. All other eigenvalues fell well below the minimum criterion for retention. Parallel analysis indicated that only the first eigenvalue is larger than its corresponding 95th percentile random data eigenvalue, and the corresponding scree plot also indicates that a single factor should be retained. Factor loadings and communality estimates for a one-factor solution are presented in Table 3. All the subtests demonstrated salient factor loadings, ranging from .49 to .76. Numerical Series, Analogic Reasoning, and Cube Design demonstrated high g loadings, Spatial Memory and Nonsymbolic Memory demonstrated medium g loadings, and Symbolic Memory demonstrated a low g loading (McGrew & Flanagan, 1998). Communality values ranged from 24% to 58% with an average of 45%.
Factor Loadings and Communalities From a One-Factor Exploratory Factor Analytic Solution for the Total Sample.
Confirmatory Factor Analysis
Model Comparisons
CFA results are presented in Table 4. Notably, RMSEA values were not indicative of acceptable fit for any of the models. RMSEA values are inflated as a result of the candidate models having few degrees of freedom (Kenny, Kaniskan, & McCoach, 2015). When degree of freedom increased, as when conducting the tests of invariance presented in the next section, RMSEA values fell within acceptable limits.
Alternative Confirmatory Factor Analysis Model Fit Statistics for the Total Sample.
Note. BIC = Bayesian information criterion; wi (BIC) = rounded Schwartz weights; CFI = comparative fit index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; Gf = fluid intelligence; Gv = visual processing; M = Memory; Q = Quantitative; R = Reasoning.
Model testing proceeded from a one-factor model to more complex models. Model 1 included a single first-order factor (i.e., g), with all six UNIT2 subtests specified as indicators, fit the data reasonably well (CFI = .96, RMSEA [90% CI] = [.08, .11], SRMR = .04). Model 2, which is consistent with CHC theory (Schneider & McGrew, 2012), included two correlated first-order factors: a fluid reasoning (Gf) factor (with Analogic Reasoning, Nonsymbolic Quantity, and Numerical Series as indicators) and a visual processing (Gv) factor (with Cube Design, Spatial Memory, and Symbolic Memory as indicators). Model comparisons provide strong evidence against the one factor model (Model 1), indicating that a two-factor model with correlated Gf and Gv factors (Model 2) provides a superior fit (ΔBIC = 6.577). In Model 2, the correlation between the Gf and Gv factors was .94, suggesting a higher order general factor should be specified as affecting lower order Gf and Gv factors. The resulting higher-order model, which included a second-order g factor and first-order Gf and Gv factors (Model 3), could not be distinguished from Model 2 based on overall fit; these models were found to have equivalent fit.
Several models were evaluated that are consistent with proposed structure of the UNIT2. Model 4 included three correlated first-order factors: Memory (with Spatial Memory and Symbolic Memory as indicators), Reasoning (with Analogic Reasoning and Cube Design as indicators), and Quantitative (with Nonsymbolic Quantity and Numerical Series as indicators). Model 4 was better fitting than Model 3 (ΔBIC = 4.148). Correlations between the Memory, Quantitative, and Reasoning factors ranged from .84 to .95. Variants of Model 4, which included a higher-order model specifying a second-order general factor and first-order Memory, Reasoning, and Quantitative factors (Model 5) and a bifactor model specifying a first-order general factor as well as first-order Memory, Quantitative, and Reasoning factors (Model 6), were found to have equivalent fit and could not be differentiated based on fit. The strong correlations between the first-order factors in Model 4 indicate common variance attributable to g, which supports a higher order or bifactor model. In addition to these models presented in Table 4, two alternative models were evaluated, but neither converged. Therefore, results for these models are not presented. These inadmissible models included (a) a bifactor model with a first-order general factor and first-order Gf and Gv factors (as in Model 2) and (b) five-factor model (consistent with UNIT2’s conceptual model, see Table 1) with a first-order Symbolic factor (with Analogic Reasoning, Symbolic Memory, and Numerical Series as indicators) and a first-order Nonsymbolic factor (with Cube Design, Nonsymbolic Quantity, Spatial Memory as indicators), in addition to the first-order Memory, Quantitative, and Reasoning factors specified in Models 5, 6, and 7.
Three CFA models (Models 4, 5, and 6), which present variations of the same model specifying three-specific ability factors (Memory, Quantitative, and Reasoning), were equally well-fitting. In such situations, researchers must turn to other results to determine which model best represents the use and interpretation of the test’s scores. In light of the EFA evidence, we argue that a bifactor model provides the best structural representation for the UNIT2. EFA results support the predominance of the general factor in accounting for variance. Model 4 ignores the general factor. In higher-order models such as Model 5, the general factor has only indirect effects on subtests, whereas in the bifactor model (Model 6), the general factor has direct effects on subtests. Thus, the bifactor model better represents the influence of the general factor at the primary level of interpretation of the UNIT2, the global intelligence score. A bifactor model also promotes decomposition of variance into common variance shared by all subtests and unique variance attributable to individual subtests and the specific ability composites to which they contribute.
Invariance Testing
We did not examine the invariance of Model 4, as this model excludes the general factor while EFA results indicate the predominance of the general factor. Model 5 tended to have negative variance when examined separately in age- and race-based groups, yielding inadmissible solutions for some groups. Inadmissible solutions for some groups preclude tests of invariance and raise questions about the generalizability of this model across groups. Invariance of the bifactor model (Model 6) was first tested across age groups and was admissible for all groups. Fit estimates for each of the five age groups are presented in Table 5. Model fit was generally acceptable across age groups. Notably, CFI (.934) and RMSEA (.153) values were not supportive of fit for the 18- to 21-year-old group. Inspection of modification indices revealed covariance between the Quantitative factor and the unique and error variance term for Symbolic Memory. Adding this covariance to the model (or, alternatively, adding a path from the Quantitative factor to Symbolic Memory) improved the CFI value to .980, the RMSEA value to .092, and the SRMR value to .033. However, given that this covariance was observed for other age groups, it was not included when testing invariance across age groups. Configural invariance was tested by simultaneously examining the fit of the bifactor model across all ages. Configural invariance was supported, as evidenced by a CFI value of .964 and an RMSEA [90% CI] of .047 [.039, .054]. As shown in Table 6, metric invariance was not supported, as evidenced by a ΔCFI of .007 when factor loadings were constrained to be equal across groups. Thus, results indicate that the constructs measured by the UNIT2 are calibrated differently across age groups.
Fit Statistics for the Bifactor Model by Age Group.
Note. BIC = Bayesian information criterion; CFI = comparative fit index; CI = confidence interval; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual.
Invariance Results for the Bifactor Model by Age, Race, Gender, and Ethnicity.
Note. CFI = comparative fit index.
Compared with previous model.
Invariance of the bifactor model also was examined across gender, race, and ethnicity. Model fit was acceptable across groups. However, as shown in Table 6, metric invariance was not supported across racial group or gender, as evidenced by ΔCFI of approximately .006 and .004, respectively. Conversely, metric invariance was supported across ethnic groups, as evidenced by a ΔCFI of .001. Scalar invariance was not supported across ethnic groups (Hispanic or non-Hispanic), as evidenced by a ΔCFI of .021 when intercepts were constrained to be equal across groups. In summary, the factors measured by the UNIT2, particularly g, appear to be calibrated differently across age, gender, and racial groups. Average standard deviations for g loadings tend to be higher across age-, gender-, and race-based groups than average standard deviations for subtest loadings on group factors—which is not surprising given that g loadings are considerably larger than the loadings for the Reasoning, Quantitative, and Memory factors.
Interpretive Relevance
Standardized loadings and squared multiple correlations (R2) for subtest scores in the bifactor model (Model 6) are presented in Figure 1. The rectangles in the center of the figure are UNIT2 subtest scores. The effects of the general factor on subtests can be examined by reviewing the values above the arrows on the left-hand side of the figure. These values are g loadings, as previously described; they are almost identical to those from the EFA. The effects of the three group factors on subtests can be examined by reviewing the values above the arrows on the right-hand side of the figure. In every case, they are lower than the g loadings. They ranged from .34 to .02. R2 values appear above in the rectangles below the names of their corresponding subtest. They are communality estimates like those reported in Table 3, but they represent the score variance accounted for across general and specific factors specified in the model (as communality estimates in Table 3 represented only general factor effects). R2 values ranged from 33% to 65%, with an average of 52%.

Bifactor model for the Universal Nonverbal Intelligence Test–Second Edition (UNIT2; Model 6) with standardized loadings and squared multiple correlations for observed variables (in parentheses).
Using values from the bifactor model presented in Model 6, ECV indicates that the general factor explains about 86% of the variance in subtest performance. The PUC index indicates that about 80% of the covariance across subtests exclusively reflects g variance. Moreover, the ARPB value of .02 suggests that treating the UNIT2 structure as unidimensional would result in only about 2% bias in g loadings. As evident in Table 7, the factor determinacy value of .91 supports the use of a factor score reflecting the general factor outside of a factor model, and the H value of. 84 suggests that the general factor is a well-defined factor in this dataset that would be likely to replicate across samples. The omega total value of .85 and omega hierarchical value of .81 support the interpretation of a unit-weighted composite score from all six subtests reflecting g. The omega hierarchical value of .81 indicates that, when averaged across ages, 81% of the variance in the UNIT2 Full Scale Battery score can be attributable to g. When considering subsets of subtests contributing to the two other global intelligence composite scores and the abbreviated global intelligence score, omega total values and omega hierarchical values were as follows: .83 and .79 for the Standard Battery without Memory, .78 and .73 for the Standard Battery with Memory, and .70 and .62 for the Abbreviated Battery. Thus, based on the omega total values, two of the four UNIT2 global intelligence composites, Standard Battery with Memory and Abbreviated Battery, failed to demonstrate sufficient levels of reliability for interpretation (.80 or higher; Kranzler & Floyd, 2013; C. R. Reynolds & Livingston, 2014).
Statistical Indices From the Bifactor Model Analysis.
Note. g = general intelligence, M = Memory, Q = Quantitative, R = Reasoning, FD = factor determinacy, H = construct replicability, Omega = model-based estimate of internal consistency for unit-weighted composite score, Omega H = percentage of variance attributable to g, Omega HS = percentage of unique, reliable variance for group factor that is independent of g.
In the same vein, omega total values for the Memory, Quantitative, or Reasoning factors were less than .80; thus, they failed to demonstrate values that indicate sufficient reliability. In addition, as indicated by omega hierarchical values, the general factor accounted for the lion’s share of variance in these factors. Omega hierarchical subscale values for these three factors were .14 or lower, which indicates that their unique variance that is independent of the general factor is insufficient for interpretation at the score level.
Discussion
The purpose of this study was to reexamine key assumptions underlying the interpretation of one of the most widely used multidimensional nonverbal tests of intelligence, the UNIT2. Specifically, we examined the structure of the UNIT2 and the invariance of the constructs measured by the UNIT2 across age, gender, racial, and ethnic groups. Additionally, we examined the interpretive relevance of factors measured by the UNIT2.
Results of this study suggest that, although the UNIT2 is indeed multidimensional, it is primarily a measure of global intelligence. Invariance testing suggests that the abilities measured by the UNIT2 are calibrated differently across age, gender, and racial groups. Metric invariance is not supported across age and gender groups. Although metric invariance is supported across Hispanic and non-Hispanic groups, scalar invariance is not. Noninvariance precludes comparisons of group means, as at least part of the difference between groups results from construct-irrelevant sources of variance.
It is unclear how the levels of noninvariance observed in this study impact interpretation at the level of individual students. Steinmetz (2013) simulated the impact of unequal factor loadings and intercepts on composite score differences between groups. Results indicated that unequal factor loadings had weak effects while unequal intercepts had strong effects on composite score differences (i.e., unequal intercepts produced spurious differences between groups when the true latent mean difference between groups was set to zero). The effects of unequal intercepts increased as the number of unequal intercepts increased and decreased as the number of indicators increased. Future research that examines how group differences in factor loadings and intercepts impacts observed test scores at the level of individual students is needed.
In the present study, the mean subtest intercept for students who identify as Hispanic is 9.41, whereas the mean subtest intercept for non-Hispanic participants is 10.13. Thus, subtests tend to be more difficult for Hispanic students due to construct-irrelevant sources of variance. As the UNIT2 Full Scale Battery consists of six subtests, given an average difference in subtest intercepts of .72 points, it is likely that the sum of scaled scores for Hispanic students is negatively biased by more than 4-scale score points. Likewise, the mean subtest intercept for Black students in 8.59, whereas the combined mean of White and Asian students is 10.23. As factor loadings were noninvariant across race-based groups the interpretation of intercepts is not straightforward. Nevertheless, given an average difference in subtest intercepts of 1.64 points, it is possible that the sum of scaled scores for Black students is negatively biased by nearly 10 points. Differences this large will produce large differences in observed IQ scores.
Based on the results of this study, we recommend focusing interpretation on global intelligence scores, representing psychometric g with high fidelity, rather than on the index or subtest scores. In addition, although the Memory, Quantitative, and Reasoning factors are good measures of g, there is insufficient unique, reliable variance to interpret them individually. The results of our study also lead us to question the alleged superiority of multidimensional nonverbal tests of intelligence over unidimensional tests. Unidimensional nonverbal tests measure g, whereas multidimensional tests such as the UNIT2 measure g and several broad cognitive abilities. Although the results of our study found that the UNIT2 is a good measure of g, once the variance related to g is taken into account, the memory, reasoning, and quantitative composites reflect insufficient unique variance to warrant interpretation. Given that unidimensional tests also tend to be good measures of g and may be more cost effective to administer, one important question is whether there is value added to the individual administration of the UNIT2. Further research on other multidimensional nonverbal tests of intelligence is needed to determine whether these findings are generalizable beyond the UNIT2.
Limitations
At least three limitations of this study should be noted. First, the goal of this study was to examine the internal structure of the UNIT2 through factor analytic methods. This test was designed to measure general intelligence and a small number of specific ability factors, and it is composed of only six subtests. Thus, these findings are not reflective of the true nature and structure of human cognitive abilities (Carroll, 1993; Schneider & McGrew, 2012) due to undersampling of the diversity of cognitive abilities. Second, as our sample was based on the large and nationally representative UNIT2 norming sample, it provides the most replicable estimates of the relations between and among the UNIT2 subtests and the psychometric properties of its scores. As these analyses were not conducted with samples of children and adolescents with educational conditions and clinical disorders, it is possible that our results may not generalize to these populations. Finally, the subtest g loadings and g saturation estimates for composite scores are not likely to be consistent across ability levels. With the possible exception of the subtests measuring Reasoning, g loadings and g saturation estimates are likely to be stronger in magnitude at lower ability levels and lower in magnitude for higher ability levels (M. R. Reynolds, 2013).
Implications for Practice
What should school psychologists do when assessing the cognitive ability of children who are not proficient in English, fully acculturated in mainstream American society, or both? Or when assessing the cognitive ability of children and youth who are deaf or have significant hearing impairments or language-related disabilities? The longstanding best practice recommendation has been to use a nonverbal test of intelligence. Results of our study support the interpretation and use of the UNIT2 as a measure of general intelligence. Two of the test’s global intelligence scores, the Full Scale Battery and the Standard Battery without Memory, met reliability standards (with omega total values .80 or higher; Kranzler & Floyd, 2013; C. R. Reynolds & Livingston, 2014) and demonstrated substantial g saturation (with g predicted to account for well more than three-quarters of their variance). Based on our analysis employing the entire UNIT2 norming sample, 81% of the variance in the Full Scale Battery and 79% of the variance in the Standard Battery without Memory can be attributable to psychometric g. These values are relatively comparable to those values reported for global intelligence scores from traditional, multidimensional intelligence tests for children and adolescents (M. R. Reynolds, Floyd, & Niileksela, 2013). In particular, M. R. Reynolds and Keith (2017) estimated that 80% of the variance in the Wechsler Intelligence Scale for Children–Fifth Edition (Wechsler, 2014) Full Scale IQ is attributable to g.
Our findings, however, do not support use of ipsative analyses of the profile of an individual’s test scores to identify cognitive strengths and weaknesses on the UNIT2. Thus, our results lead us to question whether the additional time and resources required for the administration of multidimensional nonverbal tests of intelligence is worth it when the same information might be obtained from the administration of a unidimensional test. Further research on the UNIT2 and other multidimensional nonverbal tests of intelligence is needed.
Footnotes
Authors’ Note
The authors thank PRO-ED, Inc. for allowing us to access the standardization data for the Universal Nonverbal Intelligence Test–Second Edition (UNIT2; Bracken & McCallum, 2016). Copyrights by PRO-ED, Inc.
Declaration of Conflicting Interests
The author(s) declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: We have cited a book co-written by the second and third authors of this article, which could reflect financial interest in publishing this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
