Abstract
This study compared the structure of cognitive ability (specifically, verbal/crystallized [Gc] and visual-spatial ability [Gv]), as measured in the Wechsler Intelligence Scale for Children, in youth with manic symptoms with a nationally representative group of similarly aged youth. Multigroup confirmatory factor analysis found the majority of the estimated parameters were invariant between the groups, although there was a difference in the intercepts for the Similarities subtest and difference in unique variances for the Picture Completion, Comprehension, and Picture Arrangement subtests. Thus, although there are many neurological changes associated with manic symptoms, the structure of verbal/crystallized and visual-spatial abilities appear relatively robust and invariant. As Gc and Gv are the primary domains on all the Wechsler Intelligence Scales for Children, results also indicate that clinical interpretation of the Wechsler scales may be appropriate to measure cognitive performance in youths with manic symptoms.
Keywords
Bipolar disorder among children and youth has become an important focus of clinical research. The onset of this disorder in children and adolescents causes not only disruptions in social (Goldstein, Miklowitz, & Mullen, 2006) and family functioning (Lewinsohn, Klein, & Seeley, 2000) but also a tremendous challenge in terms of academic functioning (Doyle et al., 2005; Henin et al., 2007; Lagace, Kutcher, & Robertson, 2003).
As the onset of a disorder on the bipolar spectrum appears to be accompanied by many neurobiological and neurochemical changes (Berns & Nemeroff, 2003), an obvious follow-up question is if the structure of cognitive abilities also changes. Although there is a reasonable body of literature regarding cognitive abilities in adults with bipolar disorder (for meta-analyses, see Arts, Jabben, Krabbendam, & van Os, 2008; Robinson et al., 2006), less research has been available regarding cognitive profiles of youth with this disorder until recently, and the research that has been published is often equivocal (e.g., Dickstein et al., 2004; Doyle et al., 2005; Glahn et al., 2005; for a review and meta-analysis, see Joseph, Frazier, Youngstrom, & Soares, 2008). There are some differences that appear more pronounced during acute mood states, and others that appear to persist during remission and which may be trait markers of illness. However, meta-analyses have identified some emerging patterns. More specifically, youth with this disorder may manifest the greatest impairments in the areas of verbal memory, attention, and working memory (Cohen’s [1988] ds = .60-.77), in addition to other deficits with small-to-medium effect sizes, for example, executive functioning, visual memory, and verbal fluency (Joseph et al., 2008). Some of the deficits in executive function appear to be specific to bipolar disorder and not attributable to attention-deficit/hyperactivity disorder or other comorbid conditions (Walshaw, Alloy, & Sabb, 2010).
Importantly, despite indications that deficits may exist for some neuropsychological domains, many research groups have failed to find differences in Full Scale Intelligence Quotient (FSIQ) scores between youth with a bipolar diagnosis and their nondiagnosed peers (Dickstein et al., 2004; Dickstein et al., 2005; Lagace et al., 2003; McClure et al., 2005; Rucklidge, 2006). When studies have found differences, the youth with bipolar disorder usually scored within the average range, although lower than the comparison subjects (Doyle et al., 2005; Meyer et al., 2004; Olvera, Semrud-Clikeman, Pliszka, & O’Donnell, 2005; Voelbel, Bates, Buckman, Pandina, & Hendren, 2006). In one instance, the healthy comparison group had a mean FSIQ one standard deviation above the mean, clearly creating an artifact of lower FSIQ in the bipolar group (Voelbel et al., 2006). However, in another study, inpatient youth with bipolar disorder did have strikingly lower FSIQ scores and Performance IQ scores when compared with inpatient youth with attention-deficit/hyperactivity disorder, oppositional defiant disorder, or conduct disorder (McCarthy et al., 2004), with scores from most bipolar cases falling in the below average to borderline range.
Despite these relatively recent investigations into the varying cognitive abilities in youths with bipolar disorder or different degrees of manic symptoms, there are still many unanswered questions regarding the underlying structure of cognitive abilities. Specifically, nowhere in the literature does anyone address the question of whether manic symptoms is associated with change the overall structure of cognitive abilities in youths, or if there are just mean differences, with the overall structure remaining the same. The distinction becomes important when, say, comparing youths with learning disabilities to their typically developing peers. Studies that have compared students with learning disabilities versus the general school-age population often show that students with learning disabilities tend to have lower scores on measures of cognitive ability (Fiorello et al., 2007). However, there are not just mean differences between the groups, as the few studies that have compared the factor structure between these groups have also shown differences, indicating that the ways in which knowledge is processed, remembered, and used is different in those with learning disabilities than in the general population (Pomplun & Omar, 2001; see also Hale & Fiorello, 2004). As yet, there is no study that has examined if the same holds true in youth with manic symptoms. It is plausible that mania might degrade performance in some areas, but be neutral or even enhance performance in other areas. These contrasting influences would change the relationship of the scales measuring each domain—both to each other and to any latent factors.
Comparing the structure of cognitive abilities across groups can be done through a process called multigroup factor analysis (MGCFA; Widaman & Reise, 1997). The goal of multigroup factor analysis is to see if the parameter estimates for the factor-analytic model (e.g., factor loadings, variances) are the same across multiple groups. Scholars investigate factorial invariance under two scenarios: investigating measurement invariance (MI) or factor-analytic structural equation invariance (FASEI; Bontempo & Hofer, 2007; Vandenberg & Lance, 2000). Although the methods used in both scenarios can be equivalent, the reason for choosing a particular method is usually quite different. Researchers conduct MI studies in test development as part of the test validity studies, and the major purpose is to investigate possible bias in the instrument. Consequently, they typically conduct the analyses at the item level and usually involve only one instrument. FASEI studies, on the other hand, are usually interested in more substantive theoretical issues or, perhaps, developing better models for the construct of interest. Researchers usually conduct FASEI studies with aggregated scores (e.g., subtests, index scores) and can involve one or multiple instruments.
The current study compares the cognitive ability in youth with manic symptoms with their similarly aged peers without such symptoms. It is a FASEI study, as the purpose is to assess whether the structure of cognitive ability is the same in the two groups of youths and not to assess for bias in a specific measurement instrument.
Method
Sample
Two hundred forty-eight youth were admitted to a university hospital children’s psychiatric inpatient facility for the first time between September 1987 and December 1993. Irrespective of admission symptoms, a school psychologist administered a Wechsler Intelligence Scale for Children (WISC) to all admitted youth unless he/she had had recent testing in school or the psychologist was not available on admission. A psychologist interviewed the youth and their parents with the Yale Version of the Schedule for Affective Disorders and Schizophrenia for Children (Orvaschel, Puig-Antich, Chambers, Tabrizi, & Johnson, 1982) within 2 weeks of admission. The psychologist doing the interviews was blind to preadmission data. Eighty-six of these youth either reported having experienced manic symptoms and/or had parents who reported the youth had manic symptoms sufficient to warrant diagnosis of a manic episode, and thus Bipolar I disorder. All patients met with a board-certified child psychiatrist, who confirmed the diagnoses. The clinical interview and observation on the inpatient unit ruled out instances where mania might have been substance induced or because of a general medical condition. Prior medications were discontinued at admission so the staff could understand the degree to which visible psychopathology was because of the child’s condition versus a response to unsuccessful treatment when they were hospitalized. Psychoeducational testing was done at that time. Although the Diagnostic and Statistical Manual of Mental Disorders (DSM) has gone through several revisions through the period during which data were collected, it is important to note that the symptoms and definitions of a manic episode and Bipolar I disorder remained consistent across DSM-III (American Psychiatric Association, 1980), DSM-III-R (American Psychiatric Association, 1987), and DSM-IV (American Psychiatric Association, 1994). Of those 86 with mania, 81 had subtest scores available on a WISC. The mean age for the group was 8.62 years (SD = 1.95), and 89% of them were male. Eighty-three percent were Caucasian, 12% were African American, and the other 5% were Hispanic. The patients meet criteria for the following diagnoses, although they were not necessarily given a diagnosis: attention deficit/hyperactivity disorder, 81.50%; conduct disorder, 58.8%; oppositional defiant disorder, 86.4%; schizophrenia, 4.9%; depression (any type), 43.2%; anxiety (any type), 16.0%; pervasive developmental disorder, 3.7%; learning disorder (any type), 13.6%; language disorder, 19.8%; IQ ≤ 70, 1 4.9%.
For more information about the clinical sample see other work by Carlson and colleagues (Carlson & Kelly, 1998; Carlson & Youngstrom, 2003).
Instrument
The 81 youth with manic symptoms (YMS) had scores on the following WISC subtests: Information, Similarities, Arithmetic, Vocabulary, Comprehension, Digit Span, Picture Completion, Picture Arrangement, Block Design, Object Assembly, and Coding. The model of cognitive ability used in this study for the WISC subtests follows the results from Phelps, McGrew, Knopik, and Ford (2005), whose model, in turn, follows from Carroll’s (1993) work, the most comprehensive work on the structure of cognitive abilities to date (Jensen, 2004). In Phelps et al.’s (2005) model, the Vocabulary, Information, Similarities, and Comprehension subtests are measures of verbal/crystallized intelligence (Gc). Block Design, Picture Completion, Picture Arrangement, and Object Assembly subtests are measures of visual-spatial ability (Gv), although it is important to note that these tests are also time limited. The Digit Span, Arithmetic, and Coding subtests each measure separate abilities (short-term memory, quantitative knowledge, and processing speed, respectively). Because the protocol measure did not gather any additional data on these other abilities, and the use of single indicator latent variables is tenuous (Brown, 2006), we did not use scores on these subtests for the current analysis. A pictorial representation of this study’s model of cognitive ability is given in Figure 1.

The Verbal/Crystallized (Gc)–Visual-Spatial (Gv) factor model for the Wechsler Intelligence Scale for Children–Third Edition
The WISC is particularly apt for assessing invariance of Gc and Gv abilities. All WISC editions published to date have multiple subtests that measure these constructs. As a result, these factors are well specified and there is a large accumulated body of research with these constructs measured via the WISC subtests in both clinical and nonclinical samples. It is the third most commonly administered test by practicing clinical psychologists (behind only the Wechsler Adult Intelligence Scale and the Minnesota Multiphasic Personality Inventory—both of which are used primarily with adults; Camara, Nathan, & Puente, 2000), and it is the most commonly used test in the assessment of adolescents (Archer & Newsom, 2000). Moreover, the Wechsler family of cognitive ability scales (which include the WISC, the Wechsler Adult Intelligence Scale, and the Wechsler Preschool and Primary Scale of Intelligence) are the most commonly taught tests of cognitive ability in both clinical and school psychology programs (Alfonso, Oakland, LaRocca, & Spanakos, 2000). More than 60% of these programs see the WISC as “essential” to training (Belter & Piotrowski, 2001), and roughly 90% of all training programs include the WISC in their core assessment training (Childs & Eyde, 2002). Thus, using the WISC when measuring Gc and Gv abilities maximizes generalizability and clinical relevance of the study.
WISC editions
In 1991, the hospital where the YMS were admitted moved from administering the WISC–Revised (WISC-R; Wechsler, 1974) to administering the WISC–Third Edition (WISC-III; Wechsler, 1991a) following its publication, consistent with ethical guidelines for practice. Of the 81 YMS, 38 have scores on the WISC-R and 43 have scores on the WISC-III. The covariance-and-means matrices of the two groups were compared for equivalence for the eight WISC variables included in Figure 1 and were found not to be different (Sörbom, 1974), so they were combined into one YMS group for all other analyses (see Table 1). Table 2 provides descriptive statistics for the WISC subtests in the combined YMS group.
Fit Statistics for Youth With Manic Symptoms on the Wechsler Intelligence Scale for Children Combined Data
Note. CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual.
Correlation Matrix (With Standard Deviations on Diagonal) and Means for Wechsler Intelligence Scale for Children Subtests in Youth With Manic Symptoms
For the reference group, the WISC-III standardization sample was chosen, as opposed to another WISC edition (e.g., WISC-R), because its norms were developed close in time to when the YMS data were collected, thus lessening concerns about Flynn effect artifacts (Beaujean, Sheng, & Qiu, 2010; Flynn, 2007). Moreover, using Cohen’s (1988) monikers, the correlations between the WISC-R and WISC-III subtests are large for Gc (Information, 80; Similarities, .74; Vocabulary, .77; Comprehension, .67) and Gv (Picture Completion, .57; Picture Arrangement, .42; Object Assembly, .58; Wechsler, 1991b). The WISC-III norming data (correlations, means, and standard deviations) are divided by age groups, so the current analysis used the 9-year old sample (Wechsler, 1991b) as it is closest in age to the average age of the YMS data. There were 200 children assessed at this age, with 50% of them being male. The norming samples were selected to represent the population of children in the United States in the late 1980s and thus are diverse with respect to geographic location, race, and socioeconomic status (Wechsler, 1991b).
Testing for Invariance
Testing for invariance is a multistep procedure, testing more restrictive models at each step (Widaman & Reise, 1997). When testing for factorial invariance, there are multiple ways to examine if the data are consistent with the increasingly restrictive models, but Byrne and Stewart (2006) suggest two sets of criteria. The first (“traditional perspective”) examines the change in chi-square values (Δχ2) across nested models. If, as the models grow more restrictive, the Δχ2 values do not significantly change, this is evidence that more restrictive model fits data as well as the less restrictive model; thus, the more restrictive model should be favored over the less restrictive one.
The use of Δχ2 values has been criticized because of their sensitivity to sample size (Cheung & Rensvold, 2002). Recently, Cheung and Rensvold (2002) and Meade, Johnson, and Braddy (2008) provided evidence that some alternative fit indices were not prone to this problem. Specifically, both studies found that the comparative fit index (CFI; Bentler, 1990) and McDonald’s (1989) Noncentrality Index (Mc) were robust across a variety of sample sizes. Moreover, for both statistics, the changes in these values from one model nested in another (ΔCFI, ΔMc, respectively) were relatively uncorrelated with the change in values from another nested-model comparison. This is important because it means that the test of invariance between one set of models is not dependent on a test of invariance for another set of models. Thus, Byrne and Stewart’s (2006) second line of criteria (“practical perspective”) recommends that invariance can be based on two criteria: (a) the multigroup factor model exhibits an adequate fit to the data and (b) the change in values for fit indices (e.g., ΔCFI, ΔMc) is negligible.
Fit indices used for this study
Based on Byrne and Stewart’s (2006) recommendation, this study used two sets of fit indices: one to assess overall model fit and the other to assess change in model fit between two models. As Hu and Bentler (1999) recommend, we used multiple fit indices for both. For overall model fit, we included the root mean square error of approximation (RMSEA; Browne & Cudeck, 1993), CFI, Mc, and the standardized root mean square residual (SRMR; Bentler, 1995). These indices were chosen as they represent a variety of fit criteria and they tend to perform well in evaluating different models (Marsh, Hau, & Grayson, 2005). In addition, we also inspected each model’s χ2 value and its associated p value (Barrett, 2007). For both overall model fit as well as change in model fit, we looked for patterns in the fit statistics and judged acceptance/rejection of the specific model based on the majority of the indices.
For this study’s criteria of overall model-data fit, we used the following: (a) χ2 p values greater than .02 (to correct the traditional cutoff of .05 for multiple comparisons but at the same time not be overly restrictive on power; see also Yu, 2002), (b) RMSEA values less than .075 (halfway between .050 and .100; Chen, Curran, Bollen, Kirby, & Paxton, 2008), (c) SRMR values less than .080 (Hu & Bentler, 1999; Sivo, Xitao, Witta, & Willse, 2006), (d) CFI values greater than .960 (Yu, 2002), and (e) Mc values greater than .900 (Marsh, Hau, & Wen, 2004; Sivo et al., 2006). 2
To test the change in fit between nested models, we used the ΔCFI, ΔMc, and the Δ χ2’s associated p value. As with assessing overall fit, we used p values greater than .02 as evidence of invariance for the new constrictions between the two models. Cheung and Rensvold (2002) suggested .01 as the threshold of ΔCFI and .02 as the threshold for ΔMc. Meade et al. (2008), however, suggested more restrictive values of .002 for ΔCFI and .007 for ΔMc (based on having 2 factors and 8 indicators). 3 As this issue is not yet resolved, we considered both values for the ΔCFI and ΔMc, with values less than Meade et al.’s (2008) criteria indicating stronger evidence of invariance than values only meeting Cheung and Rensvold’s (2002) criteria.
Data inspection
When using latent variable models, or at least estimating their parameters by maximum likelihood, a key assumption is that the data come from a multivariate normal distribution (Kline, 2004). Using the methods outlined in DeCarlo (1997), we investigated the data for univariate and multivariate normality. No variable departed from univariate normality using D’Agostino and Pearson’s (1973) normality test, and the variable set exhibited multivariate normality using Small’s (1980) test. Consequently, we estimated all models in MPlus (version 5.1) using the maximum likelihood estimator.
Results
Invariance Analysis
Before testing for invariance, Meade et al. (2008) recommend first assessing factor model fit for each group separately. The results are in the first two rows in Table 3 (Models 1a and 1b). All the fit indices for the Gc-Gv model were within the limits specified in the Methods section for both samples, except the RMSEA for YMS data. Consequently, there appeared to be enough evidence to support assessing for invariance using the two-factor model.
Fit Statistics for Invariance Assessment
Note. YMS = youth with manic symptoms; CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; Mc = McDonald’s Noncentrality Index; Δ = change in value between competing models.
Similarities intercepts free.
Picture Completion, Comprehension, and Picture Arrangement unique variances are free.
In the next step, we assessed for configural invariance, which examines whether the factor model is the same across both groups. Given the results from the individual group analysis, it was not surprising that the Gc-Gv model fit the (combined) data well, as shown by all the models’ overall model fit values (i.e., χ2, RMSEA, SRMR, and Mc) meeting the prespecified criteria (see Table 3, Model 2).
The third step involved assessing for metric invariance, which we did by examining if the factor loadings were the same across both groups. The results (Table 3, Model 3) indicate that the overall model fit the data relatively well, as the fit statistics’ values all met the cutoff criteria. For the change-in-model fit statistics, the p value for the Δχ2 was greater than the cutoff and the ΔCFI and ΔMc both meet the Cheung and Rensvold (2002) criteria but not the Meade et al. (2008) criteria. Consequently, while some indices were not within limits specified in the Methods section, the majority of them were within the limits; thus, there appeared to be enough evidence to continue the invariance assessment.
In the next model, we constrained all the indicator variables’ intercepts (i.e., the scales’ origins) across groups. The results are in Table 3, under Model 4. The fit indices for the overall model satisfy the a priori criteria, but all three change-in-model fit values indicated this model did not fit the data as well as the previous (i.e., metric invariance) model. We examined the groups’ subtests’ mean scores and found the largest discrepancy for the Similarities tasks, so allowed it to vary between groups. We then tested this revised model (Model 4a) and found both the overall fit and change-in-fit statistics indicated this model fit the data relatively well, although the ΔCFI and ΔMc values exceeded the Meade et al. (2008) criteria.
The next step involved examining if any of the subtests’ unique (residual) variances were invariant across the two groups. The results (Model 5) of almost all the overall and change-in-fit indices (the exceptions were the SRMR and ΔMc using Cheung and Rensvold’s, 2002, criteria) indicated this model poorly fit the data. We then examined each subtests’ variability and found that Picture Completion, Comprehension, and Picture Arrangement had the largest between-groups differences (i.e., all >2 standard score units). Subsequently, we allowed them to vary between groups, and almost all the resulting fit indices from this ameliorated Model 5, Model 5a, indicated this model fit the data relatively well (the exception was the ΔCFI value using the Meade et al., 2008, criterion).
For the last step we assessed whether the factor variances, covariances, and means were invariant between the two groups. These steps were carried out separately (Models 6, 7, and 8, respectively, in Table 3), but converge in showing that there was no between-group difference in these values as both the overall model fit statistics and the change-in-fit statistics are all within the predefined criteria. The parameter estimates from the final model (Model 8) are in Table 4.
Estimates of Intercepts, Loadings, and Unique Variances of the Partial Invariance Model
Note. YMS = youth with manic symptoms; WISC-III = Wechsler Intelligence Scale for Children–Third Edition. Parameter estimates are shown for Model 8 (see Table 3), where all factor variances are 1, all factor means are 0, and the Gv-Gc covariance is .76. The estimated values for the YMS and 9-year-old norming group from the WISC-III groups are shown separately if invariance constraints are not imposed.
Discussion
The purpose of this study was to compare the structure of Gc and Gv abilities of youth with manic symptoms (YMS) and similar-aged peers from the general population. Using the four Gv and four Gc subtests on the WISC-III (Wechsler, 1991a), a group of YMS were compared with a similarly aged group from the norming sample of the WISC-III (9-year-old sample). The comparison was conducted by examining invariance in the mean and covariance structure between the two groups via a multiple-group confirmatory factor analysis.
The results indicated that the Gv-Gc factor structure (see Figure 1) and factor loadings are invariant across both groups. Moreover, for all subtests except Similarities, the intercepts were the same between both groups. Alternatively stated, all mean differences in the subtest scores (except for the Similarities) are accounted for by the Gv and Gc factor means (Bontempo & Hofer, 2007).
When we examined the unique variances, five of the eight subtests were equal across groups; the three that were not were Picture Completion, Comprehension, and Picture Arrangement. For Picture Arrangement the YMS group exhibited more unique variance (10.36 vs. 8.32), but the norming group showed more for the other two subtests (2.51 vs. 4.83 and 2.90 vs. 5.64 for Comprehension and Picture Completion, respectively). Last, we assessed invariance at the factor level and found, when the noninvariant intercept and unique variances were unconstrained, the factor variances, covariance, and means were equivalent across both groups.
Byrne, Shavelson, and Muthén (1989) first labeled the situation where most, but not all, of the parameter estimates across groups are the same as partial measurement invariance. However, as we conducted the current study under a factor-analytic structural equation invariance framework instead of measurement invariance framework (Bontempo & Hofer, 2007), we interpret the findings from a different perspective than Byrne et al. (1989). First, while the Gc and Gv constructs are the same between the two groups, their ability to predict subtest scores somewhat differs. More specifically, for a given factor score, the predicted subtest score is the same in both the YMS and the norming groups for all the WISC-III subtests except Similarities. For this subtest, a given score on the Gv factor will produce a higher predicted score for the YMS group than the norming group. As this difference in predicted Similarities scores will be constant across all Gv factor scores, this situation is akin to an ANCOVA with equal slopes but different intercepts (see Figure 2A). Second, because there were unique variance differences between the two groups on the Picture Completion, Comprehension, and Picture Arrangement subtests, there will be differences in the accuracy of the Gv and Gc factors to predict these subtests’ scores. For example, because the norming group has a larger unique variance for Picture Completion than the YMS group, we would expected that for a given Gc factor score the predicted Picture Completion score would be more accurate for the YMS group than the norming group. This situation is akin to heterogeneity of error variance in an ANCOVA model, where for a given value of the predictor variable there is more variability in the outcome variable in one group than another group. We depict this situation graphically in Figure 2B.

Graphical depiction of measurement invariance
The results from this current study indicate that despite the many neurobiological and neurochemical changes that co-occur in people with a bipolar diagnosis (Chang et al., 2005; DelBello, Zimmerman, Mills, Getz, & Strakowski, 2004; Goodwin & Jamison, 2007), the general structure of their overall cognitive abilities, at least in their youth, does not change. This finding is reassuring from a clinical perspective, because it indicates that one of the most widely used tests in clinical assessment (i.e., the WISC) appears to remain internally valid when working with youths potentially affected by bipolar disorder. Given the importance of recommending and implementing appropriate educational accommodations for youth with bipolar disorder (Fristad, Verducci, Walters, & Young, 2009; Papolos & Papolos, 2002; Pavuluri, O’Connor, Harral, Moss, & Sweeney, 2006), it is reassuring that a widely used measure of cognitive ability appears to maintain largely invariant measurement properties, even when used with youth experiencing severe enough disturbance to warrant psychiatric hospitalization.
Limitations
There are three major limitations in the current study. The first is that the YMS sample is small (n = 81). Although this sample was representative of the population of youth with mania who are admitted to the hospital where the sample was drawn, we should be cautious about generalization to all youth with mania, especially as the demographics for this sample (i.e., race, sex, age) might not match the general YMS population. Even though the comparison group had 200 people, the size of the overall combined sample (n = 281) could be a reason why we did not find significant departures from invariance at some levels (e.g., loadings, factor variances; Meade et al., 2008). Larger and more representative clinical samples would enable us to make statements that are more definitive. The second major limitation is the use of older versions of the WISC as the measure of cognitive ability. The Wechsler intelligence instruments are among the most commonly used instruments in applied psychology, but they only measure a few areas of cognitive ability, with Gv and Gc being the two major areas, although the fourth edition of the WISC does have fewer subtests measuring Gv than its predecessors (Kaufman, Flanagan, Alfonso, & Mascolo, 2006). Future studies should look to use either a larger battery of cognitive ability batteries or should use a battery designed to measure different areas of cognitive ability, such as the Woodcock-Johnson–Third Edition (Woodcock, McGrew, & Mather, 2001) or Differential Ability Scales–Second Edition (Elliot, 2007). Although medications had been discontinued on admission and for the baseline time during which testing was done, data were not available about the medications taken by each youth prior to admission. Many of the medications used to manage bipolar disorder are known to influence performance on neurocognitive tests. Because the sample is older and the children young, the medication exposure of participants was lower than it would be in a similar cohort today. Medication effects on short- and long-term cognitive performance remain an important topic of inquiry.
Overall, practitioners and consumers of testing reports—including educators, youths, and their families—can take some comfort in the fact that the major factors underlying the WISC appear to have fundamentally equivalent factor validity in youth with mania as compared with a nationally representative sample of youths. Because the current findings are from an inpatient sample, findings from other studies are likely to be similar (or to have even smaller differences) when factor validity is examined in less severely affected samples, such as outpatient or primary care settings. Present findings also enhance confidence in the validity of research finding using the WISC, as its measurement properties appeared robust across the constructs between clinical and community samples.
Footnotes
Declaration of Conflicting Interests
The authors declared no conflicts of interest with respect to the authorship and/or publication of this article.
Funding
The authors received no financial support for the research and/or authorship of this article.
