Abstract
The ability to reason with language is a highly valued cognitive capacity that correlates with IQ measures and is sensitive to damage in language areas. The Penn Verbal Reasoning Test (PVRT) is a 29-item computerized test for measuring abstract analogical reasoning abilities using language. The full test can take over half an hour to administer, which limits its applicability in large-scale studies. We previously described a procedure for abbreviating a clinical rating scale and a modified procedure for reducing tests with a large number of items. Here we describe the application of the modified method to reducing the number of items in the PVRT to a parsimonious subset of items that accurately predicts the total score. As in our previous reduction studies, a split sample is used for model fitting and validation, with cross-validation to verify results. We find that an 8-item scale predicts the total 29-item score well, achieving a correlation of .9145 for the reduced form for the model fitting sample and .8952 for the validation sample. The results indicate that a drastically abbreviated version, which cuts administration time by more than 70%, can be safely administered as a predictor of PVRT performance.
Verbal reasoning capacity is an important cognitive domain that is often assessed using analogy problems (e.g., Ekstrom, French, & Harman, 1979). As stated in a recent evaluation of a large longitudinal data set, “Analogical reasoning is a core cognitive skill that distinguishes humans from all other species and contributes to general fluid intelligence, creativity, and adaptive learning capacities” (Richland & Burchinal, 2013, p. 87). Clinically, deficits in verbal reasoning have been linked to damage in left temporoparietal cortex and are useful in predicting recovery (e.g., Willner et al., 1976). Therefore, tests of verbal analogical reasoning are often included in neuropsychological testing (e.g., Lezak, Howieson, Bigler, & Tranel, 2012).
With the advent of large-scale clinical and genomic studies and increased demand for comprehensive neurocognitive assessments to be used as endophenotypes (Green et al., 2004; R.C. Gur et al., 2011; Insel & Cuthbert, 2009; McClay et al., 2011; Reichenberg et al., 2010; Seidman et al., 2010), there has been increased need to abbreviate tests. Such efforts have been applied to several domains of cognition, resulting in abbreviated versions of specific tests (e.g., Calamia, Markon, Denburg, & Tranel, 2011; Winegarden, Yates, Moses, & Faustman, 1997) or abbreviated batteries that cover multiple domains (e.g., Cheng & Morgan, 2013; Harvey, Keefe, Patterson, Heaton, & Bowie, 2009; Hill et al., 2008; Hurford, Marder, Keefe, Reise, & Bilder, 2011; Mansbach & MacDougall, 2012; Ventura et al., 2010).
The Penn Verbal Reasoning Test (PVRT) is a multiple-choice test in which participants answer verbal analogy problems (R. C. Gur & Reivich, 1980). It has been administered in functional neuroimaging studies (R. C. Gur et al., 1982; R. C. Gur et al., 1983; R. C.Gur et al., 2000; R. C. Gur & Reivich, 1980; Roalf et al., 2013; Wharton et al., 2000), reliably producing left hemispheric temporoparietal activation, with performance linked to anatomical measures of these regions (R. C. Gur et al., 1999). The test has also been included in many large-scale genomic studies (Almasy et al., 2008; Calkins et al., 2010; R. C. Gur et al., 2012; R. E. Gur et al., 2007), demonstrating construct and criterion validity (R. C. Gur et al., 2001; R. C. Gur et al., 2010; R.C. Gur et al., 2011; R. C. Gur et al., 2012; Irani et al., 2012; Roalf et al., 2013) as well as heritability and genetic correlations (Almasy et al., 2008; Yokley et al., 2012). The items within the PVRT are multiple choice anologies, similar to “Kitten is to Cat as Puppy is to: a. Chicken; b. Cow; c. Dog; d. Milk.” The test can take more than half an hour to administer. It was originally composed of 30 items and was modified to a 29-item scale as 1 item was used for practice and never presented in tests. The total score on the test is the number of correct items. Notwithstanding Smith, McCarthy, and Anderson’s (2000) call for caution in the application of abbreviated tests, abbreviation is justified for a central, narrowly defined domain and considering the need for efficient neurocognitive testing to provide endophenotypes in large-scale genomic studies.
We previously described a procedure for abbreviating a clinical rating scale, the Quality of Life Scale, from 21 to 7 items (Bilker et al., 2003). This method was modified to accommodate large item sets and used for reducing a matrix reasoning test from 60 items to 9 (Bilker et al., 2012). Here we describe the application of this modified method to reducing the number of items in the PVRT.
Method
Data
The study sample comprised 319 participants. Table 1 presents a summary of the sample demographics. Age ranged from 15 to 77, with a mean of 34.1 (SD = 13.0) years. There were 172 male participants (54%), and the mean number of years of education was 14.4 (SD = 2.6). Of the participants, 109 were healthy volunteers, 100 were patients with schizophrenia, 55 were relatives of patients, and 55 were classified as other.
Demographics of the Study Sample.
We split the sample into two subsets: one to fit and the other to validate the models. The “model construction set” contained 160 participants and the “validation set” contained 159 participants. No demographic measures shown in Table 2 were significantly different between the model construction and validation sets.
Demographics of the Model Building and Validation Subsets of the Study Sample.
The method of Bilker et al. (2003) considered all possible subsets of test items, up to about one half of the total test items. For each combination, the predicted total score of the omitted items was estimated using linear regression. The sum of predicted scores for the omitted items and the observed scores of the items included in the model was then used as the predicted total for the particular model. The Pearson correlation and the intraclass correlation between the predicted and observed total scores were used as metrics for assessing the predictive ability of each subset of test items.
Their procedure required modification to allow application to the PVRT. First, the outcomes of interest are counts, and, as shown in Figure 1, the total number of incorrect items appears to follow a Poisson distribution. Second, the PVRT contains 29 items, and as we wish to reduce the number of items to at most 20, there are over 530 million possible subtests, and thus the computational burden of applying a brute force procedure as the above is not practical.

Histogram of the total number of incorrect items.
Bilker et al. (2012) sought to rectify this issue by modifying the previous approach. First, Poisson regression replaced linear regression to better fit the count data. Second, an adaptive algorithm was employed that dramatically reduced the number of models considered in the reduction process. This modified scale reduction approach as applied to the PVRT reduction is described below.
The process of selecting a subset of items is as follows: Let A be a particular subset of items and let AC be all items not in A. For each subset, a Poisson regression is run predicting the total number of incorrect items in AC from the scores of the items in A. As an example, suppose A contained Items 1, 4, 7, and 13. Thus, AC contains all items except 1, 4, 7, and 13 and the predictive Poisson model is then
where R is the number of incorrect items in AC; α is the intercept; Ij is equal to 1 if item j is incorrect and 0 otherwise; and βj is the coefficient for item j. The predicted scores of the items in AC are then computed via:
As the original and reduced forms of the test share items, the Pearson correlation between the total scores from both forms is spuriously high (Levy, 1967; Tellegan & Briggs, 1967). The spuriousness is due to the variance of the reduced form scores, which are contained in both forms. Thus we utilize the modified correlation coefficient proposed by Levy (1967), which adds a negative-valued penalty term to the Pearson correlation, taking into account the shared variation (Levy, 1967). For each subset A, the modified correlation coefficient (ρ) and the intraclass coefficient (ICC) between the predicted values of total correct items with the observed number of correct items are computed for both the model construction and validation sets.
To assess the effect of the particular split on the results, we employed the two bootstrap methods presented in Bilker et al. (2012). The first considers the drop in correlation between the split sets while the second explores the particular subset of items and their frequency of inclusion in the final subtest among different splits of the data. We set an a priori rule for determining the optimal model, adding predictors until the percentage increase in correlation fell below 1% for an additional predictor.
Results
Table 3 presents the items chosen, modified correlations, and intraclass correlation coefficients for the top models for each number of predictors for the construction and validation sets. All of the models were fit using all diagnostic groups collapsed. The correlation between predicted and observed total score is above .85 for as few as four items in the construction set and five items in the validation set. The model with eight predictors was chosen as the final model using the predetermined 1% criterion. Figure 2 presents the plot of observed total score versus the predicted total score from the final model.
Selected Items, Modified Correlations, and Intraclass Correlation Coefficients (ICCs) for the Top Models for Each Number of Predictors for the Construction and Validation Sets.
The mean, standard deviation, minimum, and maximum of the predicted scores are also presented. The form with eight items is chosen as the final reduced form.

Predicted total number of correct items versus the actual total number of correct items along with the 45 degree line, for the model construction data set.
The items in the final model are 4, 6, 7, 11, 15, 16, 23, and 28. The coefficients for each of these items are found in Table 4. Letting Ik be equal to 1 if the subject’s response for item k was incorrect and 0 if correct, predictions from our final model are computed via
where
Estimated Coefficients for the Poisson Model for the Top Model for Each Number of Predictors.
We compared the performance of the top models for 1 through 15 predictors between patients with schizophrenia and healthy controls. There were 50 patients and 55 controls in the model construction set and 50 patients and 54 controls in the validation set. The modified correlations and ICC of the observed totals with the predicted totals for patients and controls are presented in Tables 5 and 6, respectively. Correlations remained greater than .85 for as few as 3 items in the construction set and 7 items in the validation set. Considering only controls, predicted from the models based on all diagnostic groups collapsed, the correlations remained greater than .85 for as few as 9 items in the construction set and 12 in the validation set. Similar comparisons were made in the subgroups of first episode schizophrenia patients and family members: For first episode patients the correlations remained greater than .85 for as few as 4 items in the constructions set and 7 in the validation set, and for family members the correlations remained greater than .85 for as few as 7 and 10 items in the construction and validation sets, respectively. For the sake of brevity, the corresponding tables for the first episode patients and family members are not shown.
Selected Items, Modified Correlations, and Intraclass Correlation Coefficients (ICCs) for the top Models for Each Number of Predictors for the Construction and Validation Sets of Schizophrenic Patients Only.
Note. The mean, standard deviation, minimum, and maximum of the predicted scores are also presented.
Selected Items, Modified Correlations, and Intraclass Correlation Coefficients (ICCs) for the Top Models for Each Number of Predictors for the Construction and Validation Sets of Healthy Controls Only.
Note. The mean, standard deviation, minimum, and maximum of the predicted scores are also presented.
We considered the Pearson correlations among age, years of education, and gender with the full and reduced PVRT. We did not expect to find sex differences or age effects, but did expect significant correlations with education, and examined whether similar correlations were obtained with the full and reduced versions. The correlation between age and the full PVRT was –.12, while the correlation between age and the reduced PVRT was –.09. The corresponding correlations between education and the full and reduced PVRT were r(317) = 0.58, p < .01 and r(317) = .55, p < .01, respectively. To investigate whether there were gender differences between the full and reduced PVRT, we examined the mean total scores for each test stratified by gender. The mean total scores for the full PVRT were 18.2 for males and 18.3 for females. For the reduced PVRT, the mean total scores were 18.5 and 18.5, respectively. Thus, no differences in the magnitude of correlations with age, years of education, and sex were apparent between the full and reduced tests.
Since the range of the total score is bounded from below and above, floor and ceiling effects of the abbreviated test can be present. To assess the presence of these effects, we estimated the correlations within tertiles of the full PVRT based on the observed values from the full sample. For the PVRT, the tertiles are 0 to 16, 17 to 23, and 24 to 29. The modified correlations between the full and reduced PVRT for the first, second, and third tertiles were .80, .51, and .37, respectively. These suggest a ceiling effect, particularly in the third quartile, as its correlation is low.
Using a bootstrap approach, we investigated the variability in the difference in correlation in the model construction and validation sets when using our final eight-item reduced test. Using 10,000 bootstraps we calculated the drop in correlation between the model construction and validation sets. The mean drop in correlation was r = .0247 (SD = .0131, 95% CI = .0244, .0249) and was the same drop after the Levy correction to both correlations. Thus, for the reduced PVRT, the drop in correlation was consistent across the random bootstrap data sets.
The above approach assessed our final reduced set of items only. We also assessed the reduced test along with the choice of split used to form the model construction and validation sets. We randomly split the original data 200 times (no new data were generated), and for each split we performed the following: fit the predictive model, choose the best 8 items, and estimate the correlations for the model construction and validation data sets. The mean correlation was r = .9191 (SD = .0037, 95% CI = .9186, .9196) for the model construction sets, and for the validation sets it was r = .8760 (SD = .0122, 95% CI = .8743, .8777). The high correlations and tight confidence intervals indicate that the particular split used did not influence the results.
As mentioned above, for each of the 200 random splits, an eight-item reduced test was chosen. Each reduced test may vary from our reduced test shown above. The percentage of times an item was chosen in the top eight is displayed in Table 7. Because the correlation remains high and varies little among the random splits, yet the set of eight varies, we conclude that our eight-item reduced PVRT is one among possibly numerous valid choices achieving a high correlation.
The percentage of Times Each Item Was Selected in the Top 8 Across 200 Random Splits of the Data.
Note. The composition of the eight-item form varies across different splits.
Discussion
The scale reduction approach of Bilker et al. (2012) was successfully applied to the PVRT. The procedure enabled reducing the test substantially from 29 items to 8 items. This amounts to a greater than 70% reduction in administration time, which greatly increases the feasibility of the test’s use in large-scale studies. The modified correlation between the abbreviated and long forms was close to unity, r(158) = .9145. Validation approaches were used to further assess the quality of the prediction, and the abbreviated form performed exceedingly well. The abbreviated form also performed well when considering first episode schizophrenia patients and healthy adults subgroups. Last, there were no apparent differences in magnitude of sex differences or correlations with age or years of education between the full and reduced tests.
The methodology utilized has limitations including computational burden and the propensity to exhibit ceiling effects. Therefore, the abbreviated scale would not be suitable to detect differences among individuals performing at the high range, which is not the primary use of the PVRT. Furthermore, the reduced set of eight items is not the only valid set of reduced items, and we did not investigate the merit of the particular set relative to the other valid sets of items. However, our analyses demonstrate that the eight items chosen represent a subset that predicts the outcome of the full set of items with very good accuracy. Importantly, the abbreviated test shares many of the limitations of abbreviated tests as enumerated by Smith et al. (2000). These include lack of the ability to reliably assess factors within the large scale and the lower boundary of validity. However, our approach and results have avoided the pitfalls of failing to demonstrate comparable validity for the short and long form and the construct we assessed, verbal reasoning, is narrowly defined by the PVRT and has been unidimensional (e.g., R. C. Gur et al., 2010). Nonetheless, use of the short form should be limited to studies that do not seek to measure factors within the domain of verbal reasoning. Another limitation of abbreviated tests is that differences between the predicted and actual scores can be large even when substantial correlations have been established, and this would limit their ability to categorize impairment as accurately as the full form (Calamia et al., 2011). A final limitation of the approach is that item selection is based entirely on numerical calculations, without any consideration of theoretical relevance or interpretability. Methods that incorporate confirmatory factor analysis and item response theory (Hays, Morales, & Reise, 2000) could achieve similar efficiency while being more theoretically informative.
These limitations notwithstanding, evaluation of verbal reasoning abilities is an important part of neuropsychological assessment, and verbal analogies have traditionally served as stimuli for such tests (Ekstrom et al., 1979). The verbal reasoning test applied here has well-established psychometric properties and taps a well-validated factor (Ekstrom et al., 1979; R. C. Gur et al., 2001; R. C. Gur et al., 2010; R. C. Gur et al., 2012). Furthermore, this test has been used in multiple functional neuroimaging studies, and its processing has been mapped into a brain circuitry that includes frontal and temporoparietal components of the left hemisphere (R. C. Gur et al., 1982; R. C. Gur et al., 1983; R. C. Gur et al., 2000; R. C. Gur & Reivich, 1980). Therefore, performance deficits in the context of abilities in other domains can be linked to dysfunction of a specific brain circuitry. Computerized testing is increasingly used in large-scale genomic and clinical studies, where biomarkers and change indices are needed (Almasy et al., 2008; Calkins et al., 2010; Green et al., 2004; Insel & Cuthbert, 2009; McClay et al., 2011; Seidman et al., 2010) . In such studies verbal reasoning may not be the focus and yet its measurement is desirable if time permits. The abbreviated version of the PVRT developed here would be appropriate in such situations. It is available online for qualified investigators who want to use it in research with institutional review board oversight (http://www.med.upenn.edu/bbl).
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by NIH Grants T32MH065218 and P50MH096891.
