Abstract
Cognitive functioning may account for minimal levels (i.e., 5%–14%) of variance of performance validity test (PVT) scores in clinical examinees. The present study extended this research twofold: (a) by determining the variance cognitive functioning explains within three distinct PVTs (b) in a sample of patients with multiple sclerosis (pwMS). Seventy-five pwMS (Mage = 48.50, 70.6% female, 80.9% White) completed the Victoria Symptom Validity Test (VSVT), Word Choice Test (WCT), Dot Counting Test (DCT), and three objective measures of working memory, processing speed, and verbal memory as part of clinical neuropsychological assessment. Regression analyses in credible groups (ns ranged from 54 to 63) indicated that cognitive functioning explained 24% to 38% of the variance in logarithmically transformed PVT variables. Variance from cognitive testing differed across PVTs: verbal memory significantly influenced both VSVT and WCT scores; working memory influenced VSVT and DCT scores; and processing speed influenced DCT scores. The WCT appeared least related to cognitive functioning of the included PVTs. Alternative plausible explanations, including the apparent domain/modality specificity hypothesis of PVTs versus the potential sensitivity of these PVTs to neurocognitive dysfunction in pwMS were discussed. Continued psychometric investigations into factors affecting performance validity, especially in multiple sclerosis, are warranted.
Keywords
Neuropsychologists are tasked with determining the validity of cognitive data obtained from examinees (Larrabee, 2014; Sweet et al., 2021). Performance validity tests (PVTs), which may be standalone or embedded within existing measures (Wodushek & Domen, 2020), function to identify suboptimal test engagement or otherwise noncredible performance, as these factors meaningfully jeopardize the accuracy of interpretations of test findings (An et al., 2017). As such, PVTs are ideally designed to be as insensitive as possible to actual brain-related impairment (Green & Flaro, 2021), and factor analytic work has suggested that PVTs and bona fide measures of cognition tend to be represented by distinct latent dimensions (Ord et al., 2021; Van Dyke et al., 2013). Recommendations for continuous sampling of validity throughout neuropsychological assessment are provided by experts (e.g., Boone, 2009; Erdodi, 2021).
Notably, some (but not all; e.g., Critchfield et al., 2019) prior research has indicated that the presence of notable cognitive dysfunction may result in higher than expected rates of abnormal scores on PVTs, and ultimately increase the risk of false positives (i.e., incorrectly concluding noncredible performance; Laurent et al., 2019; Miller & Axelrod, 2018). In particular, this is commonly reported among examinees with very low IQ (Dean et al., 2008) and those with dementia syndromes (Davis, 2018; Dean et al., 2009). The degree to which the use of traditional PVT cutoff scores in those with milder degrees of cognitive impairment is mixed, however (Loring et al., 2016; Maiman et al., 2019; Quattlebaum et al., 2019; Teichner & Wagner, 2004). As such, understanding the degree to which tests of cognitive abilities psychometrically overlap with PVT scores is essential in neuropsychological research and practice.
Recently, Resch, Soble, et al., (2021) took a unique approach and determined the extent to which cognitive functioning explained variance within a standalone PVT in credibly performing examinees. Specifically, they explored overlap between the Victoria Symptom Validity Test (VSVT; Slick et al., 1997), a well-validated, memory-based PVT (see Resch, Webber, et al., 2021), and objective measures of auditory working memory (Wechsler Adult Intelligence Scale, Fourth Edition [WAIS-IV] Working Memory Index [WMI]; Wechsler, 2008), processing speed (WAIS-IV Processing Speed Index [PSI]; Wechsler, 2008), and verbal memory (California Verbal Learning Test, Second Edition Long Delay Free Recall [CVLT-II]; Delis et al., 2000). Their mixed clinical sample included 88 adults (M age = 31.7), most of whom had primary diagnoses of attention-deficit/hyperactivity disorder (ADHD; 39%), a psychiatric condition (28%), or a mild neurocognitive disorder (mND; 21%). Participants completed an average of 3.41 PVTs, and only those who passed all administered PVTs were operationally defined as valid and included to conservatively ensure that obtained test scores likely reflected optimal performance without excessive statistical artifact.
Resch, Soble, et al.’s (2021) results showed significant, yet modest, correlations between working memory/processing speed and logarithmically transformed VSVT variables. Verbal memory did not correlate with VSVT variables. Regression models showed that combined cognitive performance explained between 5% and 14% of the variance across several VSVT variables; working memory, processing speed, and verbal memory accounted for only 13% of the variance in VSVT Total correct. Notably, WMI was the only cognitive variable that tended to account for unique variance in VSVT performance, though at a marginally significant level (ps ranged from .04 to .10). In sum, Resch, Soble, et al.’s (2022) findings suggested that performance on the VSVT appeared to be minimally impacted by working memory, processing speed, and verbal memory.
Although meaningful in its contribution, Resch, Soble, et al.’s (2021) work poses three limitations. First, Resch, Soble, et al.’s (2021) operationalization of validity may have unnecessarily yielded data that were not adequately representative of a mixed clinical sample. As above, participants completed an average of 3.41 PVTs (apart from the VSVT) and conservatively were excluded if they performed outside expected limits on even one; this definition of validity may have been overly cautious (e.g., Schroeder et al., 2019), as recent work has indicated that examinees failing 0-1 PVT may be justifiably classified as valid, and those failing ≥ 2 classified as invalid (Rhoads, Neale, et al., 2021), though there is great debate in the literature regarding number of PVT failures needed to define invalidity (Odland et al., 2015; Proto et al., 2014). Of note, one-sample t tests by the present authors using Resch, Soble, et al.’s (2021) descriptive data indicated that neither WMI (M = 99.91, p = .951), PSI (M = 100.5, p = .740), nor CVLT-II (M = −.11, p = .493) were significantly different from normative, standardization population means. 1 If the included participants were representative of a sample composed of clinically referred examinees with diverse neurodevelopmental, psychiatric, and neurocognitive disorders, it is reasonable to expect that at least subtle-to-mild decreases from normative performance may be detectable in aspects of working memory, processing speed, and/or verbal memory. Prior literature has described such differences from healthy controls in these domains in samples of participants with ADHD (Fuermaier et al., 2015; Poysophon & Rao, 2018; Theiling & Petermann, 2016), psychiatric conditions (Hermens et al., 2011; Moghadasin & Dibajnia, 2021), and mND (Economou et al., 2007; Greenaway et al., 2006; Lumpkin & Sheerin, 2019). As such, it is likely that the participants included by Resch, Soble, et al. (2021) were more closely akin to a sample of cognitively intact/healthy participants (i.e., a normative sample) rather than a mixed clinical sample commensurate with typical neuropsychological practice, limiting the extent to which these results are generalizable to other clinically referred examinees.
Second, Resch, Soble, et al.’s (2021) focus was solely on the VSVT, limiting its applicability to other PVTs, particularly those using different test paradigms. Of note, recent work has suggested that PVTs may be domain/modality specific (see Erdodi, 2019). In the context of developing novel PVTs, the domain/modality specificity hypothesis posits that “the extent to which . . . variables are similar in relevant psychometric features (test paradigm, cognitive domain, sensory modality, etc.)” (Erdodi et al., 2018, p. 108) can affect their utility, classification accuracy, and psychometric robustness. By extension, PVTs may be most sensitive to noncredible or otherwise reduced performance within the cognitive domain(s) they appear to assess (Victor et al., 2013), and memory-based PVTs may fail to detect the signal of all manifestations of invalidity (Dandachi-FitzGerald & Merckelbach, 2013). Consider the VSVT, which appears to assess both functions of working memory and recognition memory. Examinees (consciously or otherwise) demonstrating invalid performance on those domains may be more likely to perform poorly on the VSVT and bona fide cognitive tests of those abilities. Other measures, including the Test of Memory Malingering (Tombaugh, 1996) and Word Choice Test (WCT; Pearson, 2009), more distinctly appear to assess visual and verbal recognition memory, respectively. Furthermore, measures such as the Dot Counting Test (Boone et al., 2002) appear to measure attention and processing speed without any appearance of assessing memory, such that they may be more appropriate to detect noncredible performance in speed than in memory. Continued work exploring the relationships between cognitive functioning across domains and PVTs of various paradigms/modalities is needed.
Third, although Resch, Soble, et al.’s (2021) utilization of a mixed clinical sample may provide benefits for generalizability, it included relatively high proportions of patients with ADHD diagnoses who were relatively young (M age = 31.7); moreover, conclusions drawn from such samples may be limited for clinical groups with more specific presenting concerns and neurological conditions. For example, there has been a burgeoning interest in performance validity in the context of specific disease states, including multiple sclerosis (Domen et al., 2020; Galioto et al., 2020; Sanborn et al., 2021), a chronic, inflammatory, autoimmune disease that presents with various clusters of physical, cognitive, and psychological symptoms (Macaron & Ontaneda, 2019). Base rates of invalidity in such patients may be upward of 20% (e.g., Galioto et al., 2020; Nauta et al., 2021), consistent with rates in other clinical populations (McWhirter et al., 2020).
The Present Study
In light of these limitations, the present study sought to contribute to the literature by extending Resch, Soble, et al.’s (2021) prior work by (a) determining the extent to which cognitive functioning in the domains of auditory working memory, processing speed, and verbal memory accounts for variance within three separate, standalone PVTs (b) in a sample of patients with multiple sclerosis (pwMS). It was hypothesized that, in line with prior work, cognitive functioning would account for relatively minimal (i.e., ≤15%) variance in PVT scores.
Methods
Patient Characteristics
Participants were identified from an archival clinical data set of pwMS referred for clinical neuropsychological assessment within a large, Midwestern, academic medical center. Seventy-five (75) participants with complete data for variables of interest were included (M age = 48.51, SD = 12.04), and they had an average of 13.87 years of education (SD = 2.37) and disease duration of 14.60 years (n = 73; SD = 12.48). Most participants were women (70.7%), identified as White (82.7%), and were diagnosed with relapsing-remitting subtype (80.0%). Of the 74 participants with known disability seeking status, just over one quarter (26.7%) reported applying for or considering applying for disability benefits. Demographic information, which appeared well-representative of population-level characteristics of pwMS and of the entire archival data set from which participants were selected, and other descriptive data are provided in Table 1.
Demographic and Descriptive Results
Note. Total N/n is equal to that denoted in the top row of each column unless otherwise noted. Group 1 (n = 63) represents participants who passed both the WCT and DCT-E. Group 2 (n = 54) represents participants who passed both VSVT and DCT-E. Group 3 (n = 56) represents participants who passed both the VSVT and WCT. MS = multiple sclerosis; RRMS = relapsing-remitting MS; PPMS = primary progressive MS; SPMS = secondary progressive MS; Disease Duration, Years = Years since reported symptom onset, not date of diagnosis; VSVT = Victoria Symptom Validity Test; WCT = Word Choice Test; DCT-E = Dot Counting Test E-score; DS = Wechsler Adult Intelligence Scale, Fourth Edition Digit Span subtest; ACSS = age-corrected scaled score, with M = 10 and SD = 3; SDMT = Symbol Digit Modalities Test, oral trial presented as a T score, with M = 50 and SD = 10; CVLT-II LDFR = California Verbal Learning Test, Second Edition, Long Delay Free Recall presented as a z score, with M = 0 and SD = 1.
Measures
Performance Validity Tests
Detailed descriptions of each PVT’s paradigm are not provided herein to protect test security. The Victoria Symptom Validity Test (VSVT; Slick et al., 1997) is a standalone, memory-based PVT. A recent analysis indicated the best performing variable on the VSVT was Total score ≤ 40 to indicate noncredible performance (Resch, Webber, et al., 2021), and only this score was of interest to the current study. Response latency–based variables were not considered given that they tend to contribute no unique benefit for classification accuracy above and beyond accuracy scores (Cerny et al., 2021).
The WCT (Pearson, 2009) is a standalone PVT with good construct validity demonstrated in experimental and clinical work (Erdodi et al., 2014; Lace, Grant, et al., 2021). A recent meta-analysis suggested an optimal cutoff score of ≤42 to denote noncredible performance (Bernstein et al., 2021).
The Dot Counting Test (DCT; Boone et al., 2002) is a standalone PVT with good construct validity in diverse clinical samples (Rhoads, Resch, et al., 2021; Soble et al., 2018). The E-score (DCT) was the variable of interest. The authors chose a conservative cutoff for noncredible performance of ≥22 to remain with similar work (Lace, Merz, & Galioto, 2021) and because this score may be appropriate in other populations with significant neurological involvement (e.g., cortical stroke, mild dementia; Boone et al., 2002).
Of note, the administration of three PVTs was deemed appropriate for routine clinical care given the relatively high base rate of invalidity seen in pwMS (Nauta et al., 2021) and review of previous literature that has administered four to five standalone PVTs to clinical neuropsychological examinees (Pliskin et al., 2021; Soble et al., 2020; Soble et al., 2021).
Cognitive Measures
The Digit Span subtest from the WAIS-IV (Wechsler, 2008) was used as a measure of auditory working memory. It is one of the two subtests that comprises the WMI composite—which was used by Resch, Soble, et al. (2021)—within the WAIS-IV. The age-corrected scaled score (ACSS; M = 10, SD = 3) for Digit Span was of interest.
The Symbol Digit Modalities Test (SDMT; Smith, 1973) was used as a measure of processing speed. It is a substitution task that is similar in appearance to and strongly correlated with Coding (previously Digit Symbol Coding) within various iterations of the WAIS (Bowler et al., 1992; Morgan & Wheelock, 1992). The oral trial was preferred over the written trial in the present study as it is widely used as the processing speed measure of choice in neuropsychological assessment of pwMS (Arnett & Strober, 2014; Benedict et al., 2006; Langdon et al., 2012) and minimizes psychomotor burden. Of interest was the age- and education-corrected T scores (M = 50, SD = 10) for each participant calculated according to the test manual (Smith, 1973).
The Long Delay Free Recall (CVLT-II) trial of the California Verbal Learning Test, Second Edition (CVLT-II; Delis et al., 2000) was used as a measure of verbal memory. Resch, Soble, et al. (2021) used this variable to represent verbal memory, as well. The CVLT-II is widely used as a verbal learning/memory test of choice with pwMS (Benedict et al., 2006; Langdon et al., 2012; Merz et al., 2018). The age- and gender-corrected z score (M = 0, SD = 1) for CVLT-II was of interest, with prior work supporting this score as a good representation of this construct (Donders, 2008).
Procedures
Participants were identified from an institutional review board–approved registry. All participants were referred for comprehensive neuropsychological assessment between 2020 and 2021 as part of medical care and were evaluated by a board certified clinical neuropsychologist or clinical neuropsychologist trained under Houston Conference Guidelines. All neuropsychological tests were administered by a Certified Specialist in Psychometry or competently trained psychometrist, clinical psychology doctoral student, or postdoctoral neuropsychology fellow familiar with standardized test administration and scoring procedures. Only participants with complete data for the PVT and cognitive variables of interest were included.
Statistical Analyses
Except for those performed in Microsoft Excel where described, analyses were performed with SPSS 26.0. Consistent with both Resch, Soble, et al. (2021) and logical expectation, PVT variables were not normally distributed, VSVT Total D(75) = .26, p < .001; WCT D(75) = .33, p < .001; DCT D(75) = .15, p < .001. VSVT Total (skewness = −1.66, kurtosis = 2.25) and WCT (skewness = −3.21, kurtosis = 11.10) demonstrated a negative skew, and DCT (skewness = 2.08, kurtosis = 6.79) demonstrated a positive skew. As such, VSVT Total and WCT were reflected (subtracting each participant’s score from the maximum score for either variable plus one) and logarithmically transformed. DCT was also logarithmically transformed. These approaches were consistent with Resch, Soble, et al.’s (2021) methods. No PVT scores were identified as univariate outliers (i.e., ≥|3.29| SD from the mean; Field, 2013) after transformations, and no scores for cognitive variables were identified as univariate outliers either.
Bivariate Pearson’s correlations between each pair of variables for the total sample were calculated. A series of linear regressions were conducted with Digit Span, SDMT, and CVLT-II scores entered as predictor variables as one set, and (logarithmically transformed) PVT variables entered as the outcome. Tolerance (acceptable values are ≥ .10) and variance inflation factor (VIF) diagnostics (acceptable values are ≤ 10; Hair et al., 2010) were examined as indicators of possible multicollinearity for each regression model. To control for Type 1 error in the context of three separate regression analyses, conservative critical α = .017 (i.e., .05/3) was chosen a priori for evaluating models and individual predictors.
In an attempt to borrow from Resch, Soble, et al.’s (2021) methodology of utilizing only participants operationally defined as “credible,” pwMS were included in a regression analysis for a PVT only if they passed both other PVTs. For example, for the VSVT regression (Group 1), pwMS passing both the WCT and DCT were included (n = 63). For the WCT regression (Group 2), pwMS passing both the VSVT and DCT were included (n = 54). For the DCT regression (Group 3), pwMS passing both the VSVT and WCT were included (n = 56). Participants were retained even if they performed outside normal limits on the PVT of interest for the specific regression. This approach was deemed appropriate, as the mean VSVT Total scores for pwMS passing both the WCT and DCT (M = 44.98, SD = 4.18) was not significantly different from the mean VSVT Total scores reported by Resch, Soble, et al. (2021; M = 44.90, SD = 4.50, Welch’s t [130] = 0.36, p = .723), 2 indicating that the operational definitions of credibility likely did not produce systematic differences in PVT scores between samples and studies.
Results
As shown in Table 1, the mean scores for each PVT were within normal limits for the total sample and for each group. Mean scores for Digit Span tended to fall in the average range, albeit at the lower end. Mean scores for SDMT and CVLT-II tended to fall in the low average range, with mild variability among groups (Guilmette et al., 2020). One-sample t tests indicated that mean scores for Digit Span, SDMT, and CVLT-II across all three groups were significantly lower than the population mean for each (all ps ≤ .001).
Bivariate correlations between each pair of variables for the total sample (N = 75) are shown in Table 2. Correlations between each pair of variables for each Group are presented in supplemental material (Supplemental Table 1). In the total sample, correlations between pairs of transformed PVTs were low to moderate (rs ranged from .26 to .52). Correlations between transformed PVTs and Digit Span, SDMT, and CVLT-II ranged from low to moderate (rs ranged from −.58 to −.16). Correlations between Digit Span, SDMT, and CVLT-II were generally low (rs ranged from .21 to .26) in magnitude and not suggestive of meaningful multicollinearity or psychometric redundancy. Notably, the strength of some correlations differed markedly across groups. For example, VSVT-CVLT Total sample r = −.46 (suggesting approximately 21% shared variance), while VSVT-CVLT Group 3 r = −.13 (suggesting approximately 1%–2% shared variance). Nearly all correlations across groups (except those falling close to zero) shared the same directionality (see Supplemental Table 1).
Bivariate Correlations in the Total Sample
Note. Total group N = 75. Pearson’s (r) correlation coefficients are displayed. VSVT = Victoria Symptom Validity Test Total Score; WCT = Word Choice Test; DCT = Dot Counting Test E-score; DS = WAIS-IV Digit Span Age-Corrected Scaled Score. CVLT-II = California Verbal Learning Test, Second Edition, Long Delay Free Recall z score.
For correlations, variable was reflected and logarithmically transformed. As such, the directionality of Pearson’s (r) correlation coefficients should be reversed when interpreting these results from a practical/clinical perspective, except for the Pearson’s correlation between logarithmically transformed VSVT Total Score and WCT (b), which can be interpreted as is. c For Pearson’s (r) correlation coefficients, variable was logarithmically transformed.
p < .05. **p < .01.
Table 3 displays results from regression analyses. For the VSVT Total score regression (n = 63), multicollinearity diagnostics were acceptable (Tolerance ranged from .96 to .99; VIF ranged from 1.01 to 1.04). Results indicated that the overall model was significant, F(3, 59) = 9.456, p < .001, and accounted for 32.5% of the variance in VSVT Total scores. Digit Span (p < .001) and CVLT-II (p = .006) emerged as significant individual predictors. Squared semipartial correlations indicated that Digit Span accounted for more than two and one-third times as much unique variance in VSVT Total scores (22.3%) than did CVLT-II (9.4%), and SDMT accounted for essentially no unique variance (0.1%). In light of the reflected and logarithmically transformed VSVT Total score, decreases in Digit Span and CVLT-II were associated with decreased VSVT Total scores.
Regression Models
Note. VSVT = Victoria Symptom Validity Test; WCT = Word Choice Test; DCT-E = Dot Counting Test E-score; DS = Wechsler Adult Intelligence Scale, Fourth Edition Digit Span subtest; SDMT = Symbol Digit Modalities Test, oral trial; CVLT-II LDFR = California Verbal Learning Test, Second Edition, Long Delay Free Recall; sr2 = semipartial correlation squared.
Variable was reflected and logarithmically transformed. As such, the directionality of regression coefficients should be reversed when considering these results from a practical/clinical perspective. b Variable was logarithmically transformed.
For the WCT regression (n = 54), multicollinearity diagnostics were also acceptable (Tolerance ranged from .92 to 1.00; VIF ranged from 1.00 to 1.08). The overall model was significant, F(3, 50) = 5.373, p = .003, and accounted for 24.4% of the variance in WCT scores. Only CVLT-II (p = .003) emerged as a significant individual predictor. Squared semipartial correlations showed that CVLT-II accounted for just over three times as much unique variance in WCT scores (14.5%) than did Digit Span (4.8%), and SDMT accounted for minimal unique variance (1%). In light of the reflected and logarithmically transformed WCT variable, decreases in CVLT-II were associated with decreased WCT scores.
For the DCT regression (n = 56), multicollinearity diagnostics were acceptable (Tolerance ranged from .92 to .97; VIF ranged from 1.03 to 1.09). The overall model was significant, F(3, 52) = 10.806, p < .001, and accounted for 38.4% of the variance in DCT scores. SDMT emerged as a significant individual predictor (p < .001), as did Digit Span (p = .001). Squared semipartial correlations indicated that SDMT accounted for approximately one and one-third times as much unique variance in DCT scores (18.6%) than did Digit Span (13.4%), and CVLT-II accounted for minimal (2.2%) unique variance. As DCT was only logarithmically transformed and not reflected, decreases in Digit Span and SDMT were associated with increased (i.e., trending toward invalidity) DCT scores.
Discussion
Neuropsychologists are tasked to determine the validity of obtained cognitive data, and some prior work has indicated that bona fide neuropsychological dysfunction may negatively impact examinees’ performance on measures of validity. The present study sought to determine the extent to which cognitive functioning in the domains of working memory, processing speed, and verbal memory accounted for variance within three separate, standalone PVTs in a sample of pwMS. Several findings deserve further discussion.
First, regarding demographics, the present sample appeared to well represent the population-level characteristics of pwMS, such that the modal pwMS tends to be a White, middle-aged woman with relapsing-remitting subtype (Hadgkiss et al., 2013; Minden et al., 2006). Relatedly, regarding cognitive test scores, although still within broad normal limits (i.e., low average to average range), the mean performances on cognitive tests in the total sample and its groups of operationally defined credible performers were significantly lower than normative data and appeared more consistent with scores derived from a clinical sample than those reported by Resch, Soble, et al. (2021). This pattern is consistent with other work indicating that samples of pwMS demonstrate low average to average performance on aspects of working memory, processing speed, and verbal memory (Archibald & Fisk, 2000; Beatty, 2004; Lafosse et al., 2013; Lindau et al., 2022; Merz et al., 2018).
Second, the pattern of obtained findings were contrary to both hypothesized results and to Resch, Soble, et al.’s (2021) conclusions. Resch, Soble, et al. (2021) found that working memory, speed, and verbal memory significantly but minimally predicted performance on VSVT, accounting for 13% of variance in total scores when all three neuropsychological domains were included in the regression model; a similar pattern was expected herein. In contrast, despite including the same or otherwise very similar neuropsychological measures (i.e., Digit Span, SDMT, and CVLT-II) and similar methodological and statistical approaches, the present study found that performance on neuropsychological tasks accounted for approximately between 24% and 38% of variance in PVT scores, with 32.5% on VSVT Total scores. These patterns are exceptionally higher than those reported by Resch, Soble, et al. (2021) and suggest that in pwMS—and potentially other clinical populations with even subtle yet meaningful cognitive difficulties—the psychometric overlap between PVTs and cognitive domains of working memory, speed, and verbal memory appear more strongly related than previously thought. Importantly, it appeared that the WCT was the least affected (relatively speaking) by cognitive functioning out of the three PVTs of interest and may be a promising tool for the assessment of validity in this population; this finding aligns with a recent poster presentation indicating that the WCT shared the less variance with objective neuroradiological disease burden in pwMS than did DCT and VSVT (Lace et al., 2022). Notably, the difference in magnitude between some correlations across groups may be due to idiosyncrasies of PVT paradigms (including the domain/modality specificity hypothesis discussed below) or that continuous PVT variables may remain somewhat non-normally distributed and affected by outliers even after applying appropriate statistical corrections (such as logarithmic transformation as was done herein and in Resch, Soble et al., 2021).
Of particular relevance was the pattern of relationships between cognitive tests and PVTs. Specifically, verbal memory significantly influenced both VSVT and WCT scores, working memory influenced VSVT and DCT scores, and processing speed influenced DCT scores. At least two possible explanations for these findings deserve consideration. First, it may be that the PVTs explored herein are especially (even overly) sensitive to bona fide neurocognitive dysfunction, particularly in this sample of MS patients. That is, while a “robust PVT is . . . relatively insensitive to actual cognitive ability(ies)” (Resch, Soble, et al., 2021, p. 1620), the degree to which these particular PVTs (i.e., VSVT, WCT, and DCT) remain insensitive to neurocognitive problems, at least with pwMS, may be inadequate. Consider the relatively high proportion of unique variance that Digit Span (13.4%) and SDMT (18.6%) accounted for in DCT scores. In addition to its purpose as a PVT, it is possible that the DCT actually is (at least in part) a legitimate measure of attention/working memory and processing speed. Similarly, it may be that both the VSVT and WCT (again, at least in part) assess verbal memory to a degree above and beyond what would be expected if they were only measures of performance validity. If this explanation is true, it may be the likelihood of incorrectly identifying invalidity may be unnecessarily elevated with increasing degrees of neurocognitive dysfunction. Of note, a recent systematic review and cross-validation of the VSVT specifically stated that “caution is recommended among patients with . . . working memory deficits due to concerns for increased risk of false positives” (Resch, Webber, et al., 2021, p. 331). Thus, it may be that the three PVTs utilized herein are too easily contaminated by neurocognitive dysfunction and of limited appropriateness for clinical use, at least with pwMS.
Alternatively, and a more favored explanation in the authors’ opinion, these findings may provide adjunctive support for the domain/modality specificity hypothesis of PVTs (Erdodi, 2019; Rai & Erdodi, 2021) rather than argue for their conservative or limited use. The construct of invalidity does not always manifest similarly across examinees nor is it equally detectable by all types of PVTs (e.g., Dandachi-FitzGerald & Merckelbach, 2013), as “patients may fail one type of [performance] validity test more than another because of the relevance of the test material to their presenting complaints” (Gervais et al., 2004, p. 476). The domain/modality specificity hypothesis is particularly relevant in developing new PVTs, as the match between criterion and predictor PVTs across test-specific characteristics can meaningfully influence their psychometric properties (Rai & Erdodi, 2021). By extension, PVTs that appear to measure certain cognitive domains (e.g., memory, processing speed) are likely to better detect noncredible performance demonstrated by examinees on actual tests of those abilities. Thus, experts recommend continuous, multi-domain sampling of validity throughout neuropsychological assessment (Boone, 2009).
Considering these findings in light of the domain/modality specificity hypothesis, it is reasonable to expect stronger psychometric relationships between variables that tap the same “cognitive construct” when performance on tests to that end trends toward invalidity. It may be that pwMS who performed invalidly with respect to attention/processing speed but not necessarily with memory performed worse more readily on tests that appeared to assess the former—which, by extension, include SDMT, Digit Span, and DCT. Relatedly, it may be that pwMS who performed invalidly with regard to memory but not necessarily with attention/processing speed more often performed worse than their actual abilities on tests appearing to measure memory—which, by extension, include VSVT, WCT, CVLT-II, and Digit Span. The relationship between PVTs and performance on “true” cognitive tests may be a function of the role of effort (or validity broadly speaking) within those neuropsychological factors. As such, it may be that the relatively high degree of psychometric overlap between constructually similar cognitive tests and PVTs speaks to their ability to detect the “signal” of noncredible performance within their specific domain/modality, rather than an oversensitivity to overt neurocognitive dysfunction or neurobiological insult. Consider the high proportion of unique variance accounted for by Digit Span in VSVT Total scores (22.3%). This finding may certainly reflect the paradigmatic resemblance of each test (i.e., recalling/recognizing strings of digits), and also relate to prior research supporting the use of Digit Span as a PVT (e.g., Babikian et al., 2006; Hurtubise et al., 2020; Shura et al., 2020; Webber & Soble, 2018), though exploration of this latter claim extends beyond the scope of this article. Both possibilities relate to the domain/modality specificity hypothesis and may, at least to some extent, account for the psychometric overlap between Digit Span and VSVT herein. Similar arguments may be made for both symbol-digit/digit-symbol coding tests and CVLT-II, which have been identified as possible embedded PVTs (e.g., Erdodi & Abeare, 2020; Erdodi et al., 2017; Persinger et al., 2018; Wolfe et al., 2010), albeit with relatively less support than the body of work on Digit Span.
Importantly, the authors firmly acknowledge that firm support for or refutations against either of the above explanations cannot be adequately reached by these findings alone given the cross-sectional nature employed herein. Future studies may seek to employ experimental methods to parse apart reasons behind and mechanisms of invalidity.
The present study, of course, is not without its own limitations. First, the current sample size was relatively small and thus increased the possibility of Type II error. However, this factor was not fatal as the number of participants in each regression met widely accepted statistical standards (Harris, 1985; Wilson VanVoorhis & Morgan, 2007). Second, though the present study aimed to replicate and extend findings from Resch, Soble, et al. (2021), some comparisons were not exact due to differences in operationalization of select measures and differing statistical analyses. Specifically, the current study did not include additional VSVT variables (e.g., latency), though the rationale for VSVT variables included (i.e., Total Score was the best performing variable in a recent meta-analysis; response latency is not additively helpful in identifying invalidity; Cerny et al., 2021; Resch, Webber, et al., 2021) was deemed appropriate by the authors. Finally, the use of logarithmic transformations is a common practice when variables demonstrate remarkable skew (Tabachnick & Fidell, 2013) and was consistent with methodology outlined by Resch, Soble, et al. (2021). However, such transformations beget some degree of interpretive difficulty and should be understood in the context of transformation (Feng et al., 2014; Tabachnick & Fidell, 2013). Future work may seek to utilize statistical techniques, such as categorical regression techniques (CATREG) in SPSS (see van der Kooij, 2007), which may minimize the need for data transformation(s).
Conclusion
In sum, the current study provided evidence that cognitive abilities, particularly working memory, processing speed, and verbal memory, may account for a greater portion of PVT performance than previously suggested. Alternative explanations were proposed, including the possible oversensitivity of these PVTs to bona fide cognitive dysfunction and providing support for the domain/modality specific hypothesis (Erdodi, 2019). Taken at face value, the authors are likely proponents of the latter explanation, such that PVTs may be more sensitive to invalid performance demonstrated on tests within narrow domains and that psychometric overlap (even between PVTs and cognitive ability tests) increases with greater paradigmatic similarity between measures. Nonetheless, clinicians should use prudence and careful judgment when interpreting PVT failures, especially in pwMS and similar populations, and utilize all clinical information in reaching conclusions. Encouragingly, cognitive functioning appeared least meaningfully related to WCT scores (24.4% variance) compared with VSVT (32.5% variance) and DCT (38.5% variance) scores. Continued work exploring invalidity and cognitive functioning in pwMS and other clinical populations is warranted.
Supplemental Material
sj-docx-1-asm-10.1177_10731911231178289 – Supplemental material for Standalone Performance Validity Tests May Be Differentially Related to Measures of Working Memory, Processing Speed, and Verbal Memory in Patients With Multiple Sclerosis
Supplemental material, sj-docx-1-asm-10.1177_10731911231178289 for Standalone Performance Validity Tests May Be Differentially Related to Measures of Working Memory, Processing Speed, and Verbal Memory in Patients With Multiple Sclerosis by John W. Lace, Victoria Sanborn and Rachel Galioto in Assessment
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
