Abstract
This study evaluated the performance characteristics, construct validity, and reliability of two computerized, self-administered verbal and visual recognition memory tests based on the Remember-Know paradigm. Around 250 healthy control participants and 440 patients referred for neuropsychological assessment used an iPad to complete the Words and Faces recognition memory tests before or after concurrent neuropsychological testing. Performance accuracy was high but without ceiling effects. Education, but not age, was related to overall performance for both samples while the influence of gender and race differed across samples. In the clinical sample, overall performance was worse in those patients demonstrating memory impairment on clinical assessment. Words and Faces subtests demonstrated the strongest correlations with neuropsychological measures of verbal and nonverbal memory, respectively. Both showed moderate correlations with processing speed while Faces was also correlated with visuospatial skills. The memory tests showed good test–retest reliability over two testing sessions. These findings demonstrate acceptable psychometric properties in clinical and community samples and suggest that this computerized format is feasible for memory assessment in clinical contexts.
Keywords
There is great potential to refine and expand neuropsychological assessment with the use of technology. Computerized cognitive testing has many advantages over more traditional paper-and-pencil measures including reduced personnel time for administration and scoring, fully standardized administration, errorless scoring, and high-resolution behavioral metrics that cannot be captured with standard testing (e.g., millisecond reaction times and eye movements; Bauer et al., 2012). Moreover, computerized measures have significant potential for use in remote/virtual assessments, and the global COVID-19 pandemic has made it abundantly clear that the field needs neuropsychological tools that can be completed remotely. Computerized versions of many standard clinical tests (e.g., Wisconsin Card Sorting Test, Wechsler Adult Intelligence Scale: 5th Edition) are now in widespread use. However, adapting traditional episodic memory measures to a computerized format has been challenging. Here, we describe the validation study for a new iPad-based instrument designed to measure immediate and delayed verbal and nonverbal recognition memory performance.
AACN Guidelines for Computerized Test Development
While the benefits of computerized assessment are widely recognized, there are also a number of potential limitations of computerized cognitive assessment tools that must be taken into consideration. To “promote accurate and appropriate use of computerized tests in a way that maximizes clinical utility and minimizes risks of misuse,” the American Academy of Clinical Neuropsychology (AACN) and the National Academy of Neuropsychology published a joint position paper on computerized neuropsychological assessment devices in 2012 (Bauer et al., 2012, p. 362). This position paper sets forth key issues pertaining to development and the use of computerized cognitive measures in health care settings. Importantly, they highlight the need to ensure that computerized assessment devices meet the same standards required for the development of psychological and neuropsychological measures, including psychometric test development. Standards for addressing potential end-user, examinee, and technical issues as well as privacy and data security are also provided (Bauer et al., 2012). The computerized Words and Faces memory tests described below and used in this validation study were developed in accordance with these standards. The Inter Organizational Practice Committee (IOPC; www.iopc.online) has further advocated for computerized testing standards that are conducive to remote assessment. The data reported here were collected in a clinic setting on institution-owned devices but a web-based remote administration version of these tests that adheres to the IOPC guidelines is currently under development, and utility for a clinic-to-home service model will be determined in future publications.
From a broader perspective, it is important to note that innovations in cognitive testing and the associated reliance on technology may also increase barriers to care for some populations. As Alberto Fernandez (2019) noted in a special issue of The Clinical Neuropsychologist dedicated to “Modern Neuropsychological Assessment Methods,” computerized testing innovations have largely benefited the Western, Education, Industrialized, Rich, and Democratic (WEIRD) countries with widespread access to computers and the internet. In less-developed countries, or socioeconomically depressed regions of developed countries, with lower educational attainment and fewer technology resources, computer-adapted or administered measures will further exclude people from care or exacerbate cultural differences related to computer use. More work is necessary to understand and maximize the utility of computer-administered tests in special populations such as low socioeconomic status (SES) or those with limited educational attainment.
Words and Faces Test Development
The computerized memory assessment tool described here consists of four memory subtests (Words Immediate, Faces Immediate, Words Delayed, and Faces Delayed) with an intervening distractor task between the Immediate and Delayed Trials. The Words and Faces tests employ the Recollection–Familiarity procedure (Tulving, 1985), a well-established cognitive neuroscience paradigm that permits assessment of not only the accuracy of memory but also the qualitative aspects of memory (i.e., confidence that the item was or was not previously seen). High-confidence (recollection) responses are more likely to reflect episodic recall of the item and associated contextual information (i.e., mentally re-experiencing the initial exposure to that item), whereas low-confidence (familiarity) responses more likely reflect recognition of the item without the associated contextual information (i.e., the initial exposure is not re-experienced). Imaging studies show that recollection responses are associated with greater activity in hippocampal–prefrontal memory circuits relative to familiarity responses (Scalici et al., 2017; Westerberg et al., 2013). Moreover, in patient studies, older adults with mild cognitive impairment or dementia are less accurate with both recollection and familiarity judgments while normal aging has a small influence on familiarity and recollection performance (Koen & Yonelinas, 2014). These differences in memory quality are supplemented by additional differences in reaction times such that familiarity judgments are typically slower than recollection. We have previously studied the psychometric properties of this memory format in neuropsychological patients and found that it can be effectively used to identify memory impairment in a computerized context (Busch et al., 2019). Indeed, the rich data derived from this memory procedure potentially afford a greater depth of measurement to identify early signs of cognitive decline.
As such, the recognition testing in the current memory subtests includes four response options: Definitely Old, Maybe Old, Maybe New, and Definitely New. Participants are instructed to select “Definitely Old” if they are SURE that they saw the item among the studied items and to select “Maybe Old” if they THINK that the item MAY have been among the studied items. They are instructed to select “Definitely New” if they are SURE that they did not see the item among the study set and to select “Maybe New” if they THINK that they MAY NOT have seen the item among the studied set.
In the Words Immediate subtest, participants view 20 words, presented one at a time, for 3 seconds each. Immediately after the final word is presented, participants are instructed to identify the 20 previously studied “Old” words from among 20 “New” distractor words that were not previously seen. Words are presented one at a time during the recognition phase and remain on the screen until the participant responds. Words were selected from a published database of English words (Brysbaert et al., 2014; Brysbaert & New, 2009) based on frequency in the English language (zipf value average = 3.0, SD = 0.2 where 1 is lowest frequency and 7 is highest frequency), word length (average 6.9 letters, SD = 1.9), and concreteness (average rating was 2.2, SD = .4 where 1 is abstract and 5 is concrete). Distractor words were drawn from the same source and matched for frequency, length, and concreteness. All words were known to 100% of the raters in the Brysbaert and New study. The words are presented in Arial type with a 46-point font size (Figure 1A).

Representative Display of Three Cognitive Tasks.
In the Faces Immediate subtest, participants view 20 black-and-white photographs of faces, presented one at a time for 3 seconds. Faces are emotionally neutral and vary across age (younger to older adults), gender (50% female), and race (60% white). Immediately after the final face is presented, participants are instructed to identify the 20 previously studied “Old” faces from among 20 “New” black-and-white photographs of cropped faces, matched on the basis of age, gender, and race. Faces are presented one at a time during the recognition phase and remain on the screen until the participant responds. All photos were taken in a professional photography studio and cropped to an oval shape to minimize nonfacial attributes (i.e., hair, clothing, jewelry). Face stimuli were approximately 4 inches high × 3.5 inches wide (Figure 1B). All individuals provided written informed consent for use of their photographs.
At the start of each Immediate subtest, two practice stimuli are presented and the participants are then asked to identify the two practice stimuli intermixed with two distractors. The participant is not permitted to continue until the correct response is recorded. After the immediate subtests, participants then complete a self-administered, iPad-based spatial span distractor task modeled on the Corsi Blocks procedure (Corsi, 1972). For this task, a spatially distributed array of nine squares “light up” in sequences of two to nine squares (Figure 1C). The application instructs participants to observe the sequence and, when cued, to touch the same blocks in the same order. Participants complete two trials at each sequence length and the test is discontinued when both trials were performed incorrectly.
The Words Delayed and Faces Delayed subtests follow the distractor task. In both Delayed subtests, 40 stimuli are presented one at a time, and participants are instructed to identify the 20 originally studied words or faces from among 20 new, matched distractor words or faces. Stimuli remain on the screen until the participant makes a response.
The memory tests were designed for a user-friendly iPad interface to allow patients to independently complete tasks, without requiring face-to-face administration time or supervision. Auditory instructions are delivered via over-the-ear headphones. This format facilitates comprehension in patients with limited reading ability, prevents external coaching or aid, and reduces environmental distractions, which allows test completion in a variety of unsupervised settings including waiting rooms, private clinic rooms, or inpatient units. (A study to determine whether the test can be reliably completed remotely, such as in a patient’s home, is currently underway.) The computerized format also captures millisecond reaction times to complement accuracy metrics and tracks the patient’s progression through the application. Compared to standard desktop or laptop platforms, the tablet format is more portable and allows for direct interaction with stimuli by virtue of the touchscreen rather than the indirect mapping of a keyboard or mouse response. Compared to a smart phone, the tablet platform retains a relatively large screen to ensure adequate visual display size and spatial separation of response options. Performance is automatically scored by the program, which makes scoring error- and bias-free.
Study Aims and Hypotheses
The primary goal of this study was to examine the psychometric properties of these self-administered iPad memory tests. Specifically, we compared performance on the iPad Words and Faces subtests to performance on standard neuropsychological measures in a large group of community-dwelling neurologically normal adults as well as a diverse sample of patients referred for neuropsychological testing. A second goal was to compare test performance over repeated administration. We hypothesized that the iPad memory tests would demonstrate good construct validity and test–retest reliability.
Methods
This prospective study was approved by the local Institutional Review Board, and all participants provided written informed consent for study participation consistent with the Declaration of Helsinki. Participants received a parking voucher, gift card, and/or stipend as compensation for their time. Study data were managed using REDCap electronic data capture tools (Harris et al., 2009, 2019). We report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study. Deidentified data and access to the computerized tool may be made available to researchers with appropriate confidentiality/data use agreements.
Participants and Procedures
Healthy Controls
A total of 250 healthy adult participants (ages 18 and older) recruited via study flier were included in the study. Participants were medical center employees, friends or family of a medical center patient, or community members attending on-campus health fairs. All participants were English-speaking and denied personal history of neurological or psychiatric disorder. Demographic characteristics are displayed in Table 1.
Characteristics of Study Samples.
Participants completed a demographic/history questionnaire that included a question about personal comfort with using computers and a brief battery of neuropsychological measures that included: Reading subtest from the Wide Range Achievement Test: Fourth Edition (WRAT-4), Boston Naming Test—Short Form (BNT-SF; Mack et al., 1992), Trail Making Test—Parts A and B (TMT), Judgment of Line Orientation—Short Form (JOLO-SF; (Benton et al., 1983), the Digit Span, Spatial Span, and Faces subtests of the Wechsler Memory Scale: Third Edition (WMS-III), and the Rey Auditory Verbal Learning Test (RAVLT). Given time constraints, a stand-alone effort measure was not administered to healthy control participants; however, examination of an embedded effort measure (i.e., reliable digit span) revealed only four patients with scores below expectation (Schroeder et al., 2012) and excluding data from these individuals did not alter the findings. The iPad tests and research battery were administered in a counterbalanced order.
Clinical Sample
A total of 451 adult patients (ages 18 and older) who were referred for clinical neuropsychological evaluations at a large, tertiary medical center participated in the study. All participants were English-speaking and were referred for evaluation by a treating provider within the institution. Specifically, referral sources were as follows: Memory Disorders Clinic = 154 (35.0%); General Neurology = 83 (18.9%); Movement Disorders Clinic = 52 (11.8%); Epilepsy Clinic = 43 (9.8%); Multiple Sclerosis = 23 (5.2%); Brain Tumor = 17 (3.9%); Other = 68 (15.5%). Eleven patients were excluded from the analysis due to invalid neuropsychological testing results based on scores below cutoff on two or more (Jennette et al., 2022; Rhoads et al., 2021) performance validity measures (Victoria symptom validity test [VSVT] hard score <18: (Grote et al., 2000; Keary et al., 2013; Test of Memory Malingering Trial 1 <40: Denning, 2012; Reliable Digit Span <6: Schroeder et al., 2012). Demographic characteristics for 440 included patients are displayed in Table 1.
Patients completed a demographic/history questionnaire that included a question about comfort with using computers and neuropsychological measures selected by the staff neuropsychologist to address the clinical referral question. As such, cognitive measures varied across patients. However, all clinical evaluations included assessment of intellectual ability, language, visuospatial functioning, attention, processing speed, working memory, verbal and/or nonverbal memory, and executive function. Self-report measures of depression and anxiety were also completed by nearly all patients. The iPad and clinical neuropsychological measures were administered in a counterbalanced order whenever possible given clinical demands. The staffing neuropsychologist was blinded to iPad performance.
Analyses
All analyses were conducted separately for the healthy controls and clinical sample. Sample size was determined based on feasibility and the obtained sample sizes provided adequate power to detect effect sizes of Cohen’s d = 0.20 or greater. Descriptive statistics for performance metrics (accuracy, hit and false-positive rate, reaction time, and confidence) were calculated for each sample (healthy controls and clinical sample). Normality of continuous variables was assessed visually using Q-Q plots. All performance metrics demonstrated skewed distributions and are presented as median with interquartile range and subject to nonparametric tests. Continuous demographic variables were recoded into ordinal categorical variables to facilitate group comparisons. Age categories were 18 to 29, 30 to 39, 40 to 49, 50 to 59, 60 to 69, and 70 to 84. Education categories were high school or less (≤12 years), college degree (13–16 years), and advanced degree (≥17 years). Analyses of associations between iPad memory performance and demographic variables were conducted using Wilcoxon rank sum tests for binary demographic variables and Kruskal–Wallis tests with Dwass Steel Critchlow–Fligner (DSCF) post hoc tests for demographic variables with three or more levels. Computerized memory performance was also compared between neuropsychological patients with and without memory impairments on clinical evaluation (operationalized as an immediate or delayed recall or recognition score 1.5 SD or more below the mean on two or more established verbal or visual memory tests; note that two impaired scores occurred on individual “impaired” tests in almost all cases and that memory impairments represented a mixture of encoding/storage and retrieval-based deficits).
To assess convergent and divergent validity separately in the healthy control and clinical samples, correlation analyses were used to assess the strength of association between each memory subtest accuracy and raw scores on each standard neuropsychological test. Age, education, gender, and race were considered as control variables in partial correlation analyses but ultimately were not needed because inclusion of control variables did not change observed correlation effects. Correlation analyses were also used to evaluate test–retest reliability in a subsample of participants who completed the Memory subtests at two time points. Pearson’s correlations were used for linear associations, and Spearman’s rank-order correlations were used for monotonic associations. Scatterplots were used to visually screen for bivariate outliers. Given the large number of comparisons, correlation coefficients were interpreted based on effect size, rather than p values, to limit the probability of interpreting spurious findings. Effect sizes were defined using Cohen’s (1969) recommendations for small (0.10), medium (0.30), and large (0.50) effects.
Analyses were two-tailed and completed using SAS Studio version 3.5 (SAS Institute Inc.) with alpha set to 0.05.
Results
Participants completed the entire memory protocol, including the distractor task, in an average of 19.3 minutes (SD = 2.2). The immediate memory trials were completed in 10.4 minutes (SD = 1.2), the distractor task was completed in 4.4 minutes (SD = 0.98), and the delayed memory trials were completed in 4.5 minutes (SD = 1.2). Overall accuracy and reaction time data for both samples are provided in Table 2. Healthy control performance stratified by age, education, gender, and race is provided in Supplemental Tables 1 to 4.
Overall Memory Performance: Median (Interquartile Range).
Healthy Controls
Accuracy
The total percentage correct items for combined subtests was moderately left skewed (skewness = −.99, kurtosis = 1.79) but did not demonstrate ceiling effects (Figure 2A). Accuracy rates did not differ across the four subtests. Participants who self-identified as female were more accurate overall (Wilcoxon Z = −2.5, p = .012) due to slightly better performance on the Faces subtests (Immediate Z = −2.8, p = .003; Delayed Z = −2.0, p = .042). There were no gender differences on the Words subtests. Education group (Kruskall–Wallis H = 19.0, p < .001), but not age group, was related to total score. The high school or less group was less accurate than the other two groups (compared to college degree DSCF = 4.7, p = .003; and compared to advanced degree DSCF = 5.6, p < .001). Participants who self-identified as White had higher overall scores than those not identifying as White (Z = 4.4, p <.001); this was due to differences on the Faces subtests (Immediate Z = 4.7, p < .001; Delayed Z = 4.6, p < .001). There was a small correlation of total score with self-rated computer use comfort (rho = .18).

Total Proportion Correct.
Hits/False Positives
The overall number of hits and false-positive responses are displayed in Table 2; there was no difference across subtests. We also calculated the a’ discriminability index (a nonparametric version of d’ which is valid under circumstances when the false-alarm rate may be 0) and found no differences across subtests. There was no effect of age on hits or false positives. There was an effect of education on both hits (H = 20.5, p < .001) and false positives (H = 6.2, p = .044); participants with high school or less had fewer hits than both those with college degrees (DSCF = 6.3, p < .001) and advanced degrees (DSCF = 3.7, p = .024), while participants with a college degree had more false positives than those with advanced degrees (DSCF 3.4, p = .039). Participants who self-identified as female had more hits than males (Z = −2.8, p = .005). Participants who self-identified as White had more hits (Z = 3.3, p = .001) and fewer false positives (Z = −2.4, p = .017) than those who did not self-identify as White.
Confidence Judgments
Overall, participants made more “Definitely” responses than “Maybe” responses (Table 2). There were no differences across subtests. Consistent with the expected difference in memory quality between “remembered” and “known” events, “Definitely” responses were more often correct (86%) than “Maybe” responses (65%) overall and for all subtests (χ 2 = 1,338.7, p < .001). Participants who self-identified as male selected “Definitely” responses more often than those self-identifying as female, for all subtests except Immediate Faces (Words Immediate Z = 2.3, p = .024; Faces Immediate Z = 1.3, p = .183; Words Delayed Z = 2.3, p = .023; Faces Delayed Z = 2.1, p = .038). Participants who did not self-identify as White selected “Definitely” responses more often than those self-identifying as White, for all subtests except Immediate Faces (Words Immediate Z = −3.1, p = .002; Faces Immediate Z = −1.8, p = .080; Words Delayed Z = −2.7, p = .007; Faces Delayed Z = −2.2, p = .025). There were no differences in confidence among age or education groups.
Response Times
Reaction times are reported in Table 2. Correct responses were faster than incorrect responses (Z = 29.2, p < .001) and “Definitely” responses were faster than “Maybe” responses (Z = 64.6, p < .001); there was no interaction of accuracy and confidence. Reaction times were faster among younger participants (H = 49.4, p < .001), with a more pronounced age effect for correct responses (H = 56.5, p < .001) than for incorrect responses (H = 18.5, p = .002). Compared to those with more than a high school education, participants with a high school education or less had longer reaction times overall (H = 6.4, p = .040) and for correct responses (H = 6.7, p = .036), but not for incorrect responses (H = .75, p = .688). There were no gender or race differences in reaction time.
Convergent and Divergent Validity
Table 3 displays the associations between the Memory subtests and raw performance scores on the brief neuropsychological battery. In general, the Immediate subtests showed higher correlations than the Delayed subtests. The Words Immediate showed the highest correlations (r = .38–.47) with RAVLT and moderate correlations with the WMS-III Faces tests. The Immediate Faces showed the opposite trend with the strongest relationships (r = .49–.59) observed with the Faces I and II subtests and moderate correlations with the RAVLT. Note that the size of these correlations is very similar to correlations reported between existing memory measures (Wechsler, 2009). Both Words and Faces also showed moderate correlations with visuospatial skills, processing speed, and visual/verbal attention span.
Correlations Between Subtests and Neuropsychological Measures in Healthy Controls.
Note. Correlations represent Pearson’s correlations unless otherwise noted. Light green cells = moderate effect sizes (i.e., 0.30–0.49); Dark green cells = large effect sizes (i.e., ≥0.50). RAVLT = Rey auditory verbal learning test; WMS = Wechsler memory scale; BNT = Boston naming test; JOLO = judgment of line orientation.
Spearman’s correlation; correlations were generated using demographically corrected standard scores or t-scores for the neuropsychological measures and raw scores for the computerized memory tests.
Test–Retest Reliability
A subsample of 41 participants from the Healthy Control sample completed the Memory subtests at two time points. Median inter-test interval was 17 days (interquartile range [IQR] = 14, 29; range = 10–84). The sample was 67% female and 45% White, with a mean age at first testing of 48.9 (SD = 16.2) and 14.3 years of education (SD = 2.7). On average, participants correctly identified 7.2 more test items (4.5%) on the second test session. Performance between the two test sessions was highly positively correlated (r = .7; See Figure 3A) indicating good test–retest reliability.

Test–Retest Performance.
Clinical Sample
Accuracy
As with the control sample, percent correct was moderately left skewed (skewness = −0.73, kurtosis = 0.51) and without ceiling effects (Figure 2B). Accuracy rates did not differ across subtests. The accuracy findings mirrored those of the control group. Again, education group (H = 14.6, p < .001), but not age group, was related to total score. The high school or less group was less accurate than the other two education groups (DSCF = 4.0, p = .012 compared to college degree and 5.0, p = .001 compared to advanced degree). Patients who self-identified as female were more accurate overall (Wilcoxon Z > −2.9, p = .004) with slightly better performance on all subtests (all Wilcoxon Z < −2.1, p < .05). There was no effect of race on accuracy. Correlation with computer use comfort was small (rho = .21). Importantly, patients with memory impairment identified on the clinical evaluation had significantly lower total scores and lower subtest scores compared to patients without memory impairment on clinical evaluation (all Wilcoxon Z < −5.5, p < .001).
Hits/False Positives
The number of hits and false positives did not differ across subtests. We also calculated the discriminability index (a’) and found no differences across subtests. There was no effect of age or race on hits or false positives. There was an effect of education (H = 9.6, p = .008) such that patients with high school or less made more false-positive errors than those with an advanced degree (DSCF = 4.2, p = .008). Patients who self-identified as female had a higher hit rate than those self-identifying as male (Z = –2.5, p = .014). Patients with memory impairment identified on the clinical evaluation had a lower hit rate (Z = −4.6, p < .001) and made more false-positive errors (Z = 6.2, p < .001) than patient without memory impairment.
Confidence Judgments
Overall, participants made more “Definitely” responses than “Maybe” responses. Participants selected “Definitely” significantly less often during the Words Delayed subtest, when compared to all other subtests (all DSCF ≥ 3.9, p < .05). In addition, “Definitely” responses were more likely to be correct (85%) than “Maybe” responses (66%) (χ2 = 2,310.5, p < .001). This was consistent across subtests. There were no gender, age, education, race, or memory impairment status differences in the proportion of “Definitely” responses selected.
Response Times
Findings again mirrored the healthy controls. Correct responses were faster than incorrect responses (Z = 36.7, p < .001), and this was observed for every subtest. “Definitely” responses were faster than “Maybe” responses (Z = 86.1, p < .001). There was no interaction of accuracy and confidence. Reaction times were faster among younger participants overall (H = 92.3, p < .001) and for both correct responses (H = 90.8, p < .001) and incorrect responses (H = 70.7, p < .001). There was no effect of education or race on reaction time. Females were faster than males overall (Z = 2.7, p = .007) and for correct responses (Z = 2.6, p = .009), but there was no gender difference in reaction time for incorrect responses. Patients who demonstrated memory impairment on clinical testing showed longer reaction times overall (Z = 2.9; p = .004) and to correct items (Z = 3.2, p = .001) but not incorrect items.
Convergent and Divergent Validity
Table 4 displays the associations between the memory subtests and standard neuropsychological measures of memory across the entire clinical sample. Again, the size of these correlations is very similar to correlations reported between existing memory measures (Wechsler, 2009). The Words Immediate and Delayed subtests showed the strongest correlations (r = .31–.44) with select neuropsychological measures of verbal memory (i.e., RAVLT, Hopkins Verbal Learning Test [HVLT]; Logical Memory subtests from WMS-IV). Correlations with the California Verbal Learning Test: Second Edition (CVLT-II) were notably smaller. The Words subtests also demonstrated generally moderate correlations with select measures of visual memory (i.e., Brief Visual Memory Test [BVMT], Faces I & II from WMS-III).
Correlations Between Subtests and Episodic Memory Measures in the Clinical Sample.
Note. Correlations represent Pearson’s correlations unless otherwise noted. Light blue cells = moderate effect sizes (i.e., .30–.49); Dark blue cells = large effect sizes (i.e., ≥.50). RAVLT = Rey Auditory Verbal Learning Test; HVLT = Hopkins Verbal Learning Test; CVLT = California Verbal Learning Test; WMS = Wechsler Memory Scale; LM = logical memory; BVMT-R = Brief Visual Memory Test—Revised.
Spearman’s correlation; correlations were generated using demographically corrected standard scores or t-scores for the neuropsychological measures and raw scores for the computerized memory tests.
The Faces subtests showed the strongest correlations (r = .22–.50) with WMS-III Faces I & II subtest and Immediate and Delayed recall on the BVMT. The Faces subtests also showed moderate correlations with all word list learning measures of verbal memory, as well as fairly strong correlations (r = .24–.41) with the Logical Memory subtests of WMS-III and IV.
Table 5 displays the associations between subtests and neuropsychological measures that primarily index abilities other than memory. The Words Immediate and both Faces subtests showed moderate correlations with several tests dependent on visuomotor processing speed. Immediate and Delayed Words did not show consistent correlations in other domains. The Faces subtests also showed moderate correlations with several measures assessing fluency, visuospatial skills, and speeded executive function.
Correlations Between Computerized Memory Subtests and Other Neuropsychological Measures in the Clinical Sample.
Note. Correlations represent Pearson’s correlations unless otherwise noted. Light green = moderate effect sizes (i.e., .30–.49). BNT = Boston naming test; Semantic fluency = MOANS (Mayo older American normative studies) animals, fruits, vegetables; JOLO = judgment of line orientation; WAIS = Wechsler adult intelligence scale; WMS = Wechsler memory scale; DKEFS = Delis–Kaplan executive function system; WCST = Wisconsin card sorting test.
Spearman’s correlation; correlations were generated using demographically corrected standard scores or t-scores for the neuropsychological measures and raw scores for the computerized memory tests.
Test–Retest Reliability
A subsample of 40 patients from the Clinical sample completed the Memory subtests at two time points. Median inter-test interval was 56 days (IQR = 26–119; range = 7–226). The sample was 65% female and all patients were White, with a mean age at first testing of 60.4 years (SD = 14.5) and 15.1 years of education (SD = 2.3). On average, patients correctly identified 9.9 more test items on the second test session. Scores between the two test sessions were highly positively correlated (r = .76, Figure 3B) indicating good test–retest reliability.
Discussion
The computerized Words and Faces subtests employ a Remember-Know recognition format to evaluate verbal and visual memory. Patterns of accuracy and performance in both groups conformed to those identified in experimental settings with elaborate training on the Remember-Know procedure. For example, “Definitely” responses were more likely to be correct than “Maybe” responses, indicating that “Definitely” responses reflected a more reliable memory experience. Likewise, “Definitely” responses were faster than “Maybe” responses, suggesting better “access” to, or less time necessary to verify, remembered items.
In our samples of healthy controls and patients referred for neuropsychological assessment, performance was related to a number of demographic variables, similar to standard neuropsychological measures of memory (Strauss et al., 2006). In both groups, lower performance was noted in males and in individuals with high school education or less, and related normative scores are provided in the supplemental tables. There was also an effect of race in the healthy control group, which was much more racially diverse than the clinical sample. The finding that those who self-identify as White did slightly better on the faces task is likely a product of the Own-race bias (see Meissner & Brigham, 2001 for review) where unfamiliar faces from other races or cultures are harder to remember than faces from within one’s race or culture. This has also been observed in the validation studies for standard neuropsychological tests or race-specific normative studies (Lucas et al., 2005; Schneider et al., 2015) indicating that the computerized Faces task is subject to the same biases that are important to recognize in test interpretation. Sixty percent of the faces presented in the test were White while the other 40% were a mix of other races, leading to a mismatch between stimuli and the racial makeup of our normal control sample. As such, we have included a normative table to describe performance based on race.
Age group was not related to performance in either participant group which initially seems counterintuitive. However, several aspects of the task design, alone or in combination, may have minimized age differences in performance. Prior studies show that normal aging has smaller effects on recognition testing compared to free recall (Danckert & Craik, 2013), recollection–familiarity paradigms yield smaller age effects on familiarity responses relative to recollection responses (Koen & Yonelinas, 2014), and age effects on recognition testing tend to be smaller when the to-be-remembered information is individual items rather than associated or paired information (Old & Naveh-Benjamin, 2008). Furthermore, unlike many standardized memory tasks, the foils for the Words subtests were not selected to create perceptual or semantic interference with the list words and this could also lessen the effect of age on performance given that age increases susceptibility to interference in recognition measures (Wilson et al., 2018). It may also be the case that use of abstract words, as in these computerized memory tasks, may lessen the benefit for younger subjects who are more likely to engage in dual encoding of more concrete words (Peters & Daum, 2008). Other untested influences could be that low-frequency words, which are easier to recognize, may work in older adults’ favor, or that the long presentation time (4 seconds) allowed older adults to engage in more elaborate encoding than is typically permitted with shorter presentation times (1 item per second). Regardless, the data suggest that the computerized format of the memory subtests did not disproportionately disadvantage older adults in our samples. Test performance was only weakly associated with ratings of comfort with computer use (rho = .18–.21), and no participants discontinued due to difficulty understanding how to complete the task with the iPad despite standard scores on the WRAT Reading test, a measure highly correlated with Full Scale IQ, as low as 55.
The memory subtests showed good convergent validity (based on presence of medium to large effect sizes) with traditional neuropsychological measures of verbal and visual memory in a large, diverse sample of adult patients and control participants. Indeed, the subtests correlated well with most measures of free recall despite the recognition format. The magnitude of correlations observed between the Words and Faces subtests and established memory tests were similar to those reported between domain-specific subtests in the normative sample from the Wechsler Memory Scale: Fourth Edition (Wechsler, 2009). For example, moderate correlations (r = .42–.44) were observed between the WMS-IV Logical Memory and Verbal Paired Associates tests. Likewise, moderate correlations (r = .38–.47) were observed between the WMS-IV Designs and Visual Reproduction. The small correlations observed herein the clinical sample between the Words Immediate and Words Delay with the CVLT-II (but not the other word lists) are particularly surprising. It is unclear to what extent procedural differences between the remember–know recognition format and free recall tests account for the low correlations with selected delay recall of word lists in both the control and clinical samples. Aside from case studies in densely amnestic patients, there are little data to directly compare the two test formats. However, studies that measure delayed free recall and require participants to provide remember–know judgments of each recalled word have consistently reported that both remember and know processes contribute to free recall, although the majority of recalled words are judged as “remembered.” If performance on the Words delayed subtest was disproportionately weighted toward Know processes, either as function of the test demands or the shorter duration of the delay interval on the computerized test (Yonelinas, 2002), this could explain the poor correlations between the Words Delayed subtest and delayed recall from standardized word lists. However, post hoc analyses to control for proportion or accuracy of remember/know responses did not substantially influence correlation magnitudes. It is also possible the relatedness or imageability of to-be-remembered words (highest in the CVLT-II, less in the RAVLT/HVLT, and least in the Words Delayed subtest) may contribute to selectively lower delayed memory correlations. Nonetheless, the similarities in correlation magnitude between the computerized and other examiner-administered memory tests suggest that the subtests effectively measure episodic memory using an unsupervised, computerized format.
The computerized memory subtests also demonstrated moderate correlations with a number of established nonmemory measures. The vast majority of these were tests with a processing speed component including Symbol Search, Coding, versions of Trails A and B, Delis–Kaplan Executive Function System (DKEFS) Inhibition and Inhibition/Switching, Block Design, and versions of semantic fluency. The consistent association between the computerized memory tests and processing speed measures may be a result of stimulus duration. Each word or face stimulus is presented for 4 seconds. It may be the case that the extra time allows “fast processors” to more deeply encode to-be-remembered items compared to “slow processors.” This would also explain why Faces subtests correlated more consistently than Words subtests with processing speed measures; more opportunity to encode is likely to convey more benefit for faces than words given that novel faces are inherently less meaningful than words.
The Faces subtests also showed moderate associations with several speeded and unspeeded tests requiring visuospatial skills. This is not particularly surprising given that visuospatial skills are clearly necessary for facial processing. Otherwise, aside from the relation with processing speed, the Words and Faces subtests showed weak or sporadic correlations with tests in other cognitive domains, suggesting reasonable divergent validity.
Taken together with the finding that patients who demonstrated memory impairment on clinical evaluation had significantly lower scores on the Words and Faces tests compared to patients without memory impairment, the data as a whole suggest that the Words and Faces tests may have clinical utility in identifying memory impairment. Moreover, findings from the cognitive neuroscience literature suggest that specific profiles of performance on the Words and Faces tests, such as proportion or accuracy of “definitely” versus “maybe” responses, may be useful in dissociating damage or dysfunction in specific neural circuits (i.e., hippocampal versus frontal; Scalici et al., 2017; Westerberg et al., 2013) or in differentiating normal aging from Mild Cognitive Impairment (Koen & Yonelinas, 2014). Ongoing and future studies will examine the relationship between patterns of performance on the Words and Faces tests and specific neurological diagnoses (i.e., mild cognitive impairment versus dementia, right versus left temporal lobe epilepsy) or neurophysiological indices of normal and disordered network activity (i.e., functional magnetic resonance imaging, intracranial electroencephalography, magnetoencephalography) to better delineate the diagnostic utility of this tool.
The Words and Faces tests demonstrated good test–retest reliability in both the healthy controls and patient sample. Similar to other tests of memory, small practice effects were evident over these short intervals. A follow-up study of 1 year test–retest reliability in older adults is currently underway with the goal of establishing reliable change indices for clinical use. Importantly, the Words and Faces tests did not exhibit ceiling effects despite the wide range of ages and abilities in our participant sample. This is crucial for detecting small but meaningful changes, particularly in high functioning patients with significant cognitive reserve, or for identifying true improvements over time.
Finally, the well-understood characteristics of Remember–Know recognition responses can be harnessed to provide embedded measurement of test-taking effort. Preliminary analyses in a separate sample of study participants suggest that response patterns (i.e., stereotyped response selections such as repeatedly alternating between “Maybe Old” and “Maybe New” responses, accuracy performance below chance) and reaction time metrics (i.e., longer reaction times to correct versus incorrect or to “Definitely” versus “Maybe” responses, impulsive reaction times that are too fast for perceptual processing to have taken place) may be helpful for identifying performance anomalies that signal invalid test results.
There are important limitations to this study that should be acknowledged. First, study participants were recruited from a single center and, while many people seeking specialty care at our facility travel from other cities, states, and countries, the regional diversity of our sample is limited. Additional data collection in other national and international regions would be helpful in fully establishing generalizability and the tool may be made available to investigators seeking to study its performance in different groups. Second, the clinical batteries varied across our patient sample, and there is significant variability in the sample sizes for correlations between tests. Most sample sizes were robust, but the correlations with some subtests of the WMS-III (i.e., Faces I & II, Digit Span, Spatial Span) should be viewed cautiously. Third, the focus of this study was to evaluate performance in patients with or without cognitive impairment, irrespective of diagnosis. Future studies are planned to determine whether the Words and Faces tests have utility in identifying specific diagnostic categories. Fourth, the neuropsychological test battery administered to the normative sample included only a single, embedded measure of performance validity which may be insufficient to detect suboptimal effort. Fifth, self-administered measures may be particularly susceptible to error associated with poor comprehension of test instructions or insufficient attention to test stimuli, and this cannot be completely ruled out. To minimize problems with comprehension and limit distractions, the Memory subtest instructions are given via headphones in simple, plain language, and each subtest includes practice items that are repeated until the patient is able to respond to four stimuli correctly. The fact that our sample demonstrated a broad range of cognitive abilities and no participants discontinued due to problems understanding the instructions suggests that these efforts are largely successful. Finally, all participants completed the test in a hospital setting, and the current findings apply only to that context. An additional study to determine whether the test can be reliably completed remotely, such as in the patient’s home, is currently underway.
Conclusion
This computerized Remember–Know recognition memory test includes both Words and Faces memory subtests. Memory subtests show both convergent and divergent validity when compared with traditional neuropsychological measures (although processing speed also contributes to performance on the computerized measures), and show similar effects of education, gender, and race (normative tables provided). The Memory subtests also demonstrate good test–retest reliability over a brief period, and the test structure lends itself to embedded effort metrics. These data provide further evidence that automated memory testing may be a useful addition to clinical neuropsychological assessment.
Supplemental Material
sj-docx-1-asm-10.1177_10731911231195844 – Supplemental material for Validation of Self-Administered Visual and Verbal Episodic Memory Tasks in Healthy Controls and a Clinical Sample
Supplemental material, sj-docx-1-asm-10.1177_10731911231195844 for Validation of Self-Administered Visual and Verbal Episodic Memory Tasks in Healthy Controls and a Clinical Sample by Darlene P. Floden, Olivia Hogue, Abagail F. Postle and Robyn M. Busch in Assessment
Footnotes
Acknowledgements
The authors would like to thank the community members and patients who participated in this study. They would also like to thank Dillon James, Malavika (Pia) Sengupta, Sarah McCormick, Julia Biars, Lisa Ferguson, Jamie Gatesman, Danny Bermudez, Taylor Lasik, and the staff and psychometrists in the Section of Neuropsychology at Cleveland Clinic who assisted with recruitment, clinical coding, and test administration. The authors would like to thank Dr. Jay Alberts, Stacey Clemence, and their team for programming the application for iPad. Finally, we are grateful to the anonymous reviewers for their helpful comments and suggestions.
Declaration of Conflicting Interests
The author(s) declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: D.P.F. and R.M.B. would like to disclose the potential for future distributions from Ceraxis Health, Inc. as inventors.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Primary support for this research was provided by a Cleveland Clinic Neurological Institute Innovations and Discovery Award (D.P.F and R.M.B.). Additional support was provided by the National Institutes of Health (D.P.F. grant nos. 1K23NS091344; 1R61AG069729). Its contents are solely the responsibility of the authors and do not necessarily represent the official views of the NIH.
Supplemental Material
Supplemental material for this article is available online.
