Abstract
The construction of the German Auditory Wordlist Learning Test (AWLT) for the assessment of verbal memory in late-life cognitive decline was guided by psycholinguistic evidence, which indicates that a word’s linguistic characteristics influence its probability of being learned and recalled. The AWLT includes four trials of learning, short and long delayed free recall, and a recognition task. Its words were selected with taking into account their semantic content, orthographic length, frequency in the language, and orthographic neighborhood size (the number of words derived by adding, subtracting, or replacing a single letter at a time). Through this method, it was possible to better control item and test difficulty, improve the similarity between parallel forms, and reduce bias through recall advantages for certain words due to their linguistic characteristics. In two pilot studies with cognitively healthy subjects, the AWLT showed good internal consistency, split-half reliability, and parallel forms reliability and proved able to assess learning, retention, and recognition. Overall, linguistic recall effects were mitigated; however, an advantage for high-frequency words was observed.
Item difficulty is usually a central concern in the construction of cognitive tests. In verbal memory tests using the wordlist paradigm, however, this issue is often neglected or hardly considered. This is surprising given that there is evidence indicating that, among others, variables like a word’s length, frequency in the language, and neighborhood size can influence its probability of being learned, recalled, and recognized.
Word length can be operationalized as the number of a word’s syllables (Jalbert, Neath, Bireta, & Surprenant, 2011) and usually short words are better recalled than long words (e.g., Hulme, Suprenant, Bireta, Stuart, & Neath, 2004; Jalbert et al., 2011). Word frequency indicates how common a word is in a given language (Criss, Aue, & Smith, 2011). High-frequency words are often better recalled, while low-frequency words are better recognized (e.g., MacLeod & Kampe, 1996). Neighborhood size describes the number of words that can be created by replacing, adding, or deleting a single letter of a word at a time. Large neighborhood sizes are usually related to a higher recall probability (e.g., Jalbert et al., 2011; Roodenrys, Hulme, Lethbridge, Hinton, & Nimmo, 2002).
These findings have three important implications for the construction of verbal memory tests using the wordlist learning paradigm. First, linguistic characteristics could be employed to control the difficulty of such tests on the item level. This was done in the construction of the second version of the California Verbal Learning Test (CVLT-II; Delis, Kramer, Kaplan, & Ober, 2000), for which only high-frequency words were selected to render the list easier to learn and recall. Second, matching wordlists with regard to their linguistic profiles allows for increasing the similarity between parallel forms and between learning and distraction lists in the recognition task. Third, there is evidence indicating an interaction between the linguistic profile of wordlists and the cognitive status of the test takers. For example, the recall advantage for high- over low-frequency words was reduced in subjects sedated with lorazepam and alcohol compared with controls (Soo-ampon, Wongwitdecha, Plasen, Hindmarch, & Boyle, 2004). Furthermore, the German version of the CVLT (Niemann, Sturm, Thöne-Otto, & Willmes, 2008) was found to be subject to linguistic recall effects (Hessler, Fischer, & Jahn, 2016). While both controls and persons with Alzheimer’s dementia recalled more high- than low-frequency words, there was a mirrored pattern with regard to neighborhood size. The controls had significantly higher recall rates for words with large neighborhood sizes, while the opposite was true for persons with dementia. These interactions might bias the diagnostic accuracy of a wordlist learning test. A list containing many words with small neighborhood sizes, for example, could support the memory performance of persons with dementia and, thereby, obscure group differences and lead to reduced diagnostic accuracy of the test. A possible way to mitigate these effects would be to reduce the variability in the linguistic factors between the words in order to create a list that comprises highly similar words.
Controlling the linguistic properties of the selected words also allows for increasing equivalence in the translation of existing tests. While, for example, there is only little information about how the wordlist within the neuropsychological test battery of the Consortium to Establish a Registry for Alzheimer’s Disease (CERAD; Morris, Heyman, Mohs, & Hughes, 1989; CERAD-WL) was compiled, its Korean translation is linguistically closely matched (Lee et al., 2002). The Korean CERAD-WL resembles the original version with regard to the relative word frequency, semantic category, and partly in word length. For the German version, the original words were directly translated (Thalmann & Monsch, 1997).
The present study describes the construction of the German Auditory Wordlist Learning Test (AWLT; German: Auditiver Wortlisten Lerntest) that was based on psycholinguistic evidence. Furthermore, the AWLT’s psychometric qualities and the presence of linguistic recall effects were investigated in two pilot studies with cognitively healthy subjects. The AWLT was developed as a cooperative project by the Schuhfried GmbH (Mödling, Austria) and the Clinical and Experimental Neuropsychology unit of the Department of Psychiatry and Psychotherapy, Klinikum rechts der Isar, Technical University of Munich (Germany). The AWLT will be part of the tablet-based neuropsychological test battery Cognitive Functions Dementia (CFD) for the diagnosis of dementia that is currently being normed and validated by the developers. The battery is the first tablet-based test set within the well-known Vienna Test System and aims at the early identification and differential diagnosis of predominantly neurodegenerative dementia syndromes.
Construction of the AWLT
General Aims of Test Construction
The above described psycholinguistic knowledge was applied to the construction of the AWLT to (a) control the test’s difficulty on the item level (linguistic item difficulty), (b) increase the similarity between parallel forms as well as between learning and distraction lists, and (c) mitigate linguistic recall effects as much as possible. Furthermore, the AWLT was designed to be a valuable alternative to existing wordlist-learning tests in terms of its structure and linguistic profile.
Commonly employed tests of verbal memory in the context of aging, mild cognitive impairment, and dementia are the CVLT and the CERAD-WL. While the CVLT might be too exhausting for some patients, the CERAD-WL might be too easy in certain cases and produce ceiling effects. To construct a valuable alternative for these two established tests, we aimed to place the AWLT between CVLT and CERAD-WL with regard to its length and number of measures. Assuming that longer paradigms with more measures are more demanding for test takers, we aimed to develop the AWLT as an intermediate solution. Of course, the AWLT might as well be employed in the diagnostics of conditions other than dementia.
Structure of the AWLT
The AWLT has two parallel forms, each comprising 12 learning words. The AWLT’s structure consists of four measures (see Table 1 for a comparison with CERAD-WL and CVLT).
Learning Phase: During the four trials of the learning phase, the 12 words are read to the test taker, who is to recall as many words as possible after each trial with the original order being irrelevant.
Short Delayed Free Recall: After an interval of 5 minutes filled with nonverbal and nonmemory-related tests and without being previously warned, the test taker is again to recall as many words as possible with the original order being irrelevant.
Long Delayed Free Recall: After an interval of 20 minutes filled with nonverbal as well as nonmemory-related tests and without being previously warned, the test taker is again asked to recall as many words as possible with the original order being irrelevant.
Recognition: A list of 24 words containing the original 12 and 12 new but semantically matched words is read to the test taker, who has to recognize the learned words among the distractors.
Test Structures of CERAD-WL, AWLT, and CVLT.
Note. CERAD-WL = Consortium to Establish a Registry for Alzheimer’s Disease neuropsychological test battery wordlist; CVLT = California Verbal Learning Test; AWLT = Auditory Wordlist Learning Test.
The AWLT ranges between the CVLT and the CERAD-WL with regard to number of words and number of measures (Table 1). Similar to the CVLT, the words are read to the subject. As with the CERAD-WL, the AWLT does not include a second learning list.
The CERAD-WL requires that the words are read aloud by the test taker, which ensures that the words are perceived as well as encoded and prevents the use of rehearsal strategies. As a consequence, recall scores are assumed to reflect “real” recall performance that is unaffected by the use of strategies. As the AWLT was intended as a purely auditory test, we decided to have the words read to the subject by the examiner or played from an audio file on the tablet. We think that, for two reasons, auditory word presentation does not introduce more difficulties with regard to rehearsal and encoding than visual presentation: (a) Visual presentation cannot completely prevent the use of strategies. For example, it is possible to conceive a story, which develops along the wordlist as it is read and that can later be reconstructed to promote recall. (b) Encoding can also be assessed and distinguished from retrieval with a test employing auditory word presentation. Impaired encoding can be suspected when a specific profile of low performance in learning, recall, and recognition is observed (Delis et al., 1991; Miller, 1956). In the learning phase, encoding deficits are signified by no or very little improvement over the individual trials and recall rates that do not exceed the auditory working memory of 7 ± 2 items. When these deficits co-occur with low performance at delayed recall and recognition trials, the inability to encode the words and move them to long-term memory can be assumed. Also, a comparison of performance between the delayed recall and recognition measures allows for discriminating impairments in retrieval and encoding. While impaired recall and intact recognition point to a retrieval deficit, an impairment of both retrieval and recognition indicates an encoding deficit (Butters, 1985).
Word Selection and List Compilation
As with the test structure, the aim for the word selection was to place the AWLT between CERAD-WL and CVLT with regard to mean word length, frequency, and neighborhood size. Furthermore, we aimed to minimize the linguistic variability of the AWLT as much as possible in order to attenuate or even prevent linguistic recall effects.
The linguistic analyses of the German CVLT’s Wordlist A and the German CERAD-WL, as well as for the word selection for the AWLT, were performed with the dlexDB (Heister et al., 2011). The dlexDB is a German lexical database that is based on the text corpus of the Digital Dictionary of the German Language (Digitales Wörterbuch der deutschen Sprache; Geyken, 2007), which includes 122,816,010 tokens and 2,224,542 types. The sequence AABBCCDD includes 8 tokens (AABBCCDD; i.e., concrete occurrences of a word in the corpus) and 4 types (A, B, C, D; i.e., class of words). Considering a simplified example with a hypothetical corpus that includes only two sentences: (a) “Anna offers the dog a treat” and (b) “The dog eats the treat off Anna’s hand.” This corpus would have 14 tokens (Anna; offers; the; dog; a; treat; the; dog; eats; the; treat; off; Anna’s; hand) and 10 types (a; Anna; Anna’s; dog; eats; hand; off; offers; the; treat) with the tokens counting each single word in the corpus regardless whether it had occurred before or not and the types counting only the first appearance but no further ones. The Digital Dictionary of the German Language corpus comprises prose, newspaper articles, functional texts, and transcribed spoken language from the whole 20th century in equal shares. The dlexDB is free of charge and accessible online (http://www.dlexdb.de/).
Length was defined by the number of syllables. Normalized annotated type frequency, as well as normalized neighborhood size, were determined by means of the dlexDB. Normalization in the dlexDB is a form of standardization. In the case of frequency, normalized values indicate the type frequency per million tokens in the corpus. For neighborhood size, normalized values indicate the number of neighbored types per million types in the corpus. Annotation allows for analyzing orthographically similar words separately according to the different parts of speech they occupy. Through this method, we were able to extract the normalized frequency of the nouns while excluding personal names and other parts of speech. We employed normalized values for frequency and neighborhood size so that the two variables could be analyzed and interpreted on the same scale. In the remainder, “frequency” will denote annotated normalized frequency, and “neighborhood size” will denote normalized neighborhood size.
The CERAD-WL had a mean word length of 1.7 (SD = 0.67, range 1-3), mean frequency of 29.47 (SD = 27.78, range 7.19-14.22), and mean neighborhood size of 13.27 (SD = 7.98, range 2.57-25.69). The Learning List A of the German CVLT had a mean word length of 2.13 (SD = 0.89, range 1-4), mean frequency of 1.88 (SD = 1.83, range 0.02-5.28), and mean neighborhood size of 5.78 (SD = 6.17, range 0.00-8.41). As the two tests are rather opposed with regard to their mean frequency and neighborhood size, the AWLT could potentially be positioned in between. As the mean word lengths were very close to each other, we decided to select only one- or two-syllabled words, which would produce a similar mean but less variability.
The AWLT’s words for the four lists were selected by a five-step procedure:
Predefinition of 12 semantic categories: furniture, food, transportation, clothing, tools, recreation, animals, plants, buildings, musical instruments, kitchen, and daily life.
Creation of a pool of words that are nouns, easily imaginable and concrete, assignable to one of the 12 semantic categories, and mono- or disyllabic.
Analysis of all words in the pool with regard to frequency and neighborhood size with the dlexDB.
Selection of 48 words that have a word frequency and neighborhood size between 5 and 15 to produce a maximal range of 10 (i.e., smaller than in CVLT and CERAD-WL) in frequency and neighborhood size.
Distribution of these words across four lists of 12 words (12 learning words and 12 distractors for each parallel form) in a way that the lists’ mean normalized frequency and normalized neighborhood size would lie around 10 (between CVLT and CERAD-WL) and that each semantic category appears only once on each list (i.e., no semantic overlap).
It was not fully possible to select only words with frequencies and neighborhood sizes between 5 and 15. In some cases, these thresholds had to be crossed in order to fill the four lists with suitable words. The resulting lists, however, met the previously set criteria of including only one- or two-syllabled words that belong to 12 different semantic categories and having a mean frequency and neighborhood size as well as a difference between maximum and minimum around 10. The words were then randomly sorted within the lists and the lists were randomly assigned their position in the test (learning list or distraction list for the recognition trial and Forms 1 or 2).
Pilot Study 1
In the first pilot study, a paper-and-pencil version of the AWLT was administered to cognitively healthy subjects in the testing center of the Schuhfried GmbH to investigate the test’s feasibility and identify potential areas for adjustment and improvement.
Subjects and Procedures
The subjects were recruited by means of newspaper advertisements in the Vienna area. All interested persons were questioned about the presence of psychiatric disorder or neurological disease and the use of neurotropic drugs. The inclusion criteria for participation were age of 16 years or older and no previous testing with a wordlist learning test within the past year. The AWLT was administered by trained staff of the testing center. The retention intervals before short and long delayed recall were filled with nonverbal and nonmemory-related neuropsychological tests.
All persons who volunteered could be included in the study. Thirty-four persons (17 female, 50.0%) were tested with the AWLT’s first parallel form. Their mean age (SD) was 49.56 (15.86) years. Four (11.8%) finished compulsory primary education or middle school, 19 (55.9%) vocational training, 7 (20.6%) higher schools, and 4 (11.8%) university. Thirty-five persons (18 female, 51.4%) were tested with the AWLT’s second parallel form. Their mean age (SD) was 49.74 (15.38). Two (5.7%) finished compulsory primary education or middle school, 22 (62.9%) vocational training, 6 (17.1%) higher schools, and 5 (14.3%) university.
Statistical Analysis, Results, and Discussion
The AWLT’s feasibility was operationalized as the rate of tests that could be fully administered so that all test scores could be calculated. Feasibility was 100% for both forms.
Intrusions (i.e. falsely “recalled” words) were analyzed with regard to their semantic content to identify items that might produce intrusions from the same semantic category. In Form 1, 11 different intrusions were named. While “Pfeil” (arrow) was named twice, “Mühle” (mill), “Ast” (branch), “Bild” (picture), “Dampf” (steam), “Vogel” (bird), “Kugel” (ball, sphere), “Saum” (seam), “Veilchen” (violet), “Blume” (flower), and “Stuhl” (chair) occurred once. In Form 2, 11 different intrusions were named. While “Blume” (flower) occurred five times and “Kork” (cork), “Haus” (house), “Blüten” (blossoms), “Zucker” (sugar), “Puppe” (doll), “Ball” (ball), “Topf” (pot), “Dach” (roof), and “Schlüssel” (key) were named once.
With five entries in 35 persons, the intrusion “Blume” (flower; frequency = 10.37, neighborhood size = 11.13) might be a result of the word “Blüte” (blossom; frequency = 14.73, neighborhood size = 6.42) that was on the learning list of the second form. Presumably, “Blume” is more prototypical and therefore a common intrusion. In order to remove this bias, “Blüte” needed to be replaced by a more suitable alternative.
Adjustment and Final Version of the AWLT
Based on the results of Pilot Study 1, the AWLT’s second parallel form was adjusted. “Blüte” was replaced by “Blume,” which did not affect the overall linguistic profile of the list. Form 1 remained unchanged.
The AWLT’s final version was then examined with regard to linguistic item difficulty and similarity within and between parallel forms.
Statistical Analysis
Several statistical analyses were conducted to assess the AWLT’s linguistic item difficulty in comparison with CVLT and CERAD-WL, as well as the similarity between its parallel forms.
Linguistic Item Difficulty
A 3 × 3 multivariate analysis of variance (MANOVA) with the between-factor list (AWLT Form 1; CERAD-WL; CVLT List A) and the within-factor linguistic (length, frequency, neighborhood size) was employed to compare the tests’ linguistic profiles. Furthermore, using AWLT data from the pilot study, the recall rates for each word at Trials 1 and 4 of the learning phase, as well as at short and long delayed recall, as indicators of retention, were calculated and plotted in a line graph to investigate the presence of primacy and recency effects.
Similarity Within and Between Parallel Forms
The linguistic similarity between the four wordlists of the AWLT was tested with a 4 × 3 MANOVA, with the between-factor list (Form 1 learning; Form 1 distraction; Form 2 learning; Form 2 distraction) and the within-factor linguistic (length; frequency; neighborhood size). It is not possible to confirm a null hypothesis by means of statistical hypothesis testing. A failure to reject the null hypotheses of equal mean values would only suggest a high probability that the lists are linguistically not different. Therefore, interpretations were mainly based on effect sizes of the differences.
Significant main effects of MANOVAs were decomposed with t tests. For post hoc tests, 95% confidence intervals for the mean differences and Cohen’s d as a measure of effect size are given. p Values were adjusted with the Benjamini–Hochberg procedure (Benjamini & Hochberg, 1995). Data analysis was performed with SPSS 23.
Results
Linguistic Item Difficulty and Comparison With CVLT and CERAD-WL
Figure 1 displays mean length, frequency, and neighborhood size for the CERAD-WL and the learning lists of the first forms of AWLT and CVLT. While the word lengths were close to each other, the AWLT’s mean frequency and neighborhood size lay between the other two tests. The dispersion of the linguistic variables was the smallest for the AWLT, except for frequency, which showed similarly little spread in the CVLT.

Means and standard deviations of length, frequency, and neighborhood size for CERAD-WL, and the first form learning lists of AWLT and CVLT.
MANOVA comparing length, frequency, and neighborhood size of the three tests revealed a difference in the combination of the three variables, Wilks’s Λ = 0.55, F(6, 66) = 3.80, p = .003, η2 = 0.26. Tests of between-subject effects revealed significant differences in frequency, F(2, 35) = 10.54, p < .001, η2 = 0.38, and neighborhood size, F(2, 35) = 4.62, p = .017, η2 = 0.21, but not in length, F(2, 35) = 0.74, p = .253, η2 = .08. Post hoc independent t tests indicated that the AWLT’s words had a higher frequency (M = 9.59, SD = 3.27) than the CVLT’s (M = 1.88, SD = 1.89), t(26) = 7.86, p < .001, Cohen’s d = 3.01, and had a higher neighborhood size (M = 10.31, SD = 3.66) compared with the CVLT’s (M = 5.78, SD = 6.37), t(26) = 2.20, p = .037, Cohen’s d = 0.84. The tests did not differ with regard to word length, t(26) = −1.34, p =.192, Cohen’s d = 0.52. Even though the difference was only marginally significant, the effect size indicated that the AWLT’s words had a lower frequency (M = 9.59, SD = 3.27) than the CERAD-WL’s (M = 29.47, SD = 29.29), t(9.19) = −2.14, p = .061, Cohen’s d = 1.00. The degrees of freedom of the t statistic were adjusted since the variances of frequency were not equal for AWLT and CERAD-WL. The words of the AWLT and the CERAD-WL did not differ with regard to length, t(20) = 0.21, p = .838, Cohen’s d = 0.09, and neighborhood size, t(11.82) = −1.10, p = .283, Cohen’s d = 0.47.
Similarity Within and Between Parallel Forms
Learning and distraction lists had the same mean word length for both parallel forms (Table 2). All lists had mean frequencies below 10 and ranges around 10. Two lists had mean neighborhood sizes below 10 and ranges were around 10. Importantly, the aim of linguistic homogeneity within and between the lists was achieved, which increases the similarity between lists and parallel forms. MANOVA revealed no differences between the lists with regard to mean length, frequency, and neighborhood size, Wilks’s Λ = 0.91, F(9, 102.37) = 0.46, p = .901, η2 = 0.03.
Linguistic Profile of the AWLT With Regard to Word Length, Frequency, and Neighborhood Size.
Note. AWLT = Auditory Wordlist Learning Test; M = mean; SD = standard deviation.
Discussion
The AWLT was constructed with its linguistic properties in mind. The individual wordlists are linguistically similar within and between the two parallel forms, which extends their equivalence to the item level. The AWLT’s mean word frequency and neighborhood size lie between the values of CVLT and CERAD-WL, suggesting an intermediate linguistic item difficulty for the AWLT. Also, the test’s length as well as its number of measures and, thereby, its demand on the test taker lies between the two established alternatives CVLT and CERAD-WL. A subsequent pilot study was conducted with the AWLT’s final version.
Pilot Study 2
A second pilot study of the AWLT was conducted to investigate the AWLT’s psychometric qualities, the trajectories of test scores within and between the parallel forms, and the presence of linguistic recall effects. These analyses were based on the test results obtained at the two pilot studies.
Method
Subjects and Procedures
The participants of Pilot Study 2 were a subset of the sample of Pilot Study 1. Subjects who completed the AWLT’s Form 1 at Study 1 were administered the adjusted and final Form 2 at Study 2 several weeks later and those who completed the original Form 2 at Study 1 were administered Form 1 at Study 2. Otherwise the procedures were similar for the two studies.
In the second session, Form 1 was administered to 22 subjects who completed Form 2 in the first pilot. Form 2 was administered to 21 subjects who completed Form 1 in the first session. Due to the adjustments after the first session, only data from the second session will be analyzed for Form 2. As the parallel forms comprise distinct learning lists, no practice effects would be expected so that Form 1 will be analyzed with the data of both session combined. Comparisons between the parallel forms will be conducted with data from the 21 subjects who completed Form 1 and the final Form 2.
Statistical Analysis
Psychometric qualities
A range of test scores was calculated: (a) the sum of correctly recalled words at each trial of the learning phase, as well as short and long delayed recall; (b) the sum score across all trials of the learning phase; and (c) the number of true positives, false positives, true negatives, and false negatives in the recognition measure, as well as an indicator of accuracy [(true positives + true negatives)/24].
Internal consistency was assessed by Cronbach’s alpha and an odd–even split-half method. Correlations between sum scores of odd and even items were calculated for all individual trials of the learning phase, the sum of the scores at the individual learning trials (1 + 3 vs. 2 + 4), and short and long delayed recall and corrected by the Spearman–Brown formula for reduced test length. Parallel-forms reliability was examined by correlating the above described test scores as well as the recognition accuracy that were obtained in the two forms.
Test scores within and between parallel forms
Mean scores on the measures of the two forms were analyzed and compared with a 7 × 2 repeated-measures analysis of variance (RM-ANOVA) with the within-factors time (Trial 1; Trial 2; Trial 3; Trial 4; short delayed recall; long delayed recall; true positives in recognition) and form (Form 1; Form 2).
Linguistic recall effects
Frequency and neighborhood size of the words of the AWLT’s two parallel forms were correlated with their recall (Learning Trials 1 and 4 as well as short and long delayed free recall) and recognition rates using Spearman’s rho.
The influence of length, word frequency, and neighborhood size on recall and recognition performance was investigated according to a previously employed method (Hessler et al., 2016). For that purpose, the learning list of the AWLT’s first form was dichotomized to form groups of words that are relatively low or high with regard to a certain linguistic characteristic. This was done separately for length, frequency, and neighborhood size by using the median of each variable as cutoff. Recall rates for words above and below the median were then compared at Trials 1 and 4 of the learning phase, and at short and long delayed free recall. For each linguistic variable, a 4 × 2 RM-ANOVA with the two within-factors time (Trial 1; Trial 4; short delayed free recall; long delayed free recall) and linguistic (below median; above median) and the interaction between time and linguistic was conducted. Each of the RM-ANOVAs analyzes the same variance for time. Therefore, we corrected the p values for the main effect of time with the Bonferroni method.
A failure to reject the null hypothesis of equal recall rates was desired, as it would indicate a high probability that the linguistic characteristics do not influence recall performance. A significant time effect, however, would reflect the AWLT’s ability to assess learning and retention.
Significant interaction effects in RM-ANOVAs involving linguistic variables were decomposed with t tests. For post hoc tests, 95% confidence intervals for the mean differences and Cohen’s d as measure of effect size are given. p Values were adjusted with the Benjamini–Hochberg procedure (Benjamini & Hochberg, 1995). Data analysis was performed with SPSS 23.
Results
Form 1 was administered to 34 subjects (17 female, 50%) with a mean age of 49.03 (SD = 15.86) and Form 2 was administered to 35 subjects (18 female, 51.4%) with a mean age of 49.74 (SD = 15.38). Table 3 displays the characteristics of the subjects who completed the final version of the AWLT.
Characteristics of the Participants Who Completed the Final Version of the AWLT.
Note. AWLT = Auditory Wordlist Learning Test; M = mean; SD = standard deviation.
Including the 21 subjects who completed Forms 1 and 2.
Psychometric Qualities
Split-half reliabilities, internal consistencies, parallel-forms reliability, and parametric correlations between the subscores are shown in Table 4. The split-half reliability and internal consistency were low at Trial 1 but increased to Trial 4 to acceptable values. The learning sum, a core variable of the AWLT, was highly reliable. At short and long delayed recall the values were also acceptable. The parallel-forms reliability was good, except at Trial 1 and for the recognition accuracy.
Reliability of the AWLT’s Form 1.
Note. AWLT = Auditory Wordlist Learning Test.
Pearson’s r with Spearman–Brown correction for altered test length. bCronbach’s α. cSpearman’s ρ.
p < .05. **p < .01. ***p < .001.
Test Scores Within and Between Parallel Forms
Figure 2 displays the recall rates of the AWLT’s first form at Trials 1 and 4 of the learning phase, as well as short and long delayed free recall. The curves showed the expected pattern. Recall rates increased from Trial 1 to Trial 4 and were lower in the long compared with the short delayed free recall. In Trials 1 and 4, the typical U shaped association between serial position and recall probability was apparent. For the delayed recalls, the recency effect was diminished and recall rates decreased with increasing serial position.

Recall rates for the 12 words of Form 1 at Trials 1 and 4 of the learning phase, as well as short and long delayed free recall.
In general, mean scores of the AWLT in both forms showed the expected pattern (Figure 3). The mean number of recalled words increased from Trial 1 to Trial 4 of the learning phase and decreased at short and long delayed free recalled. Recognition accuracy was very high, as would be expected in a cognitively healthy sample.

Mean scores and standard deviations of Forms 1 and 2 at Trials 1 to 4 of the learning phase, short delayed free recall (SDFR), long delayed free recall (LDFR), as well as true positives (TP) and true negatives (TN) at the recognition task.
Mean scores on the measures did not differ between the two forms, as could be expected from inspecting Figure 3. The degrees of freedom for time were corrected with the Greenhouse–Geisser formula to adjust for the violation of sphericity in the RM-ANOVA model. Mean scores aggregated across form increased from Trial 1 (M = 6.31, SD = 1.31) over Trials 2 (M = 8.55, SD = 1.73) and 3 (M = 9.21, SD = 1.80) to Trial 4 (M = 9.69, SD = 1.57) and decreased at short (M = 7.91, SD = 2.55) and long delayed recall (M = 7.41, SD = 2.74), while the number of true positives was high (M = 11.02, SD = 1.18), F(2.27, 45.43) = 42.37, p < .001, partial η2 = 0.68). There were no differences in scores between the two forms aggregated across time, F(1,20) = 0.17, p = .682, η2 = 0.01, and no differences between the forms over time, F(3.81, 76.29) = 1.18, p = .327, η2 = 0.06.
Linguistic Recall Effects in the AWLT
Spearman’s ρ indicated no statistically significant correlations of the words’ frequency and neighborhood size with their rates of immediate and delayed recall as well their recognition rates of the AWLT’s first form. Frequency was not associated with recall rates at learning Trial 1 (ρ = 0.21, p = .333), learning Trial 4 (ρ = 0.28, p = .178), short delayed recall (ρ = 0.32, p = .123), long delayed recall (ρ = 0.01, p = .961), or recognition rate (ρ = 0.31, p = .136). Similarly, neighborhood size was not related to recall rates at learning Trial 1 (ρ = 0.01, p = .972), learning Trial 4 (ρ = −0.01, p = .981), short delayed recall (ρ = 0.11, p = .616), long delayed recall (ρ = 0.11, p = .594), or recognition rate (ρ = 0.05, p = .836). The results did not differ when considering both forms together or separately.
The following paragraphs report the results of the repeated measures analyses, which were performed with the test data of the AWLT’s first form. The repeated measures refer to the recall scores at Trials 1 and 4 of the learning phase as well short and long delayed recall. That is, measures within one testing session, not between pilot Studies 1 and 2. As the parallel forms were found to be linguistically similar, linguistic recall effects were only examined in the first form. The degrees of freedom for time were corrected with the Greenhouse–Geisser formula due to the violation of sphericity for all RM-ANOVAs. Aggregated across all levels of length, recall rates increased from Trial 1 (M = 53.18, SD = 18.25) to Trial 4 (M = 80.75, SD = 17.70) and decreased again at short delayed (M = 70.14, SD = 22.29) and long delayed free recall (M = 67.86, SD = 23.63), F(2.27, 124.55) = 54.02, p < .001, partial η2 = 0.50. Length had no influence on recall aggregated across time, F(1, 55) < .01, p = .984, partial η2 < 0.01. Recall rates of short and long words differed over time, F(2.36, 129.93) = 4.00, p = .015, partial η2 = 0.07.
Aggregated across all levels of frequency, recall rates increased from Trial 1 (M = 55.06, SD = 16.34) to Trial 4 (M = 81.25, SD = 16.31) and decreased again at short delayed (M = 69.20, SD = 20.65) and long delayed free recall (M = 66.37, SD = 2 3.30), F(2.16, 118.90) = 57.16, p < .001, partial η2 = 0.51. Aggregated across time, high-frequency words (M = 70.46, SD = 18.14) were better recalled than low-frequency words (M = 65.48, SD = 19.89), F(1,55) = 4.84, p = .032, partial η2 = 0.08). Recall rates differed between high- and low-frequency words over time, F(2.65, 145.64) = 4.02, p = .009, partial η2 = 0.07.
Aggregated across all levels of neighborhood, recall rates increased from Trial 1 (M = 55.06, SD = 16.34) to Trial 4 (M = 81.25, SD = 16.31) and decreased again at short delayed (M = 69.20, SD = 20.65) and long delayed free recall (M = 66.37, SD = 23.30), F(2.16, 118.90) = 57.16, p < .001, partial η2 = 0.51. Aggregated across time, neighborhood had no influence on recall rates, F(1, 55) = 0.11, p= .737, partial η2 < 0.01. Recall rates differed between words with small and large neighborhood sizes over time, F(3, 165) = 5.45, p = .001, partial η2 = 0.09.
The results suggest the expected difference in recall rates between the individual measures that demonstrate the AWLT’s ability to measure learning and retention. The significant interaction effects indicate linguistic recall effects, which seemed to vary between the AWLT’s measures. Table 5 displays the decomposed interaction effects of linguistic factors with the time of recall in the AWLT. After adjusting for multiple comparisons with the Benjamini–Hochberg (1995) procedure only the recall advantage for high-frequency of low-frequency words at Trial 4 remained statistically significant with a Cohen’s d of 0.48. A similar advantage for high-frequency words was found at Trial 1 (Cohen’s d = 0.48; however, the association was not statistically significant.
Linguistic Recall Effects in the AWLT. Decomposition of the Significant Interaction Effects of Length, Frequency, And Neighborhood Size With The time of Recall in the AWLT.
Note. SDFR = short delayed free recall; LDFR = long delayed free recall; AWLT = Auditory Wordlist Learning Test; M = mean; SD = standard deviation.
Degrees of freedom = 55.
Statistically significant according to the Benjamini–Hochberg procedure.
Discussion
The AWLT showed good to very good internal consistency and reliability, especially for core variables like the learning sum, as well as short and long delayed recall. Test scores behaved as expected in cognitively healthy persons with an increase in the number of recalled words during the learning phase and a decrease at short and long delayed free recall, as well as good recognition performance. Despite the efforts to reduce linguistic recall effects by choosing words that were linguistically as similar as possible, an advantage for words with high frequency was observed. Given that this effect has already been observed in the German CVLT (Hessler et al., 2016), which has similarly little variability with regard to word frequency, it can be concluded that the word frequency effect is prominent in wordlist learning tests and might not be fully preventable.
General Discussion
The AWLT is a new, reliable test of verbal learning, short- as well as long-term retention, and recognition. Evidence from psycholinguistic studies was used to increase the control over item and test difficulty, and the similarity between parallel forms as well as between learning and distraction lists. Linguistic recall effects that were found in the German CVLT (Hessler et al., 2016) could be reduced but not eliminated and an advantage for high frequency was still present in the AWLT.
With the AWLT, we aimed to balance a sufficient amount of diagnostic information with efficiency for the clinician and acceptability on behalf of the patients. The AWLT will be part of the neuropsychological test battery Cognitive Functions Dementia (CFD) for the early detection of dementia that is run on a tablet PC with connected external loudspeakers. To increase the standardization of word presentation, the AWLT’s words were previously recorded in a studio and can be played to the test taker in the learning phase and the recognition task. In addition, it will be possible to record the answers in order to cross-check and, if necessary correct, the results after test completion. The present study employed pilot data from a paper-and-pencil version. In future studies, the AWLT’s reliability, validity, and susceptibility to linguistic recall effects need to be investigated in its final tablet-based version.
Currently, norming and validation of the test battery are in progress so that norms for the AWLT will be available in 2017. As part of these efforts, established measures of verbal learning like the CVLT and the CERAD-WL are administered to both cognitively healthy persons and persons with mild cognitive impairment and dementia, allowing for the calculation of the AWLT’s concurrent and construct validity, as well as other psychometric qualities based on data from a large sample. For now, the learning slope and recall rates found in the pilot study might serve as preliminary indicators of the AWLT’s validity in assessing verbal memory.
The linguistic approach to word selection likely also benefits the development of new verbal memory tests in other languages. Effects like the preference for high-frequency words appear to occur at basic physiological levels of language processing (Diana & Reder, 2006; Inhoff & Rayner, 1986) and the theoretical background of the AWLT’s construction is based on international studies. Employing similar linguistic construction principles, Italian and English versions of the AWLT are currently in development.
Linguistic Control of Test and Item Difficulty
The AWLT lies between CERAD-WL and CVLT not only with regard to its structure but also with regard to its linguistic profile. Experimental studies suggest that words with high frequency (MacLeod & Kampe, 1996) and large neighborhood sizes (Jalbert et al., 2011; Roodenrys et al., 2002) have higher recall rates. Given that the AWLT’s words are more frequent and have larger neighborhood sizes than the CVLT’s, the AWLT is likely easier than the CVLT not only due to the structure but also on the item level. While the average neighborhood size was similar, the AWLT’s words were less frequent than the CERAD-WL’s, suggesting that the AWLT’s items are more difficult than the CERAD-WL’s. Importantly, these theoretical considerations need to be confirmed by empirical evidence.
The AWLT is less extensive than the CVLT, as it has fewer words, a shorter learning phase, does not include a distraction learning list (“List B” in the CVLT), and does not include cued recall according to semantic categories. As a consequence, the AWLT produces fewer diagnostic variables (e.g., no score for semantic clustering) but likely is more tolerable to patients and might have higher completion rates in persons with cognitive impairment.
Parallel Forms and Reliability
Controlling the linguistic properties of the AWLT’s words also ensures high parallelization between the two test forms and between learning and distraction lists. All four lists in the two forms are similar with regard to their words’ length, frequency, neighborhood size, and semantic group membership. This high level of similarity up to the individual items is unique in the AWLT and increases the test’s diagnostic accuracy in repeated testing and the recognition trial. In addition, mean scores did not differ between the two parallel forms.
Linguistic Recall Effects
The AWLT’s variability within the linguistic variables was smaller compared with CERAD-WL and CVLT. Linguistic homogeneity reduces the likelihood of including words that have high linguistic salience, for example, due to a very high frequency compared with the rest of the list. By reducing the variability, the words are assumed to have similar linguistic recall properties. However, as in the CVLT (Hessler et al., 2016), high-frequency words had a small advantage during the learning phase of the AWLT. The effect was markedly smaller in the AWLT than in the CVLT, possibly due to the higher mean frequency of the words compared with the CVLT. Words with small neighborhoods had a very small advantage at Trial 1 of the AWLT. This effect was reversed in the CVLT, where cognitively healthy subjects better recalled words with large neighborhood sizes, as was also suggested by experimental studies (e.g., Jalbert et al., 2011; Roodenrys et al., 2002). Length had no influence on learning and recall in both tests. Even though small effects, especially pertaining to frequency, were present, the AWLT seemed to be fairly robust against linguistic recall effects. Yet, it remains to be investigated, how linguistic recall effects present in the AWLT in persons with cognitive impairment and whether differences between diagnostic groups exist.
Clinical Implications
The AWLT was primarily developed for the diagnostic use in late-life cognitive decline, but is also suited for assessing verbal memory in other contexts. Its words were chosen with the aim of increasing diagnostic fairness through reducing linguistic variability, as it has been proposed that linguistic memory effects may present differently depending on the test takers’ cognitive status (Balota et al., 2002; Hessler et al., 2016, Soo-ampon et al., 2004; Wilson et al., 1983). The results from the present study suggest that in cognitively healthy subjects, effects like the advantage for high-frequency words, which was prominent in the German CVLT (Hessler et al., 2016), could be reduced but not completely eliminated in the AWLT. This comes as no surprise, as the AWLT showed similarly little variability in frequency between the words as the AWLT. Since there is no clinical data available yet, it remains to be investigated whether the AWLT is actually fairer than other tests of verbal memory. Ideally, these studies would examine the AWLT’s feasibility, validity, and linguistic memory effects in a variability of patient groups, including, for example, aphasia, schizophrenia, and depression. Hypothetically, the AWLT might be better suited for patients that have retrieval difficulties than the CVLT, as the former’s linguistic properties promote encoding and retrieval more than the latter’s.
Conclusions
The AWLT is a reliable test for the assessment of learning and verbal memory. The application of psycholinguistic evidence in its construction allowed for higher control of item difficulty and better parallelization between forms as well as between learning and distraction lists. Even though word frequency affected learning performance, the AWLT seems to be less linguistically loaded.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
