Abstract
Aims and objectives:
Semantic verbal fluency taps into vocabulary knowledge but often elicits poorer performance in bilingual speakers tested in a second language (L2). Few studies have examined verbal fluency tasks that target abstract words and examined test–retest reliability. The present study examined the performance gap between English speakers and Chinese-English speakers on classic semantic, action, and emotional verbal fluency.
Methodology:
English data were collected for six verbal fluency tasks from 41 English speakers and 18 Chinese-English speakers who spoke English as L2. Verbal fluency performance was reassessed one week later to examine test–retest reliability.
Data and analysis:
Verbal productivity, valence and arousal, test–retest reliability, and cross-session consistency in producing similar responses in repeated testing were analyzed. Morphological complexity of verbal responses also was compared because Chinese is morphologically simpler than English and emotional verbal fluency may trigger morphologically complex words.
Findings/conclusions:
Productivity gaps between groups were reliable for all tasks despite practice effects but were smaller for action and emotional verbal fluency. Words produced by Chinese-English speakers were more positive generally and less arousing for specific categories (e.g., “happy”). Both speaker groups were equally likely to produce variable responses with repeated testing, especially for emotional verbal fluency. Words produced by Chinese-English speakers were morphologically simpler than those produced by English speakers for emotional verbal fluency.
Originality:
Past work on semantic verbal fluency has focused on the retrieval of neutral and concrete words. This is the first study to assess the reliability in performance gaps between speakers tested in L1 and speakers tested in L2 on verbal fluency that targets the retrieval of abstract and emotional-laden words.
Significance:
Verbal fluency is one of the most frequently employed tasks in neuropsychological assessments. The findings demonstrate the effects of language proficiency on performance across verbal fluency variants, group differences in valence and arousal, the effects of repeated testing, and the effects of linguistic features on responses produced.
Semantic verbal fluency measures lexical retrieval (Ruff et al., 1997; Shao et al., 2014) and the ability to generate novel, verbal responses (Pastor-Cerezuela et al., 2016). Test administration is straightforward—participants name as many single-word responses as possible for specific semantic category (e.g., “name as many animals as you can”) without repetitions for 60 seconds. Categories targeted by semantic verbal fluency are often common in daily encounters (e.g., animals, clothing, food). Poorer performance has been associated with clinical conditions that negatively impact social, educational, and occupational functioning in children and adults, such as autism (Pastor-Cerezuela et al., 2016; Spek et al., 2009), developmental language impairment (Weckerly et al., 2001), Down syndrome (Nash & Snowling, 2008), and neurological impairments, such as traumatic brain injury (Croker & McDonald, 2005; Henry & Crawford, 2004) and dementia (Henry et al., 2004). Lexical retrieval and verbal knowledge tasks, such as lexical decision (e.g., Haebig et al., 2015) and picture naming (e.g., Gollan et al., 2012; Laine et al., 1997), are effective for assessing lexical processing, but semantic verbal fluency remains one of the most frequently employed neuropsychological measures based on its standardization (e.g., Troyer et al., 1998), time effectiveness, and relative cultural neutrality. Semantic verbal fluency has been normed primarily for speakers who acquire the testing language beginning at birth and use it predominantly in their day-to-day life (e.g., Ardila, 2020; Tallberg et al., 2008). The task often elicits fewer correct responses in bilingual speakers who acquire the testing language as a second language (L2) after early childhood (Kisser et al., 2012; Portocarrero et al., 2007). Underperformance associated with the mismatch between L2 and the testing language risks misidentifications of disorders (Bedore & Peña, 2008; Kisser et al., 2012; Portocarrero et al., 2007). Verbal experience is language-specific so that dual-language experience in bilingual speakers often leads to a smaller language-specific vocabulary (Bedore & Peña, 2008; de Villiers, 2015; Pearson et al., 1993; for a review, see Bialystok, 2009). Poorer performance also is found in highly proficient bilingual speakers on semantic verbal fluency tasks (Gollan et al., 2002; Rosselli et al., 2002; Sandoval et al., 2010) due to elevated between-language interference associated with the development of dual-language proficiency (Gollan et al., 2002; Sandoval et al., 2010).
The Revised Hierarchical Model (RHM; Kroll & Stewart, 1994) is relevant to understanding poorer semantic verbal fluency performance in L2. According to the model, the bilingual mental lexicon comprises two lexical systems in L1 and L2, and a shared conceptual system linked to the lexical systems. Language-specific proficiency plays an important role in determining the strength of interconnectivity among the lexical-conceptual systems (Kastenbaum et al., 2019; Kroll & Stewart, 1994). The conceptual system is strongly connected to the L1 lexical system due to early exposures and frequent usages, which lead to a “privileged access” to meaning in L1 (Kroll et al., 2010). The access to meaning in L2 acquired after early childhood is less direct and might require mediation via translation equivalents stored in the L1 lexical system. Semantic verbal fluency in L2 may require speakers to adopt an indirect route—the activation of a target concept may not trigger the linked nodes in the L2 lexical system effectively and may require preliminary access to the L1 lexical system. Poorer semantic verbal fluency performance in L2 may result from the needs to adopt an additional access route via L1 besides the mastery of a smaller L2 vocabulary relative to L1 vocabulary (Kastenbaum et al., 2019).
Studies of semantic verbal fluency have focused on concrete non-emotional nouns (e.g., animals, foods). Few have administered action verbal fluency tasks that target the retrieval of verbs (Piatt et al., 1999), which are more abstract and less likely to elicit poorer verbal fluency performance in speakers tested in L2 due to weakened cross-language interference in bilingual activation (Portocarrero et al., 2007). There is little study of emotional verbal fluency that targets the retrieval of emotional words (Wauters & Marquardt, 2018), which are comparable in abstractness to verbs (Lam & Marquardt, 2020) and are highly personalized (Gawda & Szepietowska, 2013). The retrieval of nouns and verbs is supported by different neurological structures; the frontal lobe is more critical for action naming and the temporal lobe for object naming (Cappa et al., 1998; Woods et al., 2005). Compared with the retrieval of neutral words, emotional words may recruit a more diverse network of brain activation (Abeare et al., 2017), including the frontal region and cingulate cortex for positive emotional words, and parietal and occipital region for negative emotional words (Gawda et al., 2017). Studying emotional word retrieval is useful for theoretical and clinical reasons with functional significance. Imagine a bilingual worker sharing her thoughts about her challenging working conditions, yet emotional words in her L2 fail to describe the difficulties that are frustrating to her. Or an immigrant grandparent who fails to use emotional words in L2 to explain adequately how proud she is for her grandchildren at their graduation ceremony. Learning vocabulary that evokes emotions (e.g., friends, achievement, loss, anxiety) is critical to developing social-emotional competence in L2. Communicating emotional experience is central to survival in that positive emotions are associated with safety and desires, and negative emotions signal threats in the environment (Schrauf & Sanchez, 2004).
Differences in target responses between emotional and non-emotional verbal fluency have implications for understanding the influence of language proficiency on lexical retrieval and its assessment. A previous study (Wauters & Marquardt, 2018) suggested that retrieval of emotion-laden words (e.g., friends, money, death, funeral) may be less dependent on language proficiency when compared with non-emotional words. Twenty-one Spanish-English speakers reported self-rated language proficiency in L1 and L2, and completed emotional verbal fluency (joy, anger) and non-emotional verbal fluency (animals, foods) in Wauters and Marquardt (2018). Correlational analyses showed that L1–L2 proficiency difference scores correlated with L1–L2 productivity differences in non-emotional verbal fluency but not emotional verbal fluency. If the retrieval of emotional-laden words is less dependent on language proficiency as suggested by the findings, emotional verbal fluency should elicit a narrowed productivity gap between speakers tested in L1 and those tested in L2.
Measuring emotional verbal fluency performance should acknowledge an important distinction between emotion-laden words that evoke emotions and emotion labels that label affective states (e.g., happy, afraid, scared). Processing studies show that emotion-laden words and emotion labels elicit distinct responses at the behavioral (e.g., Altarriba, 2006; Altarriba & Basnight-Brown, 2011; Kazanas & Altarriba, 2015, 2016a, 2016b; Knickerbocker & Altarriba, 2013) and neurophysiological levels (e.g., Wang et al., 2019; Wu et al., 2019, 2020; Wu & Zhang, 2019, 2020; Zhang et al., 2017; Zhang Teo, & Wu, 2019; Zhang, Wu, Yuan, & Meng, 2018, 2019). Increasing evidence suggests stronger emotional reactivity toward emotion labels than emotion-laden words (e.g., Kazanas & Altarriba, 2015; Zhang et al., 2017). However, methodologies were mixed in previous emotional verbal fluency research (e.g., Sass et al., 2013; Wauters & Marquardt, 2018), with some studies allowing participants to produce both emotional-laden words and emotion labels in the same trial (e.g., Abeare et al., 2017).
Emotional verbal fluency in Lam and Marquardt (2020) instructed participants to retrieve only emotion-laden words (e.g., “things/events that make people feel happy”). Their study used large-scale corpora to establish the criterion validity of the task and showed that words elicited by emotional verbal fluency were from two ends of a valence continuum and were more arousing than words elicited by non-emotional verbal fluency. Emotional verbal fluency has the potential to provide an additional measure beyond verbal productivity for evaluating performance differences between speakers tested in L1 and those tested in L2, such as valence and arousal of the verbal responses. Accumulating evidence obtained predominantly from the receptive modality (e.g., Altarriba & Basnight-Brown, 2011; Eilola & Havelka, 2011; El-Dakhs & Altarriba, 2019; García-Palacios et al., 2018; Kazanas & Altarriba, 2016a; Lam et al., 2021; Zhang, Wu, Yuan, & Meng, 2018; Zhang Teo, & Wu, 2019) suggests that there are significant effects of language dominance on emotional processing (Dewaele, 2004; for a review, see Caldwell-Harris, 2014, 2015). Some studies have shown weakened emotional reactivity in speakers tested in L2 relative to speakers tested in L1 (e.g., Eilola & Havelka, 2011; Lam et al., 2021); no information is available about the valence and arousal of verbal responses produced by L1 and L2 speakers in emotional verbal fluency tasks.
Underexamined areas in verbal fluency: reliability, consistency, and morphological complexity
Performance reliability has received relatively little attention in verbal fluency. It is unclear whether repeated testing moderates the performance gap in different variants of semantic verbal fluency. Cooper and colleagues (2004) warned that test–retest may lead to false-negative results for individuals affected by clinical conditions because of test score inflation. They found that healthy older adult speakers produced approximately three more names for the category of “animals” in repeated testing taken within a 1-week period. Test–retest reliability data are not available for variants of semantic verbal fluency for speakers tested in L1 and speakers tested in L2. For speakers tested in L1, changes in verbal productivity across sessions may result from task familiarity that moderates the ease of lexical search and retrieval. For speakers tested in L2, verbal productivity is limited by task familiarity but also language proficiency, which may be the more important determining factor in performance and limits the effects of repeated testing.
There is limited information on the extent to which participants produce identical responses across sessions also. Cross-session consistency in response produced may be influenced by language status and task properties. A smaller L2 vocabulary may limit word choices and lead to a higher cross-session consistency for speakers tested in L2 compared with those tested in the L1. Categories that have a limited set size also may elicit a higher cross-session consistency than categories that have many word choices. Compared with non-emotional semantic verbal fluency, emotion-laden words targeted by emotional verbal fluency are more highly personalized (Gawda & Szepietowska, 2013) with significant variations expected within the same individual. For instance, a participant may produce “friends” as one of the responses to the prompt “things/events that make people happy” during the first session and yet produce the same response to the prompt “things/events that make people sad” during the second session because a friend just passed away.
Few studies have examined the morphological complexity of verbal fluency responses because classic semantic verbal fluency (e.g., “animals”) and action verbal fluency (e.g., “things people do”) target basic nouns (e.g., dog, cat) and verbs (e.g., run, eat). A recent study (Lam & Marquardt, 2020) showed that words elicited by emotional verbal fluency are more morphologically complex than words elicited by classic semantic and action verbal fluency. While classic and action verbal fluency may elicit early acquired inflectional morphemes occasionally, such as “-s” (e.g., dog → dogs) and “-ing” (e.g., run → running), words elicited by emotional verbal fluency may involve complex morphological transformations acquired later in life. For example, “homelessness,” which is derived from the word stem “home,” involves two derivational morphemes “-less” and “-ness,” and “relationships” involves derivational morphemes “-ion” and “-ship,” and an inflectional morpheme “-s.” There are significant differences in morphological productivity across languages. Derivational and inflectional morphemes are productive in English but rare in Chinese where over 70% of words are compound (e.g.,
Aims, research questions, and predictions
The study examined English performance on three variants of semantic verbal fluency (i.e., classic, action, and emotion) in speakers who spoke English as L1 and Chinese-English speakers who spoke English as L2. Analyses aimed to assess the stability of the productivity gap between the two speaker groups in repeated testing, and investigated lexical properties (valence and arousal), cross-session consistency, and morphological complexity of the verbal responses, which may differ between emotional and non-emotional verbal fluency. Emotional verbal fluency was limited to eliciting emotion-laden words because they are highly functional and common in daily life; naming lists of emotion labels is relatively uncommon in daily activities. This study targeted “happy” and “sad”—they are two contrastive and fundamental emotions shared across cultures and developed early in life. For the speaker group with English as L2, the study recruited Chinese-English speakers for two reasons. First, cross-language cognates (i.e., words that share phonological/orthographical forms across languages) may influence vocabulary development (Sheng et al., 2016) and performance on semantic verbal fluency (Kamat et al., 2012). Although cognates exist between Chinese and English (e.g., Zhang, Wu, Zhou, & Meng, 2018), their relative rarity compared with cognates in other language pairs (e.g., Spanish and English) may provide a clearer picture of the effect of language proficiency in the testing language on variants of semantic verbal fluency. Second, inflectional and derivational word formation rules are rare in Chinese but highly productive in English. Comparing the morphological complexity of verbal responses will illustrate whether additional measures may reveal further performance differences between speakers tested in L1 and those tested in L2.
The study addressed three questions. The first question addressed the generalizability and stability of performance gaps across variants of semantic verbal fluency. A smaller verbal productivity gap in action verbal fluency that targeted abstract words (Portocarrero et al., 2007) and a weakened relationship between language proficiency and emotional verbal fluency (Wauters & Marquardt, 2018) predict smaller group differences in number of words produced in emotional verbal fluency. Since practice effects from repeated testing may be limited by poorer L2 proficiency, repeated testing may widen the performance gap between speaker groups during the second testing. Predictions about valence and arousal were less straightforward due to the lack of verbal fluency studies examining lexical properties. However, weakened emotional reactivity in speakers tested in L2 found in processing studies (e.g., Eilola & Havelka, 2011; Lam et al., 2021) predicts a higher likelihood for Chinese-English speakers to produce English words with lowered arousal than English speakers.
The second question addressed cross-session consistency in producing similar responses in repeated testing. Reduced knowledge of English words in Chinese-English speakers predicted a higher likelihood of similar responses in two sessions and thus a higher cross-session consistency. The final question addressed whether words produced by English speakers and Chinese-English speakers differ in morphological complexity. Speakers who spoke Chinese as the L1 were predicted to produce morphologically simpler English words than English speakers. Since emotion-laden words targeted by emotional verbal fluency are more likely to elicit inflectional and derivational morphemes than classic semantic and action verbal fluency, cross-linguistic influences on morphological productivity were expected to be more robust for emotional verbal fluency.
Method
Participants
Fifty-nine adult participants, aged 18–35, were recruited from a major-sized public university (>40, 000 enrolled students) in the United States. Included were 41 English speakers (31 female) and 18 Chinese-English speakers (14 female) who spoke Chinese as L1 (16 Mandarin, 2 Cantonese). The English speakers were born in the United States; all Chinese-English speakers were born abroad and reported English as L2. The two language groups had a comparable composition of male/female participants. Participants provided written informed consent and received monetary compensation for their participation. Participants completed a questionnaire that asked about maternal education. None of the participants reported a history of hearing, language, or neurological disorders; they all had normal or corrected-to-normal vision. Procedures were approved by the Institutional Review Board of the university.
Measures
Language questionnaire
Participants were administered the Language Experience and Proficiency Questionnaire (LEAP-Q; Marian et al., 2007). The questionnaire provides information on the age of acquisition of English (AoA), number of years residing in English-speaking countries, percentage of language usage (%usage), and self-rated speaking, listening, and reading proficiency for each language on a 10-point scale.
Lexical decision task
English language proficiency was assessed with the Lexical Test for Advanced Learners of English (LexTALE; Lemhöfer & Broersma, 2012). LexTALE is devised to measure English lexical knowledge and correlates with general English proficiency. LexTALE includes 40 real words and 20 nonwords. During the computerized test, participants were shown a series of letter strings on a computer screen and responded by pressing a computer key to indicate if the letter strings were English words or not. Lemhöfer and Broersma (2012) indicated the word frequency of some test items is so low that a ceiling effect is highly unlikely.
Working memory measure
The working memory capacity of the participants was assessed with Operation Span, a measure frequently employed to estimate complex working memory span (Conway et al., 2005; Linck et al., 2014) and found to predict verbal fluency performance (Shao et al., 2014). The test is a valid measure of memory capacity, predicts fluid intelligence reliably in non-native speakers (Sanchez et al., 2010), and has been used previously in Chinese-speaking speakers (Bulut et al., 2018). During the computerized test (Unsworth et al., 2005), participants were given a series of simple arithmetic operations (e.g., 2 + 3). For each operation, the participant pressed designated buttons to indicate whether it was true or false, after which they were provided with a letter they had to recall later. For example, in the operation (2 + 3 = 5?, M), participants should respond “true” and memorize the letter “M.” After the completion of a series, participants were prompted with a 4 × 3 matrix of letters and asked to click on the letters they had been given to memorize, in the correct sequential order.
The computerized Operation Span consists of 15 recall sequences, with each sequence ranging from three to seven letters. Memory span (maximum span: 75) was calculated by adding the total number of letters correctly recalled in sequential order. For example, if a participant correctly recalled a sequence of five letters, five points were added to the span; however, if a participant incorrectly recalled even one letter of the sequence, no points were added to the span. The design of the Operation Span resembled the conceptualization of working memory as processing (i.e., arithmetic operation), maintenance (i.e., memorizing the letter), and recall (i.e., clicking on the matrix of letters).
Verbal fluency task
Participants completed six verbal fluency tasks, which represented three variants of semantic verbal fluency: classic verbal fluency (“animals,” “insects”), action verbal fluency (“things people do,” “things people do with hands”), and emotional verbal fluency (e.g., “happy,” “sad”). “Insects” and “things people do with hands” were administered to confirm whether smaller set sizes lead to higher cross-session consistency in verbal responses for selected categories. The task administration procedures for eliciting responses in the six categories were identical. Before each trial, the instructions were displayed on a computer screen and read to the participants by a research assistant. “Please name words from the category “_____________,” as many as you can. All items MUST be one word. You should not repeat, say proper names or places (e.g., John, United States), or say the same word with different endings (e.g., cup, cups). Items that violate these rules will be counted as incorrect.” For example, for the category “happy,” participants read and heard “Please name words from the category “THINGS/EVENTS THAT MAKE PEOPLE FEEL HAPPY.” They then pressed a space bar to proceed, which was followed by a “beep” from the computer and a cross on the computer screen that signaled them to start generating words. The cross on the computer screen disappeared when 60 seconds had elapsed. The order of the six categories was randomly presented and verbal responses were audio recorded. A trained graduate research assistant transcribed the audio recordings and electronically recorded the responses onto an Excel sheet. Verbal fluency performance was reassessed 1 week later to examine test–retest reliability. In the second session, the instructions and administrations of the verbal fluency tasks were the same. The order of the categories was randomized.
Data analysis for verbal productivity
The data analysis included only correct responses. Responses were categorized as errors if they were (1) category errors (e.g., plants for “animals”); (2) repetitions (the same word repeated or produced with different morphemes within the same category); (3) proper names (e.g., Twitter, Facebook); and (4) multi-word responses. The emotional verbal fluency responses were highly personalized (Gawda & Szepietowska, 2013); that is, related to episodes of personal experience and dated events. Only repetitions, saying the same word with different endings (e.g., cup, cups), producing proper names or places (e.g., iPod), and multi-word responses were scored as incorrect for “happy” and “sad.”
Data analysis for valence and arousal
Estimation of valence and arousal of correct verbal responses followed the procedure described in Lam and Marquardt (2020). Each correct verbal response (n = 11002) was matched to the words sampled in the Norm of Valence, Arousal, and Dominance of 13,915 English Words (cf. Warriner et al., 2013). Elicited responses could not be constrained to only words sampled by the corpus and, therefore, require modifications. In this paper, modifications were performed only when (1) the verbal responses could not be found on the corpus without modifications and (2) the modification would not change the word class of the verbal response. Using these criteria, all modifications (1,385 of the 11,002 correct responses, 13%) involved removal of the plural morphemes “-s” from a noun stem. A total of 1,982 responses (i.e., 18.0% of the original data) were excluded from analyses because they were not sampled in the corpus even with modifications. In summary, the analyses on valence and arousal accounted for over 80% of the total correct responses elicited from participants.
Statistical analyses
Analyses were performed with linear-mixed effect modeling (LMER) using version 1.1-21 of the lme4 package (Bates et al., 2015). Simulation study (Schielzeth et al., 2020) highlights the robustness of LMER in handling complex data in repeated measures that may violate distributional assumptions, such as small or unbalanced samples. The first analysis examined total number of correct responses produced in 60 seconds. The analysis started with task (animals, insects, things people do, things people do with hands, happy, sad), language groups, sessions, and their interactions as fixed effects. Working memory (i.e., operation span) and maternal education were included to examine whether they predicted performance. Follow-up analyses were performed on valence and arousal of the verbal responses with task, language groups, sessions, and their interactions as fixed effects.
The second analysis examined cross-session consistency in responses. This dependent variable was operationalized as the percentage of common correct responses shared between two testing sessions. The measure of cross-session consistency was obtained by tallying the number of shared correct responses in Sessions 1 and 2, and the total number of different correct responses. Responses that shared the same word stem but different bound morphemes (e.g., run vs. running, dog vs. dogs, etc.) were treated as the same and not discarded as repetitions if they were produced in different sessions. The analysis of cross-session consistency started with task, language groups, and their interactions as fixed effects, in addition to working memory and maternal education. The total number of different correct responses was included in the models because tasks with smaller set sizes (e.g., “insects”) may yield a higher cross-session consistency. That is, participants are more likely to produce similar responses in separate sessions if selected tasks allow fewer word choices.
The final analysis examined the average number of morphemes per correct response. The analytic approach was the same as for verbal productivity. For all analyses, by-participant intercepts were included in the model as random effects. The models were refined by removing, one at a time, factors that exhibited the highest p value, but retained the hierarchical rule of interactions. Likelihood ratio comparisons were performed to confirm that including a given factor did not improve the amount of variance explained (Baayen et al., 2008).
Results
Participant characteristics
Demographic and language proficiency information for the English speakers and Chinese-English speakers are shown in Table 1. Analysis of variance of self-rated proficiency confirmed a language group effect, F(1, 57) = 103.63, p < .001, η2 = 0.65. English speakers reported higher proficiency for English than Chinese-English speakers (mean difference = 2.48, p < .001). Comparisons of percentage of daily English usage showed the same results and confirmed a language group effect, F(1, 57) = 57.70, p < .001, η2 = 0.50. English speakers used English predominantly and more frequently than Chinese-English speakers (mean difference = 28.4%, p < .001). Chinese-English speakers reported greater proficiency and more daily usage in Chinese than in English, proficiency: F(1, 17) = 17.62, p = .011, η2 = 0.51; daily usage: F(1, 17) = 14.73, p < .001, η2 = 0.46.
Demographic and language proficiency information for English speakers and Chinese-English speakers who spoke Chinese as the first language.
Note. Standard deviations are in parenthesis.
p < .001.
Examination of LexTale performance confirmed a language group effect on lexical decision accuracy, F(1, 57) = 68.96, p < .001, η2 = 0.55. English speakers were more accurate in identifying English words than Chinese-English speakers (mean difference = 18.3%, p < .001). A correlational analysis showed that higher lexical decision accuracy was related to higher self-rated English proficiency, r(57) = .77, p < .001, which supported the reliability of the subjective and objective proficiency measures.
Analyses of non-language measures showed comparable characteristics between English speakers and Chinese-English speakers in working memory (p = .365), maternal education (p = .681), and age (p = .101). To conclude, analyses of participant characteristics showed group differences in language measures only. Specifically, the Chinese-English speakers knew fewer English words, rated themselves as less proficient in English, and used English less on a daily basis than English speakers.
Verbal productivity
Verbal productivity data are presented in Table 2. Model comparisons for verbal productivity indicated a significant two-way Language Group × Task interaction, χ2(5) = 16.37, p = .006, and a significant two-way Task × Session interaction, χ2(5) = 32.66, p < .001. No other effects and interaction were significant, which included working memory (p = .231), maternal education (p = .966), Language Group × Session interaction (p = .302), and the three-way interactions (p = .807). The two-way interactions were analyzed with the lsmeans package 2.30-0 to obtain the least-squares means and to test linear contrasts for linear and generalized mixed models (Lenth, 2016).
Total number of English words in 60 seconds, valence and arousal of verbal responses, and cross-session consistency in responses for three variants of semantic verbal fluency tasks (i.e., classic, action, and emotion).
Note. Only correct responses are reported. Standard deviations are in parenthesis. Eng-L1: English speakers; Chi-L1: Chinese-English Speakers who spoke Chinese as the first language.
Valence and arousal are on a scale of 1–9. (a) Valence: 1 = unpleasant, 9 = pleasant. (b) Arousal: 1 = calm, 9 = excited (cf. Warriner et al., 2013).
Analysis of the Language Group × Task interaction showed that English speakers outperformed Chinese-English speakers on all tasks in between-group comparisons (“animals”: p < .001; “insects”: p < .001; “things people do”: p = .001; “things people do with hands”: p = .005; “happy”: p = .009; “sad”: p = .039). The group differences were task dependent. In comparisons between emotional verbal fluency and non-emotional verbal fluency tasks, group differences for “happy” were smaller than “animals” (β = −3.12, SE = 1.06, p = .003; 95% confidence interval [CI] = [−1.06, −5.19]). Group differences for “sad” were smaller than both “animals” (β = −3.73, SE = 1.06, p < .001; 95% CI = [−1.67, −5.79]) and “insects” (β = −2.32, SE = 1.06, p = .029; 95% CI = [−0.26, −4.39]) (see Table 3). “Happy” and “sad” did not differ from “things people do” or “things people do with hands” (range of p values = .208–.831). Within the non-emotional variants, group differences for “animals” were greater than “things people do” (β = 2.39, SE = 1.06, p = .025; 95% CI = [0.32, 4.45]) and “things people do with hands” (β = 2.90, SE = 1.06, p = .007; 95% CI = [0.83, 4.96]). No differences were found between “insects” and the two action verbal fluency tasks (range of p values = .162–.357). Comparing tasks of different set sizes within the same variant (i.e., “happy” vs. “sad”; “animals” vs. “insects”; “things people do” vs. “things people do with hands”) showed comparable group differences in verbal productivity (“animals” vs. “insects”: p = .187; “things people do” vs. “things people do with hands” p = .633; “happy” vs. “sad”: p = .570). In summary, group differences for “animals” were greater than all tasks except “insects.” Emotional verbal fluency elicited smaller group differences than classic verbal fluency but was comparable to action verbal fluency. Finally, set sizes within the same verbal fluency variant may not moderate performance differences between English speakers and Chinese-English speakers.
Output from best-fit linear-mixed effects models on verbal productivity.
Note. “Animals” and “English speakers” are set as the reference levels. SE: standard error.
p < .001; **p < .01; *p < .05.
The Task × Session interaction was driven by improvement in verbal productivity in “animals” and “things people do” but not the other tasks. Participants produced more words for “animals” (β = 1.44, SE = 0.69, p = .038; 95% CI = [0.10, 2.78]) and “things people do” (β = 4.25, SE = 0.69, p < .001; 95% CI = [2.91, 5.60]) in Session 2 than Session 1. The improvement was greater for “things people do” than “animals” (β = 2.81, SE = 0.98, p = .004; 95% CI = [0.91, 4.71]). Verbal productivity across sessions was stable for “insects” (p = .282), “things people do with hands” (p = .494), “happy” (p = .113), and “sad” (p = .406). The absence of a three-way interaction suggested group differences were stable across sessions and cross-session changes in performance were comparable between English speakers and Chinese-English speakers.
Valence and arousal of correct verbal responses
Valence
LMER indicated a language group effect, χ2(1) = 4.69, p < .03, and a Task × Session interaction, χ2(5) = 16.00, p = .010. No other effects and interaction were significant (range of p values = .143–.747). The group effect indicated words produced by Chinese-English speakers were generally more positive in valence than English speakers (β = 0.09, SE = 0.04, p = .034; 95% CI = [0.01, 0.18]). The Task × Session interaction was driven by “happy,” which elicited words that were more positive during the second session (β = 0.15, SE = 0.08, p = .018; 95% CI = [0.03, 0.27), and “things people do,” which elicited words more neutral during the second session (β = −0.14, SE = 0.05, p = .008; 95% CI = [−0.04, −0.24]). There were no session effects on other tasks (range of p values = .251–.554).
Arousal
LMER indicated a Language Group × Task interaction only, χ2(5) = 11.15, p = .049. All other effects were non-significant (range of p values = .388–.989). The interaction was driven by “happy” and “animals,” with English speakers producing words that were generally more arousing than Chinese-English speakers (“happy”: β = 0.12, SE = 0.06, p = .033; 95% CI = [0.01, 0.24]; “animals”: β = 0.13, SE = 0.05, p = .009; 95% CI = [0.03, 0.22]). There were no group differences in other tasks (range of p values = .108–.845).
Test–retest reliability
Correlational analyses were performed to examine test–retest reliability of verbal productivity. Verbal productivity for the three variants of verbal fluency was the average number of correct responses elicited by the two tasks for the respective variant (classic: “animals” and “insects”; action: “things people do” and “things people do with hands”; emotion: “happy” and “sad”). The test–retest reliability in performance was moderate (See Table 4) and did not differ for the three variants (range of p values: .142–.936).
Test–retest reliability for classic semantic, action, and emotional verbal fluency.
p < .001; *p < .05.
Model comparisons for cross-session consistency in responses indicated a significant Task effect, χ2(5) = 244.73, p < .001, but no effects of language group (p = .313), Language Group × Task interaction (p = .292), or working memory (p = .325). Besides tasks, the effects were significant for maternal education, χ2(1) = 5.16, p = .023, and total number of different responses associated with selected tasks, χ2(1) = 23.54, p < .001. Participants with higher maternal education produced a greater percentage of consistent responses across sessions (β = 0.50, SE = 0.22, p = .027; 95% CI = [.07, 0.93]). Tasks of greater set sizes predicted a smaller consistency in responses across sessions (β = −0.44, SE = 0.09, p < .001; 95% CI = [−0.27, −0.62]). “Insects” elicited a greater consistency than the other tasks (range of p values = .002 to <.001), as expected. Set sizes did not fully account for task differences in cross-session consistency (Table 2). “Animals” elicited greater consistency than action and emotional verbal fluency (all p values < .001). “Things people do” elicited greater consistency than “happy” (β = 9.79%, SE = 2.14, p < .001; 95% CI = [5.61%, 13.94%]) and “sad” (β = 9.73%, SE = 2.45, p = .001; 95% CI = [4.95%, 14.47%]) but was comparable to “things people do with hands” (p = .393). Finally, the two emotional verbal fluency tasks did not differ from each other (p > .90). To conclude, analyses of cross-session consistency yielded two findings. First, cross-session consistency was the highest for “animals” and “insects”; emotional verbal fluency was among the lowest. Second, cross-session consistency was not influenced by language groups but was by maternal education.
Average number of morphemes per correct response
The average number of morphemes per correct response is presented in Table 5. Model comparisons for average number of morphemes per correct response indicated a significant two-way Language Group × Task interaction, χ2(5) = 105.90, p < .001. No other effects or interactions were significant, which included working memory (p = .259), maternal education (p = .428), Language Group × Session interaction (p = .203), Task × Session interaction (p = .216), and the three-way interactions (p = .289). Analysis of the Group × Task interaction showed that English speakers produced more morphemes per response than Chinese-English speakers in “happy” (β = 0.18, SE = 0.06, p = .005; 95% CI = [0.06, 0.30]) and “sad” (β = 0.16, SE = 0.06, p = .011; 95% CI = [0.04, 0.29]) (Table 5). The reverse was found for action verbal fluency tasks that English speakers produced fewer morphemes per response than Chinese-English speakers in “things people do” (β = 0.37, SE = 0.06, p < .001; 95% CI = [0.24, 0.49]) and “things people do with hands” (β = 0.40, SE = 0.06, p < .001; 95% CI = [0.27, 0.52]). No group differences were found for “animals” (p = .908) and “insects” (p = .567) (see Table 6).
Average number of morphemes per correct response produced for three variants of semantic verbal fluency tasks (i.e., classic, action, and emotion).
Note. Only correct responses are reported. Standard deviations are in parenthesis. Eng-L1: English speakers; Chi-L1: Chinese-English speakers who spoke Chinese as the first language.
Output from best-fit linear-mixed effects models on average number of morphemes per correct response.
Note. “Animals” and “English speakers” are set as the reference levels. SE: standard error.
p < .001; **p < .01; *p < .05.
Varying group differences in response morphological complexity in verbal fluency tasks were corroborated in a follow-up analysis, which replaced Language group with lexical decision accuracy in modeling. The analysis showed a two-way Lexical Decision Accuracy × Task interaction, χ2(5) = 106.22, p < .001. Better lexical decision performance, which indicated higher English proficiency, predicted greater morphological complexity for “happy” (β = 0.011, SE = 0.004, p = .015; 95% CI = [0.002, 0.019]) and “sad” (β = 0.010, SE = 0.004, p = .035; 95% CI = [.001, 0.017]), less morphological complexity for “things people do” (β = −0.027, SE = 0.004, p < .001; 95% CI = [−0.019, −0.035]) and “things people do with hands” (β = −0.022, SE = 0.004, p < .001; 95% CI = [−0.014, −0.031]), and no effects on “animals” (p = .289) and “insects” (p = .128).
Within-group analyses showed distinct patterns in morphological complexity in English speakers and Chinese-English speakers (Table 5). In English speakers, action verbal fluency tasks elicited words with fewer morphemes on the emotional verbal fluency tasks (all p values < .001). In Chinese-English speakers, words elicited by action and emotional verbal fluency tasks did not differ in the average number of morphemes (all p values > .90).
In summary, Chinese-English speakers produced words of less complexity than English speakers in emotional verbal fluency tasks. However, Chinese-English speakers unexpectedly produced words with more morphemes than English speakers in the action verbal fluency tasks. A descriptive analysis of the action verbal fluency tasks showed that over 90% of the responses produced by English speakers contained only the word stem without bound morphemes. Responses with bound morpheme were more common among Chinese-English speakers and made up more than 40% of all correct responses. Among the bound morphemes elicited, over 95% produced by Chinese-English speakers were the inflectional morpheme “ing.” Although “ing” was also the most predominant bound morpheme produced by English speakers, which made up 81% of all bound morphemes produced, English speakers also produced words with derivational morphemes that were relatively uncommon in Chinese-English speakers (e.g., socialize, memor
Discussion
This study investigated English performance on the classic semantic, action, and emotional verbal fluency of English speakers and Chinese-English speakers who spoke English as L2. Verbal productivity, valence and arousal of verbal responses, cross-session consistency and response morphological complexity were analyzed. There were four primary findings. First, English speakers outperformed Chinese-English speakers on all verbal fluency tasks across testing sessions, but the productivity gap was smaller for action and emotional verbal fluency. Second, words produced by Chinese-English speakers were more positive in general and less arousing for specific categories (i.e., “happy,” “animals”). Third, both speaker groups were equally likely to produce a different set of responses when they were tested again in the second session. Finally, cross-linguistic differences in word formation rules exerted distinct effects on the types of words English speakers and Chinese-English speakers produced on emotional and non-emotional verbal fluency tasks.
Stability of performance gaps in variants of semantic verbal fluency
Testing semantic verbal fluency in bilingual speakers’ L2 often elicited fewer correct responses (Kisser et al., 2012; Portocarrero et al., 2007). Portocarrero and colleagues (2007) administered the Controlled Word Association Test and showed poorer L2 performance for all semantic verbal fluency tasks (“animals,” “kitchen,” “things people do”). Kisser and colleagues showed English speakers produced 15% more words than speakers who spoke English in L2 for “animals” and “supermarket items” combined. The current study extended these findings by showing stable performance gaps in classic semantic, action, and emotional verbal fluency across testing sessions.
Language-specific proficiency plays an important role in influencing lexical retrieval within the framework of RHM (Kastenbaum et al., 2019; Kroll & Stewart, 1994). The impact of language proficiency on semantic verbal fluency is robust. Chinese-English speakers performed more poorly on LexTale, which indicated poorer English proficiency, and produced fewer correct responses on all semantic verbal fluency tasks. Kastenbaum et al. (2019) showed that Chinese-English speakers also produced fewer correct English responses in classic semantic verbal fluency (“animals,” “food,” and “clothing”) than Spanish-English speakers, who reported greater exposures to English. Poorer semantic verbal fluency performance in English found in Chinese-English speakers may be attributed to weakened conceptual to lexical connections due to a lack of immersion in the testing language, the need for an indirect route to access L2 lexical systems via L1 translation equivalents (Kastenbaum et al., 2019), divided language input due to dual-language experience, and/or cross-language interferences that negatively impacts lexical access in the testing language (Gollan et al., 2002; Rosselli et al., 2000). Nevertheless, differences in productivity gaps between variants of semantic verbal fluency tasks highlight the need to consider the interaction between language proficiency and target words in understanding lexical retrieval.
There is a high likelihood that “animals” will yield a significant productivity gap between speakers tested in L1 and those tested in L2 when compared with other semantic categories. Portocarrero and colleagues showed that group differences for “animals” were greater than “things people do,” which is replicated in this study and extended to emotional verbal fluency tasks. Words targeted by “animals” are more concrete than “things people do,” “happy,” and “sad” (Lam & Marquardt, 2020), which may elevate cross-language interference and lead to a more robust demonstration of poorer performance for bilingual speakers tested in L2 (Portocarrero et al., 2007). However, the concreteness hypothesis is incomplete. “Insects,” which targets concrete, animate entities like “animals” but has a relatively small set size, produced a greater performance gap than “sad” only. This finding suggests that category set size may determine whether classic semantic verbal fluency produces a greater productivity gap than action or emotional verbal fluency. Similarly, Portocarrero and colleagues found that the performance gap associated with action verbal fluency was comparable to “kitchen,” which had a smaller set size than “animals.” An alternative explanation for a widened performance gap for “animals” is that some animal entities are unique to countries with no direct translations in English (Portocarrero et al., 2007). This explanation from a cultural perspective may apply to semantic categories that are influenced by geographical regions, such as “foods” (Peña et al., 2002) and “insects,” which may illustrate why the performance gap was comparable between “animals” and “insects.” Differences in productivity gap between “animals” and emotional verbal fluency also align with Wauters and Marquardt (2018), who found weakened relationship between language proficiency and emotional verbal fluency performance. A potential explanation is that emotion-laden words refer to daily encounters that are highly personalized (Gawda & Szepietowska, 2013) and may be learned less systematically than words targeted by classic semantic verbal fluency task in L2 classrooms (e.g., zoo animals, farm animals). Differences in how emotional and non-emotional words are learned may moderate the relationship between language proficiency and target word retrieval. Verbal fluency performance improves with repeated testing. Older English speakers in Cooper et al. (2004) produced three more animal terms and young adult English speakers in this study produced 1.2 more when tested again in 1 week. Greater improvement in Cooper et al. may relate to poorer verbal productivity for older than younger adults during initial testing (Older adults in Cooper et al.: 20.9; Younger adults in this study: 23.0). This study added that the facilitative effects were task-dependent and apply to “animals” and “things people do” only, with greater effects for “things people do.” It is possible that repeated testing may result in improved task familiarity, faster lexical-semantic activation, and/or more effective inhibition of unrelated responses when tested again in 1 week. The restricted effects to “animals” and “things people do” suggest greater facilitative effects on categories that have more representatives and results from greater functionality of verbs than animal terms. Since high frequency animal terms (e.g., dog, cat, lion) would be exhausted in initial retrieval phases of second testing, improvement in verbal productivity in the second session necessitates the search and retrieval of animal terms that occur much less frequently. Relative to verbs that are more functional in daily usage, the retrieval of infrequent animal entities would be more challenging and limit the facilitative effects of repeated testing for “animals.”
The facilitative effects of repeated testing may be time dependent. No effects were found for emotional verbal fluency in this study or Abeare et al. (2017) when participants were retested in 1 week, although productivity improved for testing repeated in 5 hours (Abeare et al., 2017). The timing effects also apply to action verbal fluency. The two speaker groups produced 20% more words when tested again in a 1-week period in this study; no effects were found when repeated testing was performed in 1 year (Woods et al., 2005). Note that improvement in verbal productivity during the second session did not result from participants’ simply recalling and repeating responses produced in the previous session. This study showed that unique responses account for over 60% of total responses to “animals” and 75% responses to “things people do” in the second session.
This study showed that English words produced by English speakers and Chinese-English speakers in semantic verbal fluency differed in valence and arousal. Words produced by Chinese-English speakers were more positive in general, and the effects were not restricted to emotion-laden words. An explanation is the predominance of positivity biases in verbal behaviors (Dodds et al., 2015; Kuchinke et al., 2005; for a review, see Kauschke et al., 2019) and semantic verbal fluency (Lam & Marquardt, 2020), which suggest a processing and retrieval advantage for positive words that may facilitate performance in L2. Positivity biases in verbal behaviors may root in lexical properties that predict ease of lexical retrieval, such as higher word frequency and earlier age of acquisition (for a review, see Goh et al., 2016). Corpus data (cf. Warriner et al., 2013) and previous study on emotional verbal fluency (Lam & Marquardt, 2020) showed that positive words tend to occur more frequently and are acquired earlier in life.
The study performed ad hoc analysis of correct responses for their word frequency (cf. Brysbaert & New, 2009) and age of acquisition (cf. Kuperman et al., 2012). Analyses of word frequency revealed no group differences except for words elicited by action verbal fluency, in which English speakers produced words that occur more frequently than Chinese-English speakers. Analyses of age of acquisition showed that Chinese-English speakers produced words that were acquired earlier than English speakers in three out of six semantic categories (“animals,” “insects,” and “sad”), and the two groups were comparable in the remaining. More positive valence for words produced by Chinese-English speakers may, therefore, be more likely a result of their biases toward producing early acquired words in half of the semantic categories.
Weakened emotional reactivity has been shown in studies of speakers tested in L2 acquired later in life using Stroop-like paradigm (e.g., Eilola & Havelka, 2011; Lam et al., 2021), presumably due to increased psychological distance with the stimuli (Costa et al., 2017; García-Palacios et al., 2018). This study showed that differences in arousal between speakers tested in L1 or L2 were present also in semantic verbal fluency. Group differences in arousal for emotional verbal fluency may depend on valence—the study showed comparable arousal between groups for words elicited by “sad” but lower arousal for “happy” in Chinese-English speakers. A potential explanation is that emotion-laden words elicited by negative emotions tend to signal dangers and threats (e.g., accidents, death, anxiety), which are relevant to survival and, therefore, are more functionally important to individuals. Potential differences in functionality between positive and negative emotion-laden words might prompt speakers of L1 and those of L2 to produce responses of comparable arousal for “sad.”
Cross-session consistency in responses on semantic verbal fluency
Verbal productivity showed moderate test–retest reliability although participants were highly likely to produce different responses when tested again. Interestingly, higher maternal education predicted higher cross-session consistency. A possible explanation is that parents with more years of education may facilitate children’s awareness of academic assessments, which leads to a higher likelihood for participants with greater maternal education to attempt comparable performance across testing sessions.
Variability was significant on all tasks sampled in this study. Significant variability is expected for emotional verbal fluency. Valenced words associated with selected emotions are highly personalized (Gawda & Szepietowska, 2013) because they are related to episodes of dated events and personal experience. The data showed that variability is not limited to inter-individual comparisons. Instead, responses to emotional verbal fluency are highly variable even for the same participant; over 80% of the responses are unique to a particular session. Among tasks, “insects” elicited the highest cross-session consistency in responses because this category has a much-limited set size. Nonetheless, over 40% of the total responses to “insects” were unique and not repeated in another session. Finally, adding an additional constraint “with hands” to “things people do” did not lead to a higher cross-session consistency. Perhaps there are more word choices for verbs than “insects” even with additional constraints (i.e., with hands), which minimizes differences in cross-session consistency between the two action verbal fluency tasks.
An unexpected finding was that Chinese-English speakers and English speakers are equally likely to produce a different set of responses in repeated testing despite differences in English proficiency. Analogous to the comparison between “insects” and other categories, a limited L2 vocabulary repertoire in Chinese-English speakers should predict a higher likelihood of similar responses when tested again. The data showed that a smaller vocabulary repertoire does not limit speakers’ tendency to produce different responses when tested again.
Potential influences of language features on word production in semantic verbal fluency
The data showed differences in the types of words Chinese-English speakers and English speakers produced in selected verbal fluency tasks. Such language effects on word production may result from differences in word formation rules between English and Chinese. Productivity of selected word formation rules in L1 influences the development of morphological awareness of the corresponding rules in L2 (Lam & Sheng, 2016). The relative complexity of derivational morphemes, their relationship with later literacy acquisition (Kuo & Anderson, 2006), and their absence in Chinese delay the mastery of such rules in English in young Chinese-English children. These rules may be particularly relevant to word retrieval in emotional verbal fluency. Specifically, valenced words elicited by emotional verbal fluency are likely more morphologically complex than action verbal fluency or “animals” (Lam & Marquardt, 2020), which tends to elicit derivational morphemes (e.g., dis- and -ment as in
An unexpected finding was that Chinese-English speakers produced words that were morphologically more complex than English speakers in action verbal fluency. Chinese-English speakers exhibited a much-elevated likelihood to attach the inflectional morpheme “-ing” to verb stem, which accounted for 40% of the total correct responses of Chinese-English speakers. A potential explanation is that inflections are often explicitly taught as grammatical morphemes in L2 classrooms. Performance of Chinese-English speakers might reflect learners’ memorization of the citation form of verbs. Alternatively, developmental studies in English-speaking monolingual children show that inflectional morphemes are acquired early in life (e.g., -s, -ing), with “-ing” and “-s” among the first to be mastered (Brown, 1973). The increased frequency of attaching “-ing” in Chinese-English speakers may result from the conceptual simplicity of the inflectional word formation rule. Findings from action verbal fluency showed that Chinese-English speakers exhibited comparable, if not greater, awareness of inflectional morphemes, compared with English speakers. Similarly, findings from “animals” and “insects” showed that Chinese-English speakers and English speakers were comparable in the awareness of inflectional morpheme “-s.” Despite comparable awareness of the inflectional morphemes, Chinese-English speakers continued to produce less morphologically complex words in emotional verbal fluency, which additionally taxed the usage of derivational word formation rules. This finding aligns with developmental studies, which suggest that derivational morphemes take much longer to master than inflectional morphemes (Berko, 1958).
Implications for practice
The results of this study have implications for verbal fluency, which is one of the most frequently employed language-mediated tasks in neuropsychological assessment (Ardila et al., 2006). First, since “animals” and “things people do” are more likely to be used repeatedly to monitor changes in neurocognitive abilities and clinical conditions, examiners need to be aware of the facilitative effects of repeated testing, although the effects weaken in time (Bartels et al., 2010; Woods et al., 2005). Second, the facilitative effect of repeated testing might be more long-lived for classic semantic and action verbal fluency compared with emotional verbal fluency; prior task exposure exerts limited effects on verbal productivity in emotional verbal fluency at a 1-week interval. Third, the effects of repeated testing were comparable for English speakers and Chinese-English speakers. Repeated testing does not moderate the gap between the two language groups in semantic verbal fluency.
The data also showed that investigation of L2 performance should move beyond verbal productivity. The analyses of cross-session consistency found that participants are highly “divergent” in the responses they produce when tested again. In other words, single testing of verbal fluency falls short in capturing the vocabulary breadth for selected semantic categories of even less proficient L2. The analyses of morphological complexity showed that language features may lead to significant differences in the types of words L1 and L2 speakers produce on action and emotional verbal fluency tasks. Consideration of types of responses may have educational and clinical importance for monitoring verbal fluency performance in L2. For example, facilitation of morphological awareness, especially derivational morphemes, may expand language-specific emotional vocabulary in bilingual speakers who are less proficient in L2 and allow for the expression of emotions with better precision.
Limitations
The study was limited by the lack of data on phonemic (letter) verbal fluency. The performance gaps between speakers tested in L1 and those tested in L2 tend to be smaller in phonemic (letter) compared with semantic verbal fluency (e.g., Kisser et al., 2012; Portocarrero et al., 2007). More importantly, semantic verbal fluency is relatively “language-neutral” compared with phonemic (letter) verbal fluency. Significant cross-linguistic differences exist in writing systems that are not alphabetical, such as Chinese. As a result, semantic verbal fluency may provide a more targeted measure of lexical-semantic and memory retrieval across language populations and allow more meaningful interpretation of the effects of language proficiency on verbal fluency performance. Future studies also may compare retrieval of emotion labels and emotion-laden words in verbal fluency. Previous studies suggest differences in processing (e.g., Altarriba & Basnight-Brown, 2011; Zhang, Wu, Yuan, & Meng, 2018) but less is understood about their differences in the expressive modality.
The current study also was limited by measuring English performance in Chinese-English speakers only. Focusing on English has theoretical and clinical implications because normative studies of semantic verbal fluency and knowledge about retrieval process were based largely on performance of English speakers (e.g., Kisser et al., 2012; Portocarrero et al., 2007). Recruitment of Chinese-English speakers avoided potential influences of cross-language cognates on semantic verbal fluency (Kamat et al., 2012) and allowed examination of whether language features may influence word production for emotional verbal fluency. An ad hoc analysis showed that only 0.2% of correct responses (21 out of 11,002) are Chinese-English cognates (golf, yoga, hamburger, guitar, cartoon, taxis, jams, cookie, salmon, tuna). Virtually all cognates refer to concepts predominant in or originated from the Western culture. Only 5 of these 21 Chinese-English cognate responses were produced by Chinese speakers, which suggests that English-Chinese cognates are relatively less accessible to speakers who speak Chinese as L1. Kamat and colleagues (2012) showed that action verbal fluency benefited less from the cognate effect compared with classic semantic verbal fluency (“animals”) because there may be fewer cognates for verbs than for nouns (Kamat et al., 2012). Compared with neutral words, emotion words may be more disparate across languages due to social-cultural differences (Dewaele & Pavlenko, 2002). As a result, the effects of cognates may be less facilitative to action and emotional verbal fluency compared with classic semantic verbal fluency. This study postulates that general “L2 disadvantages” in variants of semantic verbal fluency would generalize to other language pairs despite the potential influence of cognates. To better understand the interaction between language proficiency, language combinations, and variants of semantic verbal fluency, future studies should expand the sample size by including speakers of different language pairs and cultural background (e.g., Lorette & Dewaele, 2020) to examine bilingual performance using a within-subject design.
Conclusion
Accurate interpretation of verbal fluency assessment in research and clinical settings requires a better understanding of the impact of language-specific proficiency on lexical retrieval framed within well-established bilingual mental lexicon model (e.g., RHM). This study established the reliability of the productivity gap between speakers tested in L1 and those tested in L2 across variants of semantic verbal fluency. Future studies should examine the generalizability of the results to other speaker groups. An effective design would be to compare two L2 speaker groups who speak distinct language pairs (e.g., Chinese-English; Spanish-English).
Footnotes
Data availability
The data that support the findings of this study are available from the corresponding author upon reasonable request.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Language Learning Dissertation Grant Program of Language Learning and the Research Seed Grant of the University of North Texas..
