Abstract
This study examined the effects of the speaker’s face and accent on second language (L2) speech perception. Forty-two Chinese speakers of English immersed in the L2 environment were instructed to perform a cross-modal semantic judgement task. They saw an Asian or Caucasian face on the screen and heard word pairs in L2 in a native English accent or a Chinese accent, and were asked to judge whether the word pairs were related to each other in meaning or not. Results showed that for words presented in the native accent, there was a semantic effect in both reaction time and accuracy, irrespective of the face shown. For words presented in the non-native accent, the RT data showed a semantic effect, whereas the accuracy revealed a reversed semantic effect. The speed-accuracy trade-off suggests a relatively weak semantic effect. These patterns were not modulated by the faces accompanying the word pairs. These results suggest that the cue of accent plays an important role during bilinguals’ speech perception in L2, such that non-native accent hampers speech perception, even when it matches bilinguals’ first language. In contrast, bilinguals do not seem to depend on the social indexical cue of the face when it is not reliable. The present findings hold implications for the Bilingual Model of Lexical Access (BIMOLA) of bilingual speech perception and the monolingual models of social speech perception.
Introduction
For bilinguals, it has been shown that both first and second languages (L1 and L2) are activated in language production and comprehension even in L2 that is distant from L1 with respect to linguistic properties (e.g., Colomé, 2001; Guo & Peng, 2006; Hoshino & Kroll, 2008; Thierry & Wu, 2007). As a result, bilinguals need to select the target language in specific contexts. One interesting question arises: Do various social-indexical (e.g., an interlocutor’s face, race, culturally laden icons) and linguistic cues (e.g., accent or semantic context) affect target language selection in bilingual language processing (Hartsuiker, 2015)? According to the Bilingual Model of Lexical Access (BIMOLA; Grosjean, 1988, 1997), bilingual word perception includes three levels of units: features, phonemes, and words. Global language activation and higher linguistic information (i.e., syntax and semantics) exert top-down activation to words in both languages, which in turn activate phonemes. To be noted, this model does not address nonlinguistic social-indexical cues, such as face, and linguistic cues, such as accents. Recent research, however, has shown that these cues do play crucial roles in bilingual language production, visual word recognition, and speech perception (e.g., Bent & Bradlow, 2003; Imai et al., 2003; Li et al., 2013; Zhang et al., 2013).
Visual cue processing during bilingual language use
One line of research on bilingual language production has examined the modulation of visual social-indexical cues, including faces showing racial identity, such as Asian vs. Caucasian faces (Li et al., 2013; Zhang et al., 2013, Experiment 1), the faces of famous people associated with different languages (e.g., Elvis is associated with English; Tom Boonen is associated with Dutch) (Hartsuiker & Declerck, 2009), the faces of individual interlocutors associated with one language or both languages of a bilingual speaker (Woumans et al., 2015), and culturally laden images (e.g., a typical Chinese mailbox vs. a typical Canadian mailbox) (Jared et al., 2013; Roychoudhuri et al., 2016; Zhang et al., 2013, Experiments 2 & 4). Results from these studies have demonstrated that bilinguals make use of visual social-indexical cues associated with the target language to inhibit cross-linguistic interference (see Hartsuiker, 2015, for a recent review).
Notably, only a few recent studies (Declerck et al., 2018; Grainger et al., 2017; Martin et al., 2016; Molnar et al., 2015; Zhang et al., 2013) have examined the role of social-indexical cues in bilingual language comprehension. For example, Zhang et al. (2013, Experiment 3) reported that Chinese culturally laden icons, such as images of China’s Great Wall and the Temple of Heaven, increased the activation of Chinese lexical structures, prompting Chinese (L1) literal translations during English (L2) word processing. Similarly, Grainger et al. (2017, Experiment 2) tested French-English bilinguals in a lexical decision task, in which national flags preceded the stimuli as cues. Results showed that bilinguals were faster when responding to words in the language (e.g., English) paired with the congruent flag (e.g., the UK flag) than with an incongruent flag (e.g., the French flag). To the best of our knowledge, only one study investigated how social-indexical cues would modulate the activation of two languages during bilingual speech perception. Molnar and colleagues (2015) tested early and late Basque-Spanish bilinguals using an audio-visual lexical decision task. Participants were asked to judge whether the stimuli produced by the interlocutors were words or not. Results showed that early bilinguals made faster responses to words in both L1 and L2 when the interlocutors’ faces matched the languages in which audio stimuli were produced. In contrast, no such effect was found in the late, unbalanced bilinguals. These findings led to the conclusion that in early bilinguals, the social-indexical cue of the interlocutor’s face/interlocutor identity facilitated bilingual lexical access in spoken word recognition. In late, unbalanced bilinguals, no congruence effect was observed, although they responded faster to interlocutors who used one language than to those who used two languages, suggesting that they were somewhat sensitive to the social-indexical cue. As this is the only piece of evidence available to date, it is important to investigate further whether similar patterns would be observed in other tasks and other bilingual populations.
Impact of non-native accent on speech perception
Research on whether non-native accents affect bilingual speech perception has been scarce. In the monolingual literature, several studies demonstrate that a non-native accent hampers speech perception (Barker & Turner, 2015; Bent, 2014; Creel et al., 2016; Trude et al., 2013). For example, Barker and Turner (2015) and Bent (2014) reported that monolingual speakers of English recognised fewer words presented in the non-native accent than those produced in the native accent.
Only several studies (Bent & Bradlow, 2003; Imai et al., 2003; Major et al., 2002; Peng & Wang, 2016) examined the impact of non-native accents on bilingual speech perception by comparing bilinguals’ performance on comprehending native and non-native accented speech. These studies demonstrated that non-native accents generally interfered with bilingual speech comprehension in both languages, except when the non-native accent was mild and matched with the bilinguals’ L1. For example, Major and colleagues (2002) and Peng and Wang (2016) observed that Chinese-English bilinguals exhibited better performance when comprehending English passages read in an American accent than those in a Chinese accent. Extending the findings of monolinguals to the bilingual population, these results suggest that non-native accents cause difficulties for comprehension in bilinguals’ L2 speech. Furthermore, Imai et al. (2003) showed that during auditory word recognition, Spanish-English bilinguals were more accurate when responding to native Spanish speech than to American-accented Spanish speech.
Major et al. (2002) also observed that Spanish-English bilinguals showed higher comprehension when listening to English speakers with a mild Spanish accent than when listening to native speakers. However, Bent and Bradlow (2003) found no significant advantage of the native accent in L2 sentence comprehension when the non-native accent was mild. Considering that only a small number of previous studies examined the role of non-native accents in bilingual L2 speech comprehension and that they focused on the sentence and passage levels, it is less clear whether the disadvantage of non-native accents also exists at the word level during bilingual L2 speech comprehension.
Interaction between face and accent during speech perception
It has been shown that social-indexical cues and linguistic cues tend to be highly associated with each other. For example, the social-indexical cue of an Asian face is often associated with a non-native accent. The manipulation of consistency between faces and accents has been utilised in a few studies, providing a valuable window into exploring the representation and processing of various cues and their interactions during speech perception. In the monolingual literature, some sociolinguistic models, such as the exemplar model (Johnson, 1997, 2006) and the reversed linguistic stereotyping (RLS) hypothesis (Kang & Rubin, 2009; Rubin, 1992), have considered the roles of the speaker’s social-indexical characteristics (e.g., age, gender, race) and their interactions with linguistic traits (e.g., accent) during speech perception. Two models that make contrasting assumptions on the role of perceived speaker identity have been proposed. On the one hand, the exemplar model (Johnson, 1997, 2006) posits that listeners actively use social indexical cues to facilitate speech recognition. Specifically, an exemplar stored in the listener’s long-term memory has established bi-directional links with categories that are important for speech perception. The categories include linguistic representations (e.g., lexical items) and social indexical representations (e.g., age, race). The presentation of the indexical cue of an Asian face would positively prime the recognition of speech in the non-native accent. On the other hand, the RLS hypothesis (Kang & Rubin, 2009; Rubin, 1992) holds that listeners may be biased against non-privileged speakers, which negatively impacts speech comprehension. Thus, social indexical cues associated with non-privileged identities (e.g., a non-native speaker of English) would hinder speech comprehension.
So far, the two models have received empirical support. McGowan (2015) asked Caucasian native speakers of English to listen to mini-lectures presented in a Chinese accent. When the participants listened to passages, they either saw a Chinese face or a Caucasian face. As discussed earlier, the exemplar model predicts that the social indexical cue of a non-native speaker (i.e., the Chinese face) should facilitate speech perception, while RLS predicts that the Chinese face should interfere with speech perception. Results showed that listeners who saw a Chinese face achieved better comprehension performance than those who saw a Caucasian face, providing support for the exemplar model. In other studies, however, support for the RLS model was reported. For example, Kang and Rubin (2009) found that both native and non-native listeners showed divergences between different face conditions (i.e., lower comprehension in the Chinese face condition than the Caucasian face condition), except that when seeing an Asian face, native speakers showed higher scores than non-native counterparts. It is noteworthy that in their study, the non-native participants were from diverse L1 backgrounds. As mentioned previously, the matching between the perceived foreign accent and the listener’s L1 modulates L2 speech perception (e.g., Bent & Bradlow, 2003). To avoid confounding factors associated with different L1 backgrounds, it would be important to examine bilinguals whose L1 matches with the foreign accent.
The present study
To summarise, the roles of the social-indexical cues of faces and the linguistic cues of accents in L2 word comprehension as well as their potential interactions warrant a further investigation. From a methodological point of view, previous studies have examined accuracy only as a result of paper-and-pencil data collection. We believe that research on the effects of non-native accents on bilinguals’ word-level processing would benefit from including not only accuracy rates data but also reaction time (RT) data. Indeed, Kang and Rubin (2009) have pointed out that more demanding measures than paper-and-pencil reports were needed to test current models. Thus, the present study examines whether the social-indexical cue of the face and the linguistic cue of accent associated with L1 would hinder bilingual speakers’ L2 spoken word comprehension using a semantic judgement task. In this task, participants judge whether the pair of words is related to each other in meaning or not. If related word pairs are judged faster and more accurately than unrelated word pairs, there would be a semantic effect, which has been taken as an indicator for retrieval of lexical meaning (e.g., Kojima & Kaga, 2003; Wu & Thierry, 2010).
Moreover, a larger semantic effect is widely believed to reflect easier access to word meaning (e.g., Faust & Lavidor, 2003). Based on the face-language congruency effect reported in the bilingual language production studies (e.g., Hartsuiker & Declerck, 2009; Li et al., 2013; Zhang et al., 2013), we predict that if the Asian face yields slower semantic judgement in English (L2) than the Caucasian face, there would be an interference effect (i.e., longer RT, lower accuracy rates, and/or smaller semantic effect). Based on previous research on accented speech perception (e.g., Barker & Turner, 2015; Bent & Bradlow, 2003; Imai et al., 2003), we hypothesise that if L2 word recognition is hampered by the non-native accent, there would be an interference effect for the Chinese accent, compared with that for the American accent.
If faces and accents interact with each other, the exemplar model of social speech perception predicts that there would be a consistency effect, similar to the visual cue-language congruency effect reported in the bilingual language production literature (e.g., Jared et al., 2013; Li et al., 2013; Roychoudhuri et al., 2016; Zhang et al., 2013). In other words, the semantic effect would be larger when word pairs are presented in the Chinese accent associated with the Asian face than in the Chinese accent with the Caucasian face. Similarly, the semantic effect should be more prominent when word pairs are presented in the American accent accompanied by the Caucasian face than those accompanied by the Asian face. In contrast, the RLS hypothesis predicts that the Chinese face would lead to deteriorated speech perception for both Chinese-accented and American-accented speech (i.e., longer RTs, lower accuracy rates, and a smaller semantic effect), compared with the Caucasian face.
Method
Participants
Fifty Chinese speakers of English studying at a university in the United States participated in the experiment and received an incentive for their participation. All participants self-reported normal or corrected-to-normal vision. Data from eight participants were excluded because their accuracy rates in one or more experimental conditions (see below) were lower than 50%, resulting in data from 42 participants included in the final analyses. The mean age of the participants was 25.67 (SD = 4.50; range = 18–39). As sequential bilinguals, they started learning English as a foreign language in the classroom setting in China at approximately 9.12 (SD = 2.50) years of age, and they had been immersed in the English-dominant environment for 42.69 (SD = 39.84) months at the time of the experiment. All participants took the Test of English as a Foreign Language (TOEFL) and achieved an average score of 91.21 out of 120 (approximately equivalent to 578 on the scale of the paper-based exam; SD = 8.74), indicating that they were relatively proficient in L2. They self-rated their proficiency levels in both languages on a 10-point Likert-type scale (1 = not proficient, 10 = near native) in listening, speaking, reading, and writing. Their averaged L1 rating (M = 9.13; SD = 0.88) was significantly higher than their L2 rating (M = 7.16; SD = 1.82), t(41) = 14.18, p < .001, suggesting that they were more proficient in their L1 than L2.
Materials
The materials consisted of 136 primes (the first word of each word pair), 136 semantically related targets (the second word of each word pair), and 136 semantically unrelated control targets. The Forward Associative Strength values for the semantically related pairs (e.g., baby–child) were retrieved from the University of South Florida Free Association Norms (Nelson et al., 2004), with the mean value of 0.17 (SD = 0.07). The semantically unrelated pairs had zero association values (e.g., baby–south). The semantically related and unrelated targets were matched in frequency, number of letters, number of syllables, orthographic neighbourhood sizes, and phonological neighbourhood sizes, ps > 0.10 (see Table 1). The following criteria were used to select both related and unrelated word pairs: the prime and the target in each word pair did not overlap in phonology; words consisting of more than three syllables were rejected; words with more than one pronunciation (e.g., present) and homophones (different words sharing the same pronunciation, e.g., flour and flower, eye and I, son and sun, very and vary) were not selected; words whose Chinese translations overlap in orthography (e.g., desk-table, 书桌-桌子; look-see,看-看见; quick-fast, 快-快; prison-jail, 监狱-监狱) were not selected to rule out the possible cross-linguistic influence resulting from L1 translation word activation (e.g., Guo et al., 2012; Ma et al., 2017); gerunds (e.g., eating) and plural forms of nouns (e.g., legs and feet) were also excluded.
Means (standard deviations) for the properties of the stimuli in the conditions.
Word frequency, number of letters, number of syllables, orthographic neighbourhood, and phonological neighbourhood were based on the frequency norms of Kučera and Francis (1967) and retrieved from the English Lexicon Project website (Balota et al., 2007; http://elexicon.wustl.edu/). Ortho_N: number of orthographic neighbours; Phono_N: number of phonological neighbours.
A female Caucasian native English speaker with a standard North American accent and a female fluent Chinese speaker of English with a noticeable Chinese accent read the word stimuli and were recorded with a high-quality recorder in a quiet room. The duration of recorded words was longer in the Chinese accent (772 ms) than in the American accent (636 ms), t(407) = 21.12, p < .001. It has been well documented that word duration in the non-native accented speech is longer than in the native accented speech (e.g., Guion et al., 2000; Munro & Derwing, 1995). 1 Two images of a female Caucasian face and a female Asian face from the Fu et al. (2015) study were used as visual stimuli. Eight lists were constructed for counter-balancing purposes, with each list including 17 trials for each of the eight conditions (i.e., Semantically related pairs in the native accent presented with an Asian Face, Semantically related pairs in the native accent presented with a Caucasian Face, Semantically related pairs in Non-native Accent presented with an Asian Face, Semantically related pairs in Non-native Accent presented with a Caucasian Face, Semantically unrelated pairs in Native Accent presented with an Asian Face, Semantically unrelated pairs in Native Accent presented with a Caucasian Face, Semantically unrelated pairs in Non-native Accent presented with an Asian Face, and Semantically unrelated pairs in Non-native Accent presented with a Caucasian Face). Participants were evenly distributed and randomly assigned to the eight counterbalanced lists so that each participant heard each prime word only once.
Procedure
The experiment was carried out in a quiet room. After signing an informed consent form, every participant first completed a language history questionnaire, in which he or she provided one’s self-reported L1 and L2 learning experience and proficiency ratings.
To familiarise the participants with the semantic judgement task, a separate practice session with 24 trials (three trials by eight conditions) was provided prior to the formal experimental session. None of the 48 words used in the practice session was presented in the formal experiment. The procedure was identical to that in the formal experiment (see below). If a participant’s accuracy rate was below 50% in the practice session, he or she was asked to repeat the practice session until the accuracy rates reached 50% or higher.
Next, the participants completed the formal experiment. Each trial started with a fixation (+) in the centre of the screen for 500 ms. Next, an image (Asian face or Caucasian face) appeared, staying on the screen until the end of the trial. Five hundred milliseconds later, the first word was presented binaurally via a pair of high-quality earphones. One thousand milliseconds after the onset of the first word, the second word was presented. The participants were asked to judge whether the pair of English words they heard was related to each other in meaning (e.g., baby-child) or not (e.g., baby-south), by pressing one of two keys on the keyboard (i.e., D or K, respectively). The slide disappeared upon the participant’s response or after 3,000 ms with no response. Finally, a blank screen appeared for 1,000 ms before the next trial began. The participants were provided with instructions both orally from the experimenter before the task and visually on the computer screen at the beginning of the task.
Statistical analyses
Linear mixed-effect modelling was conducted using the lme4 package (Bates et al., 2015) in R (version 3.4.3). At first, a linear mixed-effect model was fit for the RT data, with type (semantically related vs. semantically unrelated), accent (Native vs. Non-native), face (Asian vs. Caucasian), and all their possible interactions as fixed effects. For random effects, we included by-participant and by-item random intercepts. We also fitted by-participant and by-item random slopes for type, accent, and face, as they were within-unit factors. Following the “gold standard” suggested by Barr et al. (2013), a maximal random effects structure was adopted. To obtain results analogous to those from ANOVA analyses, type, accent, and face were coded using mean-centred contrast coding (i.e., semantically unrelated = −0.5, semantically related = 0.5; non-native accent = −0.5, native accent = 0.5; Caucasian face = −0.5, Asian face = 0.5). Specifically, the model was constructed as follows: Model 1 <- lmer (RT ~ type*accent*face + (1 + type + accent + face|Subject) + (1 + type + accent + face|Item), data = data, REML = F). A linear mixed-effect model with the same structure was fit for the accuracy data, as well: Model 2 <- lmer (Accuracy ~ type*accent*face + (1 + type + accent + face|Subject) + (1 + type + accent + face|Item), data = data, REML = F). We then excluded the non-significant effects and fit a final linear mixed-effect model for the RT data, with the main effects of type and accent as fixed effects: Model 3 <- lmer (RT ~ type + accent + (1 + type + accent|Subject) + (1 + type + accent|Item), data = data, REML = F). Likewise, a final linear mixed-effect model was fit for the accuracy data, with type, accent, and type by accent interaction as fixed effects: Model 4 <- lmer (Accuracy ~ type + accent + type: accent + (1 + type + accent|Subject) + (1 + type + accent|Item), data = data, REML = F). 2
Results
For data trimming, not only were RTs shorter than 200 ms and longer than 2,500 ms removed, but also RTs that were 2.5 standard deviations above or below the mean of the participants’ performance were excluded. In total, 7.34% of the data were excluded. Only correct responses were included in the RT analyses. The mean RTs and accuracy rates across eight experimental conditions are presented in Table 2 and Figures 1 and 2.
Mean reaction times (ms) and accuracy rates (%) (standard deviations in parentheses).
RT: reaction time; S + NA: semantically related word pairs pronounced with a native accent accompanied with an Asian face; S + NC: semantically related word pairs pronounced with a native accent accompanied with a Caucasian face; S + NNA: semantically related word pairs pronounced with a non-native accent accompanied with an Asian face; S + NNC: semantically related word pairs pronounced with a non-native accent accompanied with a Caucasian face; S − NA: semantically unrelated word pairs pronounced with a native accent accompanied with an Asian face; S − NC: semantically unrelated word pairs pronounced with a native accent accompanied with a Caucasian face; S − NNA: semantically unrelated word pairs pronounced with a non-native accent accompanied with an Asian face; S − NNC: semantically unrelated word pairs pronounced with a non-native accent accompanied with a Caucasian face.

Mean reaction times (ms) in eight critical conditions.

Mean accuracy rates (%) in eight critical conditions.
As shown in Table 3, results for the RT analyses in the final model revealed a significant main effect of type, t = −11.45, p < .001. The semantically related word pairs were judged more quickly than non-related pairs, showing a semantic effect of 278 ms. In addition, the main effect of accent was significant, t = −6.60, p < .001. The word pairs presented in the native accent were responded to faster than those in the non-native accent, indicating an accent effect of 101 ms.
Mixed-effects model for reaction times.
Only significant effects are presented based on the final model. SE: standard error.
As presented in Table 4, analyses on the accuracy data in the final model showed a significant main effect of type, t = −2.07, p < .05. Also, there was a significant main effect of accent, t = 1.65, p < .05, indicating that the Chinese speakers of English responded to the word pairs in the native accent more accurately than in the non-native accent. There was also a significant two-way interaction between type and accent, t = 4.63, p < .001. Pairwise comparisons using the lsmeans package (Length, 2016) showed that semantically related pairs were judged with lower accuracy than unrelated pairs when the word pairs were produced in the non-native Chinese accent, t = −4.52, p < .001. However, such a pattern was not present when word pairs were produced in the native American accent, t = 0.32, p > .10.
Mixed-effects model for accuracy rates.
Only significant effects are presented based on the final model. SE: standard error.
Discussion
The present study examined the effects of the speaker’s face and accent on L2 speech perception using a semantic judgement task. One major finding is that there was a significant effect of semantic relatedness in the RT data. In line with previous studies (e.g., Thierry & Wu, 2007; Wu & Thierry, 2010), this finding indicates that the unbalanced sequential bilingual speakers can access the meaning of L2 words in speech perception.
The participants responded to the word pairs presented in the native accent more quickly and accurately than to those in the non-native accent, indicating that accent also modulated the semantic judgement of Chinese speakers of English. This pattern is consistent with previous findings that apparent non-native accent resulted in lower performance in bilinguals’ L2 speech comprehension, even when the bilinguals’ L1 was congruent with the non-native accent (e.g., Bent & Bradlow, 2003; Peng & Wang, 2016). Considering that previous studies only examined sentence or passage comprehension in L2, the current results provide the first piece of evidence for the negative impact of non-native accent on bilingual L2 speech comprehension at the word level.
Accent also influenced the semantic effect in the accuracy data. Specifically, when word pairs were presented in the non-native accent, the accuracy rates were lower in the semantically related word pairs than unrelated pairs. When words were presented in the native accent, however, there was no such effect. Together with the RT data (i.e., faster responses to semantically related word pairs than for unrelated pairs), the current findings in the non-native accent condition suggest a speed-accuracy trade-off. For the native accent condition, however, no trade-off was observed. Collectively, these results indicate that non-native accent disrupts bilingual speakers’ speech perception, even when it matches with the bilinguals’ native language.
The current findings have implications for the models of bilingual word comprehension (e.g., Grosjean, 1988, 1997), which currently do not characterise how accent modulates bilingual word comprehension. In the BIMOLA model, there are three levels of representation, including features, phonemes, and words. In addition, global language activation (i.e., language mode and listener’s base language) and higher linguistic information (i.e., syntax and sentence semantics) modulate the activation of words in both languages. Our results have several implications for the BIMOLA model. Above all, the processing of non-native accented words has not been considered in the model. To incorporate the situation of how the bilingual listener comprehends spoken words produced by another bilingual speaker in the model, two factors need to be considered. First, the bilingual listener may not have the same representation of units (e.g., features, phonemes, and words) in L2 as native speakers. For example, the representation of phonemes in the L2 (particularly those which do not exist in the listener’s L1) may be characterised with replacement or assimilation with a corresponding phoneme in the L1. Likewise, the bilingual speaker may speak with a non-native accent of the L2. There are possibilities for bilinguals to replace a certain phoneme in the L2 with a similar one in the L1. In this case, the features and phonemes differ from those produced by native speakers. Although accented speech is not characterised in the BIMOLA model, the model does make assumptions on the role of similarity of words across languages in word recognition of the guest language (i.e., the language used in the form of code-switches and borrowings, as opposed to the main language for communication) in the bilingual mode, where two languages are mixed. Following the same logic and expanding it to the monolingual mode (one language being used) at the phoneme level (and possibly at the word level), the non-native accented speech can slow down the recognition of words in L2. Indeed, our results support such an assumption. Thus, in order to capture the aspects of accented L2 speech, the overlap between the subsystems of phonemes in the two languages should be allowed. This might be modulated by the degree of non-native accent in the speaker and listener. For the listener, if the accent is weak, he/she should be able to distinguish the individual phonemes that could be easily confusable. At the same time, the phoneme system is flexible enough to identify the phoneme with replacement. These issues remain to be further investigated. By accommodating the foreign-accented speech, the BIMOLA model could then explain the modulations of non-native accent during bilingual speech perception.
Critically, contrary to our hypotheses, the social-indexical cue of the speaker’s face did not modulate our bilinguals’ speech perception. Specifically, the face cue did not affect the overall RT or accuracy rates. It did not interact with the accent, as indicated by the absence of face-accent consistency effect. These results showed that when the face was congruent with the accent (i.e., non-native accent associated with an Asian face, or native accent associated with a Caucasian face), meaning access in bilingual L2 spoken word recognition was comparable to meaning access to the L2 lexicon involved when face and accent were inconsistent with each other (i.e., non-native accent associated with a Caucasian face, or native accent associated with an Asian face). The absence of the face-accent congruency effect is consistent with the finding of late, unbalanced bilinguals in the study of Molnar et al. (2015).
Molnar and colleagues (2015) reported that during spoken word recognition, unbalanced bilinguals’ responses to words were not modulated by the interlocutors’ faces they saw (i.e., faces previously associated with their L1, or faces associated with their L2). Similar to Molnar et al.’s finding, the unbalanced bilinguals in the present study did not show slower responses or smaller semantic effects when listening to words accompanied by the Asian face than the Caucasian face. Thus, our data present tentative novel evidence for the lack of interaction between the social-indexical cue of the face and the linguistic cue of accent during bilingual speech perception. Taken together, the current evidence suggests that bilinguals are likely to divert their attention from the social-indexical cue of the face to task-oriented information during L2 speech perception.
Interestingly, the current evidence for bilinguals’ insensitivity to the social-indexical cue of face contrasts with previous results showing that visual L2 word recognition was interfered by the presence of social-indexical cues associated with the L1, such as culturally laden icons (e.g., image of Great Wall of China in Zhang et al., 2013) and its corresponding national flag (Grainger et al., 2017, Experiment 2). Specifically, Grainger et al.’s (2017) results show that the social-indexical information indicating language membership plays a role in bilingual visual word recognition. There was a significant flag-language congruency effect (i.e., longer RTs during flag-language incongruent trials than congruent trials). In the current study, however, no face-accent congruency effect was present. We speculate that the modality of the tasks might have played a crucial role in the divergent findings. In other word, bilinguals may rely more on the visual cue of the face during visual word recognition than speech perception. Thus, the modulating role of the face on L2 word comprehension could not be extended from the visual modality to the auditory modality.
Furthermore, the current evidence is at odds with the previous results of visual cue-language congruency effects observed in bilingual word production (e.g., Hartsuiker & Declerck, 2009; Jared et al., 2013; Li et al., 2013; Roychoudhuri et al., 2016; Woumans et al., 2015; Zhang et al., 2013). Different from the findings of bilingual word production suggesting that bilinguals use social-indexical cues to select the target language, the present data suggest that when the social-indexical cue of the face associated with the non-target language is not reliable, it does not impede the activation of the target language during bilingual speech perception. This is in line with Hartsuiker’s (2015) proposal that nonlinguistic visual cues may play a less important role in language comprehension, as compared with language production. With the main goal of understanding the utterance, the listener is less likely to rely on the social-indexical cues in speech comprehension, whereas the speaker needs to choose the language at hand by exploiting various cues in language production.
Moreover, the absence of the modulation of the face during bilingual speech perception is inconsistent with the findings among monolinguals. McGowan (2015) examined whether the face cue would affect monolingual speech perception presented in the non-native accent. English monolinguals achieved higher accuracy in transcribing English sentences in a non-native (Chinese) accent when a Chinese face was shown, relative to when a Caucasian face was presented. In the current study, however, no sensitivity to the face cue was observed in bilinguals. Based on the assumption that the cue of the face would be processed by the participants, we conjecture that bilinguals exert cognitive control to divert their attention from the unreliable social-indexical cues to task-relevant information in order to achieve optimal performance in speech perception.
The lack of modulation of the face cue in the present study could not be accounted for in the social speech perception models in the monolingual literature. The RLS model (Kang & Rubin, 2009; Lippi-Green, 1997; Rubin, 1992) assumes that monolingual listeners’ negative social bias towards non-native speakers results in reduced attention paid to the acoustic signals and consequently poor performance. Thus, the RLS model predicts that the Asian face would impede speech perception. The current data, however, show that bilingual speakers are not affected by the manipulation of speakers’ perceived ethnicity. According to the exemplar model (e.g., Johnson, 1997, 2006; Munson, 2010; Sumner et al., 2014), listeners actively make use of social cues, which invoke (stereotypical) phonetic expectations, to facilitate their speech perception. Therefore, the exemplar model predicts that seeing an Asian face will enhance the perception of foreign-accented speech. However, our data did not show such a facilitation effect. It is noted that the RLS model and the exemplar model were originally proposed to account for social speech perception in monolinguals. Our findings suggest that these hypotheses may not apply to the bilinguals in our study. Rather, it seems that bilinguals pay an equal amount of attention to the acoustic signals accompanied by the Asian and Caucasian faces. As our bilinguals’ native language matches with the non-native accent, it still awaits further investigations on whether the same patterns would hold true for bilingual listeners whose native languages do not match with the non-native accent in the accented speech.
In conclusion, the present experiment provides the first piece of evidence on the varying roles of the social-indexical cue of the face and the linguistic cue of accent in bilingual L2 word comprehension. It is one of the initial steps to further our understanding of the modulations of various cue processes during bilingual language comprehension. In the future, it would be important to examine the role of familiarity of the non-native accent by testing bilinguals whose native languages do not accord with the non-native accent. In addition, future studies should investigate whether and to what extent the degrees of non-native accents (e.g., L2 speech with heavy- vs. mild-non-native accent), language immersion, the semantic constraints of the sentential context, and training with the non-native accent play significant roles in bilingual word comprehension.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The study was supported by the Faculty Incentive Award from College of Education, Criminal Justice, and Human Services, University of Cincinnati and the Faculty Development Research Grant from the University Research Council, University of Cincinnati awarded to Fengyang Ma, and the National Natural Science Foundation of China (31871097) to Taomei Guo.
