Abstract
Language-specific restrictions on sound sequences in words can lead to automatic perceptual repair of illicit sound sequences. As an example, no Spanish words begin with /s/-consonant sequences ([#sC]), and where necessary (e.g., foreign loanwords) [#sC] is repaired by inserting an initial [e], (e.g. foreign loanwords, cf., esnob, from English snob). As a result, Spanish speakers tend to perceive an illusory [e] before [#sC] sequences. Interestingly, this perceptual illusion is weaker in early Spanish–English bilinguals, whose other language, English, allows [#sC]. The present study explored whether this apparent influence of the English language on Spanish is restricted to early bilinguals, whose early language experience includes a mixture of both languages, or whether later learning of second language (L2) English can also induce a weakening of the first language (L1) perceptual illusion. Two groups of late Spanish–English bilinguals, immersed in Spanish or English, were tested on the same Spanish AX (same–different) discrimination task used in a study by Carlson et al., (2016) and their results compared with the Spanish monolinguals from Carlson et al.’s study. Like early bilinguals, late bilinguals exhibited a reduced impact of perceptual prothesis on discrimination accuracy. Additionally, late bilinguals, particularly in English immersion, were slowest when responding against the Spanish perceptual illusion. Robust L1 perceptual illusions thus appear to be malleable in the face of later L2 learning. It is argued that these results are consonant with the need for late bilinguals to navigate alternative, conflicting representations of the same acoustic material, even in unilingual L1 speech perception tasks.
1 Introduction
Language-specific restrictions on possible sound sequences can lead listeners to perceive illicit sequences as if they had been repaired (Berent, Steriade, Lennertz, & Vaknin, 2007; Davidson, 2011; Dupoux, Kakehi, Hirose, Pallier, & Mehler, 1999; Polivanov, 1931, cited in Lentz & Kager, 2015). For instance, Spanish does not allow word-initial [s]-consonant clusters [#sC], and when necessary such sequences are repaired via prothetic [e] (e.g., English snob became Spanish esnob). Consequently, Spanish speakers tend to perceive an illusory [e] preceding [#sC] clusters (Cuetos, Hallé, Domínguez, & Segui, 2011), thus perceptually “repairing” the illicit cluster.
For bilinguals, however, the language-specific conditions underlying an illusion in one language may differ or be absent in the other language (universal properties of language may be an additional source of perceptual illusion, cf. Berent et al. (2007), but the focus here is on language-specific patterns). This raises interesting questions about how bilinguals’ perception may or may not reflect the contrasting properties of each language encountered in their experience. Interactions between phonetic category structures have been amply documented in perception and production, revealing both first language (L1) influence on second language (L2) categories (Best & Tyler, 2007; Flege, 2003; Kuhl et al., 2008; Pallier, Bosch, & Sebastián-Gallés, 1997) and L2 influence on L1 categories (Caramazza, Yeni-Komshian, & Zurif, 1974; Chang, 2012; Flege & Eefting, 1987; Sancier & Fowler, 1997). Similar crosslinguistic effects occur for phonotactics (Anisfeld, Anisfeld, & Semogas, 1969; Sebastián-Gallés & Bosch, 2002), at least concerning the relative probability or legality of sound sequences. Illusory perceptual repair, however, does not concern phonetic category boundaries or the legality of a sound sequence per se, but rather the perceptual insertion of sounds that are absent from the speech signal, but required based on higher-order phonotactic constraints and phonological repair processes. Perceptual illusions related to phonotactics thus provide a way to examine how bilinguals negotiate conflicting patterns at a more abstract level.
Interestingly, early Spanish–English bilinguals exhibit a weaker illusory [e] effect when perceiving Spanish speech, in proportion to their relative dominance in English (Carlson, Goldrick, Blasingame, & Fink, 2016). These findings are from early bilinguals highly fluent in both languages, who had grown up (and were tested) in an intensely bilingual environment on the USA–Mexico border. While their strong tendency to perceive a brief, ambiguous initial vowel as [e] rather than [a] demonstrates knowledge about the Spanish prohibition against [#sC] and its repair via [e]-prothesis, the strength of this asymmetry was weaker compared to Spanish monolinguals, and they were much less likely to perceive initial [e] when no vocalic material was present. Their speech perception thus appears to be intermediate between Spanish and English, in which [#sC] is plentiful (e.g., speak, snack).
How does such hybrid behavior arise? For early bilinguals, it may be that their substantial early experience with English /#sC/ words, in proportion to the mixture of languages in their early environments, resulted in an enhanced ability to perceive [#sC] veridically, whether it occurs in Spanish or English speech. Alternatively, bilinguals’ speech perception may reflect the interaction of two contrasting phonotactic systems, which provide conflicting ways of representing the same acoustic material that must be sorted out.
The present study seeks to gain some traction on this issue by asking whether the illusion is also weaker for late Spanish–English bilinguals, who learned English after establishing a robust, monolingual Spanish system. This would provide evidence that learning to represent L2 structures that are prohibited in the established L1 system (namely, [#sC]), can allow late bilinguals to perceive the acoustic signal more veridically when those illicit sequences are presented experimentally in the L1. Such changes may also depend on the amount and type of L2 experience, and a larger deviation may be observed during English immersion, where English is more salient (cf. Baus, Costa, & Carreiras, 2013; Linck, Kroll, & Sunderman, 2009). A second goal of this study is therefore to explore the effects of immersion by testing participants who are currently in either an English- or a Spanish-speaking environment.
The L2 effects on the L1 depend, of course, on learning the L2 patterns, and L2 speech sounds, prosody, and phonotactics tend to be hard to learn in general (Dupoux, Pallier, Sebastián-Gallés, & Mehler, 1997; Dupoux, Sebastián-Gallés, Navarrete, & Peperkamp, 2008; Pallier et al., 1997; Sebastián-Gallés & Bosch, 2002). L1 perceptual illusions in particular may suppress changes to the L1 by filtering L2 exposure to such an extent that the contrasting patterns are not perceived robustly. Just such a filtering effect of Spanish prothetic repair on L2 English [#sC] can be found in both speech production, for example, producing sports as [ɛspɔɾts] (Carlisle, 1991; Daland & Norrmann-Vigil, 2015) and perception, for example, [ɛspɔɾts] primes sports (Freeman, Blumenfeld, & Marian, 2016; and a similar finding for Spanish learners of Dutch, Lentz & Kager, 2015). Despite this apparent filtering, Spanish speakers learn to accept L2 (Dutch) phonotactic sequences that are illicit in Spanish, including [#sC] (Trapman & Kager, 2009), suggesting that the filtering is not complete enough to preclude learning the conflicting L2 pattern.
Whether this can influence L1 perceptual repair in late bilinguals remains relatively unexplored, apart from a study of Japanese–Brazilian Portuguese bilinguals (Parlato-Oliveira, Christophe, Hirose, & Dupoux, 2010), whose languages share a prohibition against word-internal CC sequences but repair them with different epenthetic vowels. Monolinguals in these languages perceive their corresponding illusory repair vowel (Dupoux, Parlato, Frota, Hirose, & Peperkamp, 2011), but Parlato-Oliveira et al. (2010) found that early Japanese–Brazilian and simultaneous bilinguals showed evidence of both illusions on an explicit vowel categorization task. Interestingly, (Brazilian) late learners of L2 Japanese also performed differently from Brazilian monolinguals on the same task. Neither early nor late bilinguals were distinguishable from Brazilian monolinguals on a more implicit test, suggesting that changes to perceptual repair routines may be restricted to explicit, metalinguistic tasks. However, the weaker perceptual illusion found by Carlson et al. (2016) for the early Spanish–English bilinguals occurred for both explicit and implicit tasks. Differences in the task or testing environment may have led to the Brazilian repair pattern dominating the implicit task results in the earlier study, or it may be that substituting one illusory vowel for another to repair the same illicit sequence (Brazilian Portuguese vs. Japanese) is not the same as weakening the illusion by learning that the sequence can be licit (Spanish vs. English).
In any event, the learnability of [#sC] in late L2 English and the robust weakening of the illusion for early bilinguals suggest that late Spanish–English bilinguals may also learn to perceive acoustic [#sC] more veridically, even on implicit perception tasks presented in their L1 Spanish. To test this prediction, late bilinguals were tested using the same AX (same–different) discrimination task employed by Carlson et al. (2016).
2 Methods
2.1 Participants
Fifty-five proficient, late Spanish–English bilinguals were tested, 30 in a highly Spanish-dominant environment (Granada, Spain), and 25 in a highly English-dominant environment (Pennsylvania, USA) lacking any large Spanish-speaking community. Participants were paid. They were recruited from university communities, and most were working on degrees either in English, or in fields where they must rely on English extensively in their research. Those tested in the USA were from various countries, primarily in Latin America, but differences in dialect are not expected to affect the results. All had learned English in a classroom setting before any immersion in English. None reported English dominance or high proficiency in a third language. Three additional bilinguals were excluded because they responded incorrectly to more than 25% of the control trials in which the A and X were identical. Due to a software error, one participant only completed half the discrimination task, but the recorded trials were retained for analysis. The Spanish monolingual results from Carlson et al. (2016) 1 were used as a control group.
To assess language background and provide objective measures of proficiency, the late bilinguals completed the Multilingual Naming Test (MINT) picture naming task in both languages (Gollan, Weissberger, Runnqvist, Montoya, & Cera, 2012), cloze tests from the Diplomas of Spanish as a Foreign Language (Ministry of Education, Culture, and Sport of Spain, 2006) and Michigan English Language Institute College English Test (English Language Institute, 2001), and a language history questionnaire (Li, Zhang, Tsai, & Puls, 2014). MINT responses were not recorded for one participant, but this participant had language history and cloze performance consistent with the remaining bilinguals and was retained.
Table 1 compares the bilingual groups, showing comparable age of exposure to English and English proficiency. MINT and cloze scores within the same language were correlated (using Spearman’s rho), English r = 0.59, p < 0.0001; Spanish r = 0.51, p = 0.0001, but L2 and L1 scores were unrelated, all |r| < 0.16 and p > 0.27. This is expected, as achieving high L2 proficiency need not entail negative effects on the L1.
Means (standard deviations) of language and other differences between the bilingual groups. Due to skewed distributions, Wilcoxon rank-sum tests are used for Spanish Multilingual Naming Test (MINT) and cloze scores and years spent in an English-speaking country, and medians are reported for these measures.
As can be expected, the bilinguals in the USA had studied or resided in an English-speaking country longer than those in Spain, although 17 of the latter group reported at least 3 months in an English-speaking environment, with 4 reporting over a year. The number of years spent in an English-speaking country correlated significantly with both MINT scores, English r = 0.33, p = 0.02; Spanish r = -0.49, p = 0.0002, but not with cloze scores, both |r| < 0.2, p > 0.1. The positive correlation between immersion and English MINT is unsurprising, but the negative relationship with Spanish MINT is noteworthy. Similarly, while the bilingual groups did not differ significantly in English proficiency, the bilinguals tested in the USA scored significantly lower in their L1 Spanish, suggesting that current or lengthy L2 immersion may negatively impact L1 performance on these tasks (cf. Baus et al., 2013; Linck et al., 2009). 2 This may be related to group differences in perceptual repair, as will be discussed below.
2.2 Materials and task
The speech perception task was the AX discrimination task used by Carlson et al. (2016) and it was configured to emphasize discrimination based on fine acoustic detail (see below). This was done because the perceptual illusion, though motivated by a phonological repair process (Dupoux, Pallier, Kakehi, & Mehler, 2001), has well-documented effects on listeners’ ability to detect acoustic detail (Dupoux et al., 2011). If late bilinguals are able to detect that detail more reliably than monolinguals, this would indicate that learning English is able to reduce the effects of prothetic repair on speech perception.
The stimuli were Spanish pseudowords of the form VsCid, with final stress. The V was either [e], the default repair vowel, or [a], which occurs in this position but is not used to repair [#sC]. The C was one of the 10 consonants that occur in this position in Spanish, [b, d, g, p, t, k, f, m, n, l]. Stimuli were recorded by a native speaker of Mexican Spanish, and care was taken to ensure Spanish pronunciation (e.g., vowel quality and final /d/ spirantization).
Two versions of each stimulus were created, differing in initial vowel duration. The long version retained the final 10 periods of initial vowel phonation (about 40ms, just under half the original vowel), and the short version retained the final 2.5 periods, which is short enough to preclude reliable perception of formant structure. This resulted in 10 quartets of stimuli sharing the same post-[s] consonant and differing in initial vowel quality ([a] or [e]) and duration (short or long).
Each trial included two stimuli from the same quartet. The 120 “different” trials consisted of one presentation of every possible pair from each quartet (e.g., long-estid/short-estid, long-estid/short-astid, etc., that is, 6 possible pairs per quartet × 10 quartets, × 2 presentation orders). For the 120 “same” trials, each stimulus was paired with itself, with each of the 40 pairs (10 quartets × 4 stimuli per quartet) appearing three times. Thus, 50% of the 240 trials were acoustically identical, and 50% different. Since a difference in vowel length is necessary to establish whether perceptual repair renders the short vowel more similar to [e] or [a], the 80 critical trials were those on which the length of the initial vowels differed.
On each trial, the first stimulus (A) was played simultaneously with a fixation cross, and the second (X) followed after a silent 250ms interval. The short interval was used to favor acoustic judgments (Davidson, 2011), because the goal of the study was to test whether a low-level L1 perceptual illusion can change after late L2 learning. Following the second stimulus, participants responded “same” (right index finger) or “different” (left index finger) on a button box, after which a 500ms blank screen preceded the next trial. Response times (RTs) were measured from the onset of the X. Participants listened to the stimuli on studio quality headphones at a comfortable level.
2.3 Procedure
Following the consent procedure, conducted in Spanish, late bilinguals completed the AX task, then the MINT and cloze tests, and the language history questionnaire. The order of the languages for proficiency testing was counterbalanced. Testing was conducted in quiet rooms. The AX and MINT tasks were presented using E-Prime (Psychology Software Tools, n.d.), the cloze tests using Microsoft Word, and the questionnaire via its online interface (Li et al., 2014).
3 Results
3.1 Predictions
Spanish perceptual repair is expected to lead listeners to perceive the short initial vowel as [e], making it more difficult to distinguish from the longer [e] than from the longer [a]. This perceptual repair effect should lead to less accurate discrimination of long-[e] pairs than long-[a] pairs, and it is expected to be smaller for late bilinguals than monolinguals, and possibly smaller still during English immersion.
The quality of the long vowel may also affect RTs, although Carlson et al. (2016) reported no RT effects for monolinguals or early bilinguals. However, a unique feature of this task for bilinguals, is that the “correct” response is somewhat ambiguous. While the acoustically accurate response to the critical pairs is always “different”, the most Spanish-like response for critical pairs with long [e] is actually “same”. Both acoustically correct and incorrect RTs will therefore be analyzed, but no predictions are made concerning the direction of effects. More veridical perception of [#sC] may allow bilinguals to discriminate long-[e] pairs more quickly, but they may also vacillate between the acoustically accurate response and the response indicated by Spanish phonotactics, leading to slower RTs.
3.2 Data preparation and analysis
Sixty-eight trials with RT < 200ms or > 4000ms were removed, leaving 15,228 valid trials (4,897 critical). Accuracy and RTs were analyzed with logistic and linear mixed effects regression. For accuracy, positive effects indicate increasing probability of “same” responses (“misses” in traditional AX parlance). RTs for (acoustically) correct and incorrect responses were analyzed separately. Based on inspection of quantile–quantile plots, RTs were log transformed to correct for non-normality.
Pairs with the same initial vowel quality were taken from the same recorded token and were thus acoustically identical except for the initial vowel length, but those with different initial vowels were recorded separately. Since exact acoustic match of material following the initial vowel plays no role in the hypotheses considered here, same-vowel and different-vowel pairs were analyzed separately for both accuracy and RTs.
All models shared the same basic structure. 3 Since the critical trials always included one long and one short vowel, pairs were identified based on the quality of the long vowel (vowelCode), sum-coded as -0.5 for [a] and 0.5 for [e]. This predictor captures the effects of perceptual repair
Participants’ language group was a three-level factor. The first (orthogonal) contrast, bilingualCode, compared monolinguals and bilinguals, and the second, ambientLang, compared bilinguals tested in Spain with those tested in the USA. Interactions between language group and vowelCode compare the strength of the perceptual illusion across the three groups. By-participant random intercepts and slopes for vowelCode were included and allowed to correlate. By-item random effects were not included because the items were not randomly sampled from a population, and the differences among items were almost entirely captured by the fixed effects. 4
3.3 Baseline performance (control trials)
Responses to the “same” AX pairs were highly and uniformly accurate across participant groups (mean proportion correct [95% confidence interval (CI)] for monolinguals = 0.96 [0.95, 0.98], late bilinguals in Spain = 0.94 [0.92, 0.96], and late bilinguals in the USA = 0.95 [0.92, 0.97]). The “different” pairs with two long initial vowels contrasting in quality were expected to be the easiest to discriminate and insensitive to the perceptual illusion, because vowel quality is comparatively unambiguous in the long initial vowels. Responses to these pairs were less accurate than for the “same” pairs but were consistent across groups, monolinguals = 0.81 [0.71, 0.90], late bilinguals in Spain = 0.85 [0.80, 0.90], and late bilinguals in the USA = 0.81 [0.75, 0.86]. This indicates a general bias to respond “same,” which was expected because the illusory vowel effect was predicted to increase the subjective proportion of “same” trials by making many of the “different” pairs hard to discriminate.
The remaining non-critical pairs had short initial vowels of contrasting quality. These were expected to be difficult to discriminate, but since both members of these pairs were maximally ambiguous they provide no useful test of perceptual repair. All groups showed similarly low discrimination accuracy, monolinguals = 0.32 [0.24, 0.40], late bilinguals in Spain = 0.39 [0.33, 0.46] and late bilinguals in the USA = 0.33 [0.28, 0.39]. No group differences were significant in any of these three baseline conditions, all p > 0.1.
3.4 Critical trials
The mean proportions of incorrect “same” responses for the critical trials are shown in Figure 1. Table 2 shows the fixed effects estimates for discrimination accuracy on same- and different-vowel pairs. The widely separated CIs for long-[e] versus long-[a] pairs in Figure 1 reveal that short initial vowels were more difficult to discriminate from long [e] than [a], as confirmed by the highly significant main effects of vowelCode in Table 2.

Proportions of “same” responses for critical trials with the same initial vowel (a) and with a different initial vowel (b), by group, with 95% confidence intervals (CIs) computed via nonparametric bootstrap. Values for the means and CIs shown here are given in the Appendix.
Fixed effects estimates (in logit units) for discrimination accuracy.
For same-vowel pairs, the significant interaction of bilingualCode with vowelCode (Table 2) reveals a smaller perceptual repair effect for bilinguals than for monolinguals, visible in Figure 1a. Separate models by long vowel quality confirmed that bilinguals outperformed monolinguals in discriminating long-[e] pairs, β = -1.94, standard error (SE) = 0.60, χ2 = 10.70, p = 0.001, but not long-[a] pairs, p > 0.2. For different-vowel pairs (Figure 1(b)), a similar trend is observable, but the group differences were not significant.
Mean RTs are shown in Figure 2, and Table 3 shows the fixed effects estimates for RT. 5 For incorrect responses, participants were slightly faster when the long vowel was [e], but no group differences were observed. For correct responses, however, group differences in the sensitivity of RT to the long vowel quality are reflected in significant interactions of bilingualCode with vowelCode for different-vowel pairs, and ambientLang with vowelCode for both pair types. Separate models of each group’s correct responses confirmed the pattern seen in Figure 2 (top panels). Bilinguals responded more slowly to long-e than to long-a pairs, in Spain: same-vowel β = 0.09, SE = 0.03, χ2(1) = 8.99, p = 0.03; different-vowel β = 0.09, SE = 0.02, χ2(1) = 13.33, p = 0.0002; in the USA: same-vowel β = 0.27, SE = 0.06, χ2(1) = 16.52, p < 0.0001; different-vowel β = 0.20, SE = 0.03, χ2(1) = 24.12, p < 0.0001. Crucially, the significant interactions of ambientLang by vowelCode confirmed that this effect was larger for bilinguals in the USA than in Spain. No differences emerged for monolinguals, all p > 0.2, although for same vowel pairs it is hard to evaluate the monolinguals’ RTs because they rarely discriminated the long-[e] pairs correctly.

Mean (log) response times for acoustically correct and incorrect responses to critical trials with the same initial vowel (a) and with a different initial vowel (b), by group. Error bars show 95% confidence intervals of the by-subject means. Exact values are reported in the Appendix.
Fixed effects estimates (log-transformed) for response times.
3.5 Proficiency effects
As noted above, while the late bilinguals were all able to use English proficiently for personal and professional purposes, their proficiency scores were not uniform. To explore possible effects of this variability on perceptual repair a set of exploratory models was constructed using only the bilinguals’ data, so that the MINT-English and MINT-Spanish (which were uncorrelated, permitting their effects to be separated) could be added. 6
For discrimination accuracy, three-way interactions of MINT-Spanish, vowelCode, and ambientLang, same-vowel β = 13.31, SE = 6.01, χ2(1) = 4.41, p = 0.04; different-vowel β = 11.45, SE = 5.52, χ2(1) = 4.03, p = 0.04, emerged. The estimated CIs in Figure 3(a) and (b) suggest that these interactions reflect a collapse of the perceptual illusion for bilinguals tested in the USA with lower L1 (Spanish) MINT scores. Note, however, that low L1 MINT scores were not observed among bilinguals in Spain (see Table 1), a point taken up in the discussion.

Proportion “same” responses by bilinguals as a function of first language (Spanish, panels (a), (b)) and second language (English, panels (c), (d)) Multilingual Naming Test (MINT) scores. Points show the empirical by-subject means. Lines and ribbons show predicted means and 95% confidence intervals based on the fixed effects variance in the models. The estimated effects on each plot are shown at the mean MINT score for the other language.
A two way interaction of MINT-English with vowelCode also emerged, different-vowel only, β = -3.29, SE = 1.42, χ2(1) = 5.04, p = 0.02, same-vowel p > 0.18, suggesting a weaker perceptual illusion for bilinguals with higher English MINT scores, but inspection of Figure 3(d) suggests that although the three-way interaction of ambientLang, vowelCode, and MINT-English was not significant, p > 0.23, this only holds for bilinguals tested in the USA. The common thread here is that, in the USA, lower Spanish scores and higher English scores, appear to be associated with less Spanish-like discrimination performance.
No effects of proficiency on RT were found, all p > 0.1, except on correct RTs to same-vowel pairs, apart from a weak interaction of vowelCode, ambientLang, and MINT-English emerged, β = -1.28, SE = 0.59, χ2(1) = 4.65, p = 0.03.
4 Discussion
This study yielded clear evidence for a weakening of the robust L1 Spanish perceptual illusion repairing [#sC], following later learning of L2 English. While late bilinguals’ performance still reflected the illusion, unlike English-dominant early bilinguals (Carlson et al., 2016), they outperformed monolinguals at discriminating short from long initial [e] stimuli for the same-vowel pairs. This confirms the hypothesis that learning English confers an advantage in perceiving distortions to Spanish speech when they match structures found in the L2. No bilingual advantage was found for the different-vowel pairs, but all groups were considerably more accurate in discriminating these pairs, suggesting that they were much easier to discriminate, which may have obscured any effects of bilingualism. Participants’ superior discrimination could be because they were able to detect differences in the quality of the short vowels, but it seems more likely that the slight acoustic differences in the remainder of each stimulus (recall that the different vowel stimuli were created from different recorded tokens, although they were segmentally identical, following the initial vowel) are responsible for the improvement in discrimination.
The effects of language proficiency on discrimination accuracy found in the exploratory models agree with the conclusion that knowing English reduces the perceptual repair effect. That English proficiency was negatively related to the strength of the perceptual illusion (particularly during English immersion) seems straightforward. More interesting are the effects of Spanish proficiency, suggesting that the Spanish perceptual illusion was weakest among participants tested in the USA whose Spanish was a bit “rusty.” Lower Spanish proficiency scores were, unsurprisingly, not observed among participants in Spain, such that the groups were not completely comparable in this respect, but it seems likely that lower L1 proficiency scores and weaker perceptual prothesis are related, both stemming from the reduced contact with Spanish (though self-reported Spanish use was still substantial, mean = 44%, standard deviation = 14%), and/or greater inhibition of the L1 (Baus et al., 2013; Linck et al., 2009), associated with immersion.
The significant group differences in the dependence of acoustically correct RTs on the quality of the longer initial vowel provide further evidence for L2 English effects on L1 Spanish speech perception. Interestingly, effects of initial vowel quality were not found for monolinguals, but rather for bilinguals, particularly those tested in the USA, who were slower to respond “different” when the long vowel was [e]. Bilinguals and monolinguals were uniformly fast, however, when responding “same” across groups. The results can thus be summarized as follows: bilinguals generally favored the same responses as monolinguals, and rendered those responses quickly, as did monolinguals. However, some of the time, on long-[e] pairs, bilinguals deviated from that favored response, and when they did so, they responded more slowly, particularly during English immersion. While monolinguals also occasionally answered “different” for the long-[e] pairs, they did so only rarely on the most difficult same-vowel pairs, and on the different-vowel pairs there was no slowing of responses, as found in the bilinguals, suggesting that the reasons for these “different” responses were not the same for monolinguals and bilinguals.
While the goal of this study was not primarily to identify the mechanisms whereby late L2 learning might lead to the observed weakening of the L1 perceptual illusion, taking the accuracy and RT results together offers some clues for more targeted inquiry. One possibility is that bilinguals, particularly during or after L2 immersion, have simply learned to recognize the absence of the initial vowel more easily (recall that the 2.5 periods of phonation present on the short vowel stimuli are, plausibly, not enough to unambiguously indicate the presence of a vowel). But if learning English simply reduces the confusability of the long-[e] pairs, then we would expect bilinguals’ correct responses in this condition to be faster relative to monolinguals, not slower.
On the other hand, the weakening of the perceptual illusion could result from competition between different language-specific encodings of the short-vowel stimuli (Altenberg & Cairns, 1983; Goldrick, Runnqvist, & Costa, 2014; Gonzales & Lotto, 2013). Specifically, the Spanish encoding is almost certain to be [e]-initial, whereas for English it is most likely to lack any initial vowel (in English, /#sC/ is far more common than /#VsC/; Carlson et al., 2016). Neither encoding matches initial [a], so there is no competition for long [a] pairs, but when compared to long [e], they lead to different responses. Under this interpretation, bilinguals’ responses were slower precisely when the non-dominant, English-like representation of the stimulus not only competed with the Spanish-like representation, but also prevailed in determining the eventual response. This resembles similar RT results in bilingual lexical processing, where the presence of non-target language words as fillers slows responses to interlingual homographs in unilingual lexical decision, where the bivalence of the homographs leads to conflicting responses, but to faster responses in bilingual lexical decision, where words in either language are to be accepted (Dijkstra, van Jaarsveld, & ten Brinke, 1998).
That this slowdown was greater for bilinguals in the USA could be due to greater exposure to nativelike English, or to inhibition of Spanish during immersion (Baus et al., 2013; Green & Abutalebi, 2013; Linck et al., 2009), but the correlation between current and cumulative immersion (see Table 1) makes it hard to say which. Nonetheless, variability in the duration of immersion in the two bilingual groups affords some purchase on this issue. First, 16 of the 30 bilinguals tested in Spain had been immersed in English (3–18 months, with one having 7 years), but their results did not differ from the 14 who had zero immersion experience, all p > 0.6. In addition, replacing ambientLang in the models reported above with (log) years immersed in English yielded inferior goodness-of-fit (assessed using Akaike information criterion and Bayesian information criterion) in all cases except incorrect RTs to same-vowel pairs, where immersion did not modulate the effect of vowelCode anyway.
Moreover, the suggestion in the present data that the largest RT effects were found for low, not high English proficiency bilinguals in the USA, together with the lack of RT effects in early bilinguals (Carlson et al., 2016), argues against a role for cumulative immersion in the RT results. It may be that more proficient bilinguals are better able to juggle conflicting responses in a task such as this, but further study, for example, using a longitudinal design, would be required to test this possibility.
5 Conclusions
The present findings show that learning an L2 with more permissive phonotactics can attenuate L1 perceptual illusions, even when processing the L1, yielding improved discrimination accuracy. Crucially, this occurs even when L2 learning commences after L1 perceptual routines are stable. However, the full pattern of results, encompassing discrimination accuracy, RTs, and the exploratory findings concerning proficiency in both languages, leads to interesting questions about how bilinguals navigate conflicting patterns in their languages as their long-term experience and current conditions change.
The emerging picture is that exposure to L2 patterns that conflict with the L1 provides alternative ways of representing the same acoustic material that can compete during processing. This may cause interference, but it may also be useful given specific task goals and conditions, such as for accurately distinguishing pairs of highly similar stimuli. This resonates with the idea that L2 learning involves not only the addition of a second linguistic system, but a reorganization of both L1 and L2 as a compound system (Cook, 2003; Gildersleeve-Neumann, Peña, Davis, & Kester, 2009; Hall, Cheng, & Carlson, 2006; MacWhinney, 2005, 2008) in which both source systems are constantly active (Kroll, Bobb, & Wodniecka, 2006), but under efficient control depending on a variety of contextual factors, including ambient linguistic conditions (Green & Abutalebi, 2013). Studying this process from a developmental perspective in late bilinguals promises to shed new light not only on how bilinguals come to navigate two systems, but also on the crucial question of how two systems are integrated in the bilingual mind.
Footnotes
Appendix
Mean response times [95% confidence intervals] (computed using the log-transformed values and back-transformed to milliseconds), and number of trials in each cell.
| Monolinguals |
Late bilinguals |
Late bilinguals |
||||
|---|---|---|---|---|---|---|
| Long a | Long e | Long a | Long e | Long a | Long e | |
| Correct responses | ||||||
| Same-vowel | 952ms [864, 1050] n = 115 |
1078ms [767, 1516] n = 14 |
1033ms [961, 1112] n = 288 |
1149ms [1040, 1270] n = 107 |
1017ms [957, 1082] n = 236 |
1464ms [1212, 1768] n = 75 |
| Different- vowel |
976ms [877, 1086] n = 210 |
1022ms [898, 1163] n = 95 |
991ms [937, 1048] n = 459 |
1094ms [1016, 1178] n = 268 |
1006ms [947, 1067] n = 370 |
1230ms [1127, 1342] n = 207 |
| Incorrect responses | ||||||
| Same-vowel | 1092ms [972, 1227] n = 134 |
1017ms [921, 1124] n = 237 |
1092ms [1018, 1172] n = 237 |
1053ms [982, 1129] n = 419 |
1117ms [1050, 1189] n = 211 |
1119ms [1041, 1203] n = 374 |
| Different- vowel |
1238ms [1009, 1519] n = 39 |
1061ms [961, 1171] n = 155 |
1165ms [1054, 1287] n = 69 |
1119ms [1048, 1194] n = 258 |
1321ms [1163, 1501] n = 79 |
1180ms [1106, 1259] n = 241 |
Acknowledgements
The author wishes to thank Matt Goldrick and audiences at the Center for Language Science at Penn State, the Linguistic Society of America, and Sound to Word 16 for their valuable feedback during various stages of the project, along with Teresa Bajo for graciously providing facilities for data collection in Spain, and Felix Huitian and Kyra Krass for their assistance in data collection. Data and models can be obtained from the author.
Funding
Portions of this work were supported by the National Science Foundation Partnerships for International Research and Education grant to the Center for Language Science at Penn State [OISE-0968369].
