Abstract
This study compared individuals from two first language (L1) backgrounds (Italian, Mandarin) to determine how they may differ in their perception of Japanese consonant length (i.e. singleton vs. geminate) according to the phonemic status of length in L1 and experience with Japanese. The participants included two groups of non-native learners of Japanese: 14 native Italian speakers (NI + Japanese), 18 native Mandarin speakers (NM + Japanese) and two control groups: 14 native Italian (NI – Japanese) speakers naïve to Japanese, 10 native Japanese (NJ) speakers. The participants’ length perception accuracy was examined in a forced-choice identification task. The NJ listeners hardly misperceived any tokens, but the non-native listeners were generally accurate (> 85%) in identifying the Japanese length category. The NI – Japanese group was slightly (albeit non-significantly) more accurate than the NM + Japanese group, suggesting the possibility that the use of phonemic length in L1 was facilitative. The direction of misperception (i.e. singleton as geminate or geminate as singleton) differed according to different group. Non-native learners’ results also provided evidence for plasticity in cross-linguistic perception in adulthood.
I Introduction
Languages such as Japanese and Italian use consonant length (i.e. singleton vs. geminate) to distinguish meaning. For example, in Japanese, ika 以下 and ikka 一課 mean ‘below’ and ‘lesson one’, respectively. This type of contrast is a well-known feature of Japanese phonetics and phonology (e.g. Fujisaki et al., 1973; Kawahara, 2015; Kubozono, 2013; Vance, 2008) and is an area of pronunciation that poses difficulties to non-native learners from diverse first language (L1) backgrounds (e.g. English: Han, 1992; Harada, 2006; Hayes-Harb, 2005; Toda, 2007; see, however, Darcy et al., 2013; Korean: Sonu et al., 2013; Mandarin: Hung, 2012; Lu et al., 2016; Uchida, 1993; Mongolian: Liu, 2014; Thai: Ishikawa, 2013; Minagawa and Kiritani, 1996; Vietnamese: Le et al., 2010; Đỗ, 2012, 2015). Research on foreign language (FL) speech learning focusing on non-segmental features such as length is increasing (e.g. Altmann et al., 2012 for German and Italian; McAllister et al., 2002 for Swedish; Meister and Meister, 2011 for Estonian; Ylinen et al., 2005 for Finnish; Hallé et al., 2016 for Tashlhiyt Berber). However, as pointed out in Altmann et al. (2012) and Hallé et al. (2016), most of it focuses on English (e.g. Han, 1992; Hayes-Harb, 2005) or Korean learners (e.g. Sonu et al., 2013) of Japanese.
In earlier work, we examined the role of FL learners’ previous linguistic experience in the processing of a contrast absent in the L1 (Tsukada et al., 2018). The Australian English and Korean learners of Japanese identified consonant length in known/familiar (Japanese) and unknown/unfamiliar (Italian) languages with greater than 80% accuracy. However, native Italian speakers naïve to Japanese were just as accurate as the two groups of learners in identifying the Japanese consonant length.
This study aims to contribute to the understanding of non-native perception of Japanese consonant length by directly comparing two groups of learners from typologically different L1 backgrounds, Italian and Mandarin. Further, Italian-speaking learners and non-learners were compared to determine if there is additional benefit of FL experience in the perception of Japanese consonant length. Finally, Italian speakers who were naïve vis-à-vis Japanese (i.e. non-learners) and Mandarin-speaking learners of Japanese were compared to determine whether familiarity with consonant length in L1 Italian may be more or less beneficial than advanced FL knowledge in the perception of Japanese consonant length. Unlike Italian or Japanese, consonant length is not contrastive at the level of words in Mandarin (see, however, Meng et al., 2021 for ‘fake geminates’ that occur across morpheme boundaries as a result of concatenation). As described below in Section II.2, we expanded learners’ L1 backgrounds in the present study in addition to Australian English and Korean as reported earlier (Tsukada et al., 2018) in order to address the influence of L1 sound systems more broadly.
Apart from Minagawa and Kiritani (1996), this is one of the few studies on Japanese that compared participants from typologically different L1s including one with a consonant length contrast such as Italian. Adding Italian as a target language is expected to be informative from a theoretical perspective as explained below. Since Italian, unlike Mandarin, uses consonant length contrastively (e.g. eco ‘echo’ vs. ecco ‘here (it is)’), native Italian speakers, regardless of FL experience, were expected to possess firm cognitive representations for singleton and geminate categories. Indeed, Tsukada et al. (2018) did find that native Italians naïve to Japanese were able to efficiently utilize their L1 Italian length and did not differ from the L1 Australian English and L1 Korean learners of Japanese in processing the Japanese singleton and geminate consonants. Given that Japanese has traditionally been a popular additional language to learn in many Asian countries, including in China, and Mandarin is ranked within the top 3 languages spoken by teachers and learners of Japanese (The Japan Foundation, 2018), it would be useful to examine how Mandarin learners of Japanese may perform in comparison with learners from other L1 backgrounds.
In the present study, a group of Italian learners of Japanese was added in order to directly compare their Japanese consonant length perception with that of non-native learners from Mandarin background on the one hand and native Italian speakers naïve to Japanese (i.e. non-learners) on the other hand. This approach was designed to enable us to assess whether L1 experience with consonant length in combination with FL Japanese learning might further boost native Italians’ cross-language perception to a level that exceeds that of learners from other L1 backgrounds. All the Mandarin learners had passed the JLPT (Japanese Language Proficiency Test, according to which, the easiest level is N5 and the most difficult level is N1) at N1 level and were considered highly advanced learners. As detailed in Section II.2. below, none of the Italian learners had attained a level of proficiency comparable to the Mandarin learners. It was of interest, therefore, to also examine if experience with L1 consonant length may compensate for lower Japanese proficiency for the Italian learner group. As noted elsewhere, inclusion also of a group of L1 Italian speakers with no learning exposure to Japanese was designed to allow us to explore the effect of L1 expertise vs. FL learner exposure with regard to contrastive consonant length.
We note that Altmann et al. (2012) reported that the German learners (and non-learners) of Italian had trouble with consonant length perception, but Italians with no knowledge of German (which has contrastive vowel length) had no trouble with vowel length perception. The authors concluded that there is asymmetry in cross-linguistic perception for different classes of sounds and that consonant length is difficult even for experienced learners. In this connection, we previously reported that the native Japanese speakers who were naïve to Italian had little difficulty identifying the length category of Italian consonants (in particular, voiceless consonants) (Tsukada and Hajek, 2019) unlike the German learners of Italian in Altmann et al. (2012). Furthermore, the native Italian speakers with no Japanese experience did not differ from the Australian English and Korean learners of Japanese in identifying Japanese consonant length (Tsukada et al., 2018). In other words, both Italian and Japanese speakers naïve to each other’s language were able to utilize their L1 experience to efficiently process unfamiliar consonant length. Based on these findings mentioned above, we predict that the Italian learners of Japanese may positively transfer L1 experience with consonant length and outperform more advanced L1 Mandarin learners. Whether or not the L1 Italian listeners with no Japanese experience outperform the Mandarin learners of Japanese is an open question.
In sum, the research questions addressed in this study were as follows. Firstly, is there a difference in the accuracy of Japanese consonant length perception between native speakers of Italian and Mandarin who were actively engaged in learning Japanese (i.e. L1 effect for learners of Japanese)? Secondly, is there a difference between two groups of Italians who differed in their experience with Japanese (i.e. FL effect for native Italian speakers)? Finally, is there a difference between native Italian speakers who were naïve vis-à-vis Japanese and native Mandarin speakers who were advanced learners of Japanese (i.e. L1 vs. FL effect)? Understanding how listeners’ L1 backgrounds and experience with the target language affect the efficiency of spoken language processing is important from practical and pedagogical perspectives in increasingly diversified, multilingual learning contexts. We hope to gain valuable insights into whether and how different L1s might impact the learning process of Japanese consonant length. The results obtained would provide useful knowledge for improving listening and communication skills in Japanese.
II Method
1 Stimuli preparation
a Speakers
Seven (4 males, 3 females) native speakers of Japanese in their 20s–40s participated in the recording sessions lasting between 45 and 60 minutes. All speakers spoke standard Japanese, having been born or having spent most of their life in the Kanto region surrounding the Greater Tokyo Area. The first author (native speaker of Japanese originally from Tokyo) with expertise in phonetics/phonology auditorily confirmed that all the speakers clearly differentiated the singleton and geminate consonants by duration. The speakers were recorded in a recording studio at a university in Australia or at a research institute in Japan. They received $20 (or equivalent in Japanese yen) for their participation. None of these speakers participated in the perception experiments. According to self-report, they had normal hearing.
b Speech materials
Table 1 shows 78 unique Japanese (non)words analysed in this study. With the exception of 12 (non)words, each item was spoken by two different speakers, resulting in 144 (66 words × 2 speakers + 12 (non)words × 1 speaker) tokens. Of the 144 tokens presented to the participants, 140 (or 97%) were real words. The items included (C)V
(Non)words (n = 78) used in this study.
It would have been desirable to confirm if the non-native learners of Japanese knew all the test words or not. However, as described below, the task was primarily phonetic in nature and did not require lexical familiarity. Thus, we consider our results to be interpretable. This is not intended to imply, however, that familiarity would not affect learners’ response patterns.
To record the stimuli to be used in the perception study, each (non)word was presented on a computer screen in random order and was produced in two separate conditions: one in isolation and the other in a carrier sentence (/sokowa X to jomimasu/ ‘You read it as X there’). The pace of presentation was manually controlled by the experimenter (the first author). This helped to prevent excessive variations in speaking rate by the speakers. The speech materials were digitally recorded at a sampling rate of 44.1 kHz and the target (non)words were segmented and stored in separate files. To avoid inter-speaker variation in fluency (specifically, the duration of a pause before and after the target (non)words), only tokens produced in isolation were used as experimental stimuli in this study.
Table 2 shows some durational characteristics of the stimuli presented to the listeners. For both stops and affricates, the interval between the end of the preceding vowel (i.e. where formant structures indicating periodicity end in wide-band spectrograms and time domain waveforms) and the end of consonantal closure was measured. The values for singleton and geminate consonants are very close to what has been reported in previous research (e.g. Hayes-Harb, 2005) with geminates being more than 2.5 times as long as singletons. While the difference in vowel duration in the two contexts (i.e. preceding singleton vs. geminate) was not statistically significant, the trend pattern of longer vowel duration before geminate consonants is not unexpected: other studies have shown pre-geminate vowels to be phonetically longer in Japanese (e.g. Han, 1992; Hussain and Shinohara, 2019; Idemaru and Guion, 2008).
Durational characteristics of the stimuli (in ms) (standard deviations in parentheses).
2 Participants
Table 3 summarizes the languages spoken by our participants in the present study and their L1/FL experience with consonant length. The four groups of listeners participated in forced-choice consonant length identification experiments as described in Section II.3.
Listeners in this study and their experience with Japanese.
Notes. f = female. m = male.
The first two groups consisted of L1 Italian (NI + Japanese, 7 males, 7 females, mean age = 23.4, sd = 1.6) and L1 Mandarin (NM + Japanese, 3 males, 15 females, mean age = 24.7, sd = 4.5) learners of Japanese. The NI + Japanese group consisted of students at a university in Italy. Two of the NI + Japanese learners had passed the JLPT at N2 level, six had passed N3 and one had passed N4, respectively. Five (1 male, 4 females) of the NM + Japanese learners were originally from Taiwan. The NM + Japanese learners took part in the study in Japan and had a mean length of residence of 1.3 years (range = 0.1–5.0) in Japan. As indicated above, all the NM + Japanese learners had passed JLPT at N1 and were considered as generally more advanced learners than the NI + Japanese. Crucially, however, and also as previously noted, only Italian uses consonant length contrastively. All learners were recruited from the student/staff populations at each institution.
The third group consisted of L1 Japanese (NJ, 5 males, 5 females, mean age = 25.5, sd = 6.9) listeners who previously served as native controls (Tsukada and Hajek, 2019; Tsukada et al., 2018). The NJ listeners were recruited from the student/staff populations at universities or from the local communities in Australia or Japan. The fourth and last group consisted of L1 Italian (NI – Japanese, 7 males, 7 females, mean age = 28.9, sd = 7.6) listeners who served as non-native controls, as they had no knowledge of Japanese. Except for three NI – Japanese listeners whose results were included in our earlier study (Tsukada et al., 2018), the NI – Japanese listeners were recruited at the same Italian university as the NI + Japanese listeners.
3 Procedure
The listeners participated in a forced-choice identification task and listened to a total of 252 (84 × 3 blocks) Japanese tokens. For the present study, a subset of these tokens (n = 144) with intervocalic /t k ʧ/ (72 tokens each with singleton and geminate) was included in the analyses. The presentation of the stimuli and the collection of perception data were controlled by the UAB (University of Alabama at Birmingham) software (Smith, 1997) for the NM + Japanese learners tested in Japan. The Praat program (Boersma and Weenink, 2016) was used for the participants tested elsewhere more recently, because the UAB software was no longer available.
The listeners’ task was to decide whether the medial consonant was short/singleton or long/geminate and indicate their choice using response categories appropriate for their linguistic background on the computer. Specifically, the NJ listeners and the NM + Japanese learners were given response categories with a geminate symbol clearly marked in Japanese scripts (i.e. small ‘tsu’ or「っ」) whereas the NI + Japanese and NI – Japanese listeners were given response categories labelled Singola ‘Single’ and Doppia ‘Double’. None of the groups saw the stimulus words spelled out in written form. The listeners made their responses by clicking on the response categories of their choice on the computer monitor with a mouse. The listeners were allowed to replay the stimulus tokens multiple times in order to reduce their anxiety and were asked to guess if uncertain. Once a choice was made, the listeners were unable to change their decision and the next token was presented automatically. The listeners were tested individually in a sound-attenuated booth or quiet classroom on the university campus in their country of residence. The experimental session was self-paced, but typically it lasted between 30 and 40 minutes. They heard the stimuli at a self-selected, comfortable amplitude level over the headphones on a computer and received $20 (or equivalent) for their participation.
III Results
We used R version 3.6.0 for statistical analyses and data visualization reported below (R Core Team, 2019).
1 Overall results on consonant length perception
Figure 1 shows the distributions of percentages of correct length identification by the four groups of listeners. Table 4 shows the mean correct identification and the number of listeners whose identification accuracy reached the range set by the NJ group (97.9%–100%).

The distributions of correct length identification (%) of four groups of listeners (NJ, NM + Japanese, NI + Japanese, NI – Japanese).
Mean correct length identification (standard deviations in parentheses) of four groups of listeners (second row) and the number of listeners whose identification accuracy reached the range set by the NJ group (97.9%–100%) (third row).
Table 4 shows that the mean correct identification ranged from 86.3% (sd = 10.8) for the NM + Japanese group to 99.7% (sd = 0.7) for the NJ group. As expected, the NJ group was homogenous and highly accurate. All non-native groups (including the NI – Japanese group with no Japanese learning experience) were generally accurate (> 85%) in perceiving Japanese consonant length. This is in contrast to the native German listeners who were found by Altmann et al. (2012) to have much lower accuracy rates (i.e. learners of Italian (mean dʹ = 1.95, sd = 0.76) and non-learners of Italian (mean dʹ = 1.31, sd = 0.76), respectively) than native Italian listeners (mean dʹ = 3.04, sd = 0.44). It is striking in Figure 1 that the NI + Japanese group was much less variable compared to the NI – Japanese and NM + Japanese groups.
According to a one-way analysis of variance (ANOVA), the between-group difference in consonant length identification was significant [F(3, 52) = 8.2, p < .001, η2G = 0.32]. Dunnett’s Modified Tukey–Kramer pairwise multiple comparison post hoc tests (for mean differences with unequal sample sizes and no assumption of equal population variances) showed that the NJ group was significantly more accurate than the NI + Japanese, NI – Japanese and NM + Japanese groups and that the NI + Japanese group was significantly more accurate than the NM + Japanese group. No other two comparisons differed significantly. It is notable that the NI + Japanese group showed very little variability in addition to being more accurate than the highly advanced NM + Japanese learners despite its more limited Japanese proficiency.
The test of proportion was used to assess the between-group differences in the number of listeners whose identification accuracy reached the range set by the NJ group. The overall difference was significant [χ2(3) = 23.6, p < .001] and pairwise comparisons indicated that all three non-native groups had significantly lower proportions of native-like listeners than the NJ group [p = .00073 – .02178].
2 The effect of contrastive consonant length
Figure 2 shows the distributions of percentages of correct length identification by the four groups of listeners as a function of the length category of the target consonant. As we noted earlier for Figure 1, it is striking how much smaller the range of results is for the NI + Japanese group compared to the NI – Japanese and NM + Japanese groups for both target geminates and singletons. While the NM + Japanese group had higher mean identification accuracy for the geminate (90%) than for the singleton counterparts (83%) [t(17) = −2.2, p < .05], the other groups were relatively balanced in their mean identification accuracy for the two length categories (NI – Japanese: 89% vs. 90%, NI + Japanese: 95% vs. 97%, NJ: 99% vs. 100% for geminate and singleton, respectively). The bias towards hearing geminates has been reported in previous studies involving various Asian languages (e.g. Mongolian (for stops but not fricatives): Liu, 2014; Thai: Ishikawa, 2013; Vietnamese: Đỗ, 2012, 2015).

The distributions of correct length identification (%) of four groups of listeners (NJ, NM + Japanese, NI + Japanese, NI – Japanese) as a function of length category of the target token (geminate, singleton).
A two-way repeated-measures ANOVA with group (G: NJ, NM + Japanese, NI + Japanese, NI – Japanese) as a between-subjects factor and consonant length (L: long/geminate, short/singleton) as a within-subjects factor yielded a significant main effect of G [F(3, 52) = 8.2, p < .001, η2G = 0.26] and a two-way interaction effect [F(3, 52) = 3.6, p < .05, η2G = 0.05], but not of L. A significant two-way interaction was expected due to a different pattern of length identification for different groups as clearly seen in Figure 2.
The simple effect of Group was significant for both length categories. Table 5 shows the results of one-way ANOVA which assessed the effect of Group (not assuming equal variances) and of Dunnett’s Modified Tukey–Kramer pairwise multiple comparison post hoc tests. For geminates, the NJ group outperformed the NI + Japanese, NM + Japanese and NI – Japanese groups. The three non-native groups did not significantly differ from one another despite the difference in Japanese proficiency and/or experience. For singletons, the NJ listeners never misperceived any tokens and could not be included in the analysis. Of the three non-native groups, the NI + Japanese group identified Japanese singletons significantly more accurately than did both the NI – Japanese and NM + Japanese groups, who did not differ from each other. Thus, while the NI + Japanese learners diverged from the NJ listeners in identifying geminates, they were more accurate than the other two groups of non-native listeners in identifying singletons. Even with advanced knowledge of Japanese, the NM + Japanese learners identified singletons and geminates less accurately than did the NI + Japanese learners (significantly so for singletons) and did not outperform the NI – Japanese listeners who were naïve to Japanese. Welch’s t-tests assessing the effect of Length did not reach significance for any of the groups except for the NM + Japanese group as mentioned above.
Results of one-way ANOVA assessing the effects of Group and multiple comparison tests.
Note. Significance level at 0.05.
As for the two Italian groups who were expected to possess firm, long-term representations for L1 Italian consonant length, the NI + Japanese group outperformed the NI – Japanese group in identifying Japanese singletons, but not geminates. Thus, actively learning another quantity language such as Japanese was beneficial for native Italian speakers’ cross-language speech perception at least to some extent.
IV Discussion and conclusions
We investigated the perception of Japanese consonant length by native and non-native listeners from diverse L1 backgrounds focusing on how their identification accuracy might also be affected by the phonemic status of length in L1 and experience with Japanese. Two of the non-native groups had Japanese learning experience with the NM + Japanese group being more advanced than the NI + Japanese group. The NI – Japanese listeners had no experience with Japanese, but they had experience with L1 consonant length.
As expected, the NJ listeners were the most accurate and hardly misperceived any tokens regardless of the length category of the target token. Despite diversity in L1 and FL Japanese experience, the non-native groups were generally accurate (> 85%) in their Japanese consonant length identification. Averaged across the two length categories, the NJ group was significantly more accurate than all three non-native groups. Of the non-native groups, the only between-group difference that reached significance was the comparison between the NI + Japanese and NM + Japanese groups with the former being more accurate than the latter. While the latter group was formally more proficient in Japanese than the former, it appears that the effect of prior L1 Italian experience of a phonemic length contrast combined with some exposure to Japanese had a greater positive effect than high proficiency alone in Japanese.
When we analysed the results for singletons and geminates separately, the between-group differences were observed only for singletons. While all three non-native groups were significantly less accurate than the NJ group on geminates, the NI + Japanese group outperformed the NI – Japanese and NM + Japanese groups on singletons. Here, again, it appears that a combination of L1 length experience and FL Japanese learning helped the NI + Japanese group more than the NM + Japanese group’s high proficiency in Japanese. Further, it is notable that, even with advanced knowledge of Japanese, the NM + Japanese group did not outperform the NI – Japanese group who was naïve to Japanese, confirming a profound influence of L1.
Our results were in contrast with those in Altmann et al. (2012). The German learners of Italian in their study had trouble with Italian consonant length, maybe because they were not proficient enough in Italian even though ‘they had all studied Italian at the university level for at least 11 months’ (p. 395) and had experience with an L1 vowel length contrast. Alternatively, the speeded same-different task these researchers used was more challenging than the forced-choice identification task employed in this study. Yet another possibility is Italian consonant length is less discriminable than Japanese consonant length due to cross-linguistic phonetic differences. Based on their Figure 1 (p. 400), none of the German learners of Italian appears to have reached the range set by the native Italian group. On the other hand, at least 17% of our listeners in the two learner groups attained native-like Japanese consonant length perception in our study (Table 4). As all the learners started learning Japanese as young adults, our results show that it is possible, at least for some adults, to learn difficult non-native sound contrasts to a level comparable to native speakers.
In sum, all three non-native groups of listeners including the NI – Japanese group which was naïve to Japanese identified consonant length with mean accuracy of 85% or higher. Of the learner groups, the NI + Japanese learners significantly outperformed the NM + Japanese learners in the overall identification of the Japanese consonant length category, which suggests that a combination of L1 consonant length experience and Japanese learning experience is necessary for attaining more accurate cross-linguistic consonant length identification.
In future work, it would be useful to examine what other phonetic factors such as manner and place of articulation affect listeners’ consonant length perception (e.g. Arango et al., 2021; Liu, 2014). As Japanese uses both consonant and vowel duration contrastively, we would like to assess listeners’ length perception of different classes of sounds to confirm or disconfirm whether vowel length contrasts do indeed ‘pose little problems’ (Altmann et al., 2012) and consonantal length contrasts prove ‘to be difficult’ (Altmann et al., 2012). In the case of Japanese, perceptual accuracy of vowel vs. consonant length appears to differ for different L1 groups (e.g. vowel length less accurate than consonant length for Vietnamese learners in Đỗ, 2012). The findings would have theoretical (i.e. cross-language speech perception) and pedagogical (i.e. Japanese pronunciation teaching/learning) implications. Our result that some learners in each group attained native-like Japanese consonant length perception suggests that, with sufficient experience, non-native learners from a wide variety of L1 backgrounds are capable of learning even the difficult length contrast accurately (> 85%).
Footnotes
Acknowledgements
We thank Valentina De Iacovo for research assistance and participants for their co-operation. We also thank the Editorial Team and three anonymous reviewers for their time and input.
Author Note
Portions of this study were presented at the 19th International Congress of the Phonetic Sciences in 2019 and the 10th International Conference on Speech Prosody in 2020.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by the 11th Hakuho Foundation Japanese Research Fellowship (2016–2017) to the first author.
