Abstract
Recent research has proposed that listeners use prosodic information to guide their processing of phonemic contrasts. Given that prosodic organization of the speech signal systematically modulates durational patterns (e.g., accentual lengthening and phrase-final (PF) lengthening), listeners’ perception of durational contrasts has been argued to be influenced by prosodic factors. For example, given that sounds are generally lengthened preceding a prosodic boundary, listeners may adjust their perception of durational cues accordingly, effectively compensating for prosodically-driven temporal patterns. In the present study we present two experiments designed to test the importance of pitch-based cues to prosodic structure for listeners’ perception of contrastive vowel length (CVL) in Tokyo Japanese along these lines. We tested if, when a target sound is cued as being PF, listeners compensatorily adjust categorization of vowel duration, in accordance with PF lengthening. Both experiments were a two-alternative forced choice task in which listeners categorized a vowel duration continuum as a phonemically short or long vowel. We manipulated only pitch surrounding the target sound in a carrier phrase to cue it as intonational phrase final, or accentual phrase medial. In Experiment 1 we tested perception of an accented target word, and in Experiment 2 we tested perception of an unaccented target word. In both experiments, we found that contextual changes in pitch influenced listeners’ perception of CVL, in accordance with their function as signaling intonational structure. Results therefore suggest that listeners use tonal information to compute prosodic structure and bring this to bear on their perception of durational contrasts in speech.
1 Introduction
When listeners process speech, they need to extract (among other things) a segmental message and a prosodic message from the speech signal (e.g., Cho et al., 2007; Kim et al., 2018). That is, listeners need to compute what segmental contrasts and lexical items are intended by the speaker, while also computing prosody, which will tell listeners how items are grouped in phrases, and what the prominence relations are among them. Sometimes the acoustic information which specifies these segmental and prosodic structures may be the same. We can consider this in light of a body of literature that shows higher-level prosodic structure is encoded in a detailed way by various properties of individual speech segments and articulations (e.g., Byrd et al., 2000; Cho, 2015, 2016; Cho & Keating, 2009; De Jong, 1995; Fougeron & Keating, 1997; Keating, 2006; Keating et al., 2003). One well-attested pattern in this literature is the role of prosodic boundaries in fine-tuning timing and amplitude patterns for articulations.
For example, segments are realized as “stronger” initial to prosodic domains (e.g., Cho, 2016), produced with increased articulatory contact, duration of closure, etc. (e.g., Keating et al., 2003). This introduces systematic patterns in the acoustic properties of speech sounds as a function of prosodic organization. Consider an example that illustrates the dual-function of acoustic cues domain-initially. Voice onset time (VOT), which serves to differentiate stop voicing categories (e.g., Abramson & Lisker, 1970) is also systematically modulated by prosodic structure. At the beginning of phrasal domains VOT is lengthened in some languages as a manifestation of domain-initial strengthening (Cho & Keating, 2009; Keating et al., 2003). If we consider prosodic position as a source of contextual variability for VOT along these lines, its relevance in the perception of, for example, voicing contrasts becomes evident: if VOT varies as a function of prosody, listeners would benefit from reconciling a given VOT value with the prosodic context in which it occurs, making reference to whether VOT is lengthened as a function of initial strengthening, or is simply long to cue a voicing contrast at the lexical level. Put differently, listeners would benefit from using “prosodic information in determining whether segmental information is driven lexically or post-lexically (prosodic-structurally)” (Mitterer et al., 2019, p. 14). Recent research, discussed below, suggests this may indeed be the case.
In similar fashion to documented “initial strengthening” at the left edges of domains, one well-described pattern in the literature that occurs at right edges is final lengthening, also called pre-boundary lengthening (e.g., Cho, 2015, 2016; Turk & Shattuck-Hufnagel, 2007; Wightman et al., 1992). Generally speaking, this refers to the temporal expansion of linguistic units preceding a prosodic boundary (Cho, 2015, 2016). Though domain-initial prosodic effects in speech perception have received recent attention in the literature (Kim & Cho, 2013; Mitterer et al., 2016, 2019), perception of segmental contrasts in phrase-final position remains relatively less studied.
Domain-final prosodic effects and their relationship to other aspects of linguistic structure can be considered more broadly as part of the prosodically-driven temporal organization of the speech signal. Turk and White, for example, conceptualize prosody as “hierarchical structure [that] influences the domain and distribution of durational effects” (Turk & White, 1999, p 171; see also Turk & Sawusch, 1997; Turk & Shattuck-Hufnagel, 2007). As discussed for the example of domain-initial VOT, prosodic organization along these lines might be thought to play an important role in speech perception and spoken language processing. Prosodic patterning in the temporal domain can be taken to systematically modulate the duration of segmental structure, and acoustic cues over time, and may accordingly exert a mediating influence in listeners’ perception of various durational contrasts in speech.
In the present study we test the influence of pre-boundary lengthening on the perception of contrastive vowel length (CVL) in Tokyo Japanese. We also manipulate only fundamental frequency (F0) as a contextual cue to intonational structure, which allows for control over other possible durational context effects, discussed below. The present study therefore presents a new testing ground for the relevance of intonational structure, and its temporal encoding, in listeners’ processing of durational contrasts in speech.
1.1 Previous studies
Several previous studies have presented evidence for the role prosodic boundaries play in listeners’ perception of segmental contrasts, and more generally in other domains of speech processing such as word segmentation (see e.g., Cho et al., 2007). Kim and Cho (2013) tested the aforementioned pattern of initial strengthening and perception of VOT in American English. They carried out an experiment in which listeners categorized a VOT continuum as /p/ or /b/. The target sound was placed in a carrier phrase “let’s hear pa/ba again,” and the presence/absence of a preceding intonational phrase (IP) boundary before the target word was manipulated (among other things). The phrase boundary was cued by a low boundary tone, and phrase-final (PF) lengthening. The authors predicted that listeners may use the boundary to adjust their categorization of VOT, in line with domain-initial strengthening. Specifically, if listeners expect phrase-initial lengthening of VOT, they should effectively require longer VOT for a voiceless aspirated /p/ response. The authors found this effect, though the picture is complicated by another possible explanation. As discussed by Mitterer et al. (2016), lengthening preceding a target sound could modulate perception of VOT independent of prosody. Because a preceding boundary was encoded by pre-boundary lengthening, perception of VOT (a durational cue) may be expected to shift on the basis of local speech rate normalization, that is, perceptual adjustments for changes in durational context (e.g., Miller & Liberman, 1979; Newman & Sawusch, 1996; Summerfield, 1981). This presents a possible explanation because preceding lengthening may cause subsequent VOT to be perceived as relatively short, as a function of durational contrast (e.g., Diehl & Walsh, 1989). This effect would therefore decrease listeners’ /p/ responses in the post-boundary condition, on the basis of adjustments for preceding speech rate. This alternative explanation does not implicate prosodic structure, and thus offers an interesting illustration of the complexity in testing for prosodic effects in perception, especially in the case of durational cues (see Mitterer et al., 2016; Steffman, 2019a for further discussion).
One promising way to extend this work which has recently been explored in the literature is testing perception of non-durational contrasts (Mitterer et al., 2019) or manipulating contextual cues that are not in the temporal domain (Kim et al., 2018). Mitterer et al. (2019) tested domain-initial effects on vowel realization in Maltese. Glottalization in Maltese is phonemic such that there exist minimal pairs such as /ʔɑ:m/ “he woke up” and /ɑ:m/ “he swam.” At the same time, glottalization in vowel-initial words is a manifestation of initial strengthening, occurring at the beginning of higher-level phrasal domains. Thus, in similar fashion to VOT, glottalization cues a phonemic contrast, while simultaneously varying along a prosodic dimension. Unlike VOT however, glottalization is not an inherently temporal cue. Accordingly, in one experiment the authors created a glottalization continuum ranging from /ʔɑ:m/ to /ɑ:m/, while holding duration constant. The authors were curious about how perception of this contrast would vary based on whether the target was cued as being phrase-initial, or not. Listeners categorized the target in a carrier phrase, with the presence/absence of preceding pre-boundary lengthening manipulated, to signal the target as phrase-initial. The crucial prediction was that listeners should be more likely to interpret glottalization as phonemic (as in /ʔɑ:m/), when the target was not preceded by boundary cues, that is, when domain-initial glottalization would not occur. On the other hand, a target cued as phrase-initial may be glottalized to encode prosodic structure (in lieu of a lexical contrast), and would therefore be more likely to be perceived as /ɑ:m/. The authors find this pattern, importantly with a cue that is not temporal such that speech rate normalization does not offer an alternative explanation. We can take this finding to highlight the role of prosody in the perception of domain-initial words on the basis of the way they are influenced by initial strengthening, here encoded with glottalization.
Another relevant study, Kim et al. (2018), tested if listeners reference prosodic structure, cued by pitch alone, in a phonological inferencing task. They tested Korean post-obstruent tensing (POT), whereby lax stops and affricates become tense following another obstruent. To use an example from the paper: /puri/ “beak” will become tensified [p*uri], when following an obstruent, as in the sequence /porasɛk # puri/ “purple beak,” making it confusable with /p*uri/ “root.” Importantly, the domain of this process is within the Korean accentual phrase (AP). Kim et al. (2018) used a visual world eye-tracking experiment in which listeners heard a color term and target word, as in the example above, and looked to colored orthographic representations as they listened to speech. The relevant test case is one in which listeners hear a tense obstruent: the question is if they will infer that POT has applied and look to an underlying lax form, for example, /puri/, or an underlying tense form, for example, /p*uri/. Crucially, the authors predicted that this process should be modulated by prosody: because the domain of POT is within the AP, listeners would necessarily need to reference phrasing to determine whether POT may have applied. For example, AP-internal (porasɛk p*uri), where parentheses indicate an AP boundary, may be “beak” or “root,” however an intervening AP boundary disambiguates the meaning: (porasɛk) (p*uri), can only be “root.” Kim et al. (2018) found that when listeners heard a phonetically tense target [p*uri], they looked more to an underlyingly lax word /puri/ when that word was in a context that licensed POT (e.g., / porasɛk # puri/), as compared to when it was not (when a non-obstruent preceded the target sound). In these cases, the target and preceding color term were additionally phrased together in a single AP. This evidenced the predicted phonological inferencing effect. To test how a perceived AP boundary might modulate this effect, Kim et al. (2018) manipulated F0 alone to signal an AP boundary between the color term and following target word. Due to the role of AP boundaries in restricting application of POT, the authors predicted that the effect observed for an AP-internal color-target sequence should not be observed when an AP boundary intervenes, showing listeners’ sensitivity to the phrasal domain of the phonological process. To cue an AP boundary, Kim et al. (2018) used rising F0, while leaving the temporal properties of their stimulus unaltered. As expected, the phonological inferencing effect disappeared in this case, presenting evidence that listeners used tonal cues to compute the phrasing for the target word such that it modulated phonological inferencing. This study can therefore be taken to show the relevance of tonal cues to prosodic domains, in this case, for the purpose of inferring whether a phonological process has applied. As Kim et al. (2018) note, an effect observed with prosodic structure signaled only by pitch, offers a strong argument for the role of language-specific prosodic structural effects in listeners’ processing of speech, as compared to a speech rate normalization account which is language-general.
These two previous studies therefore suggest that, independent of normalization for durational changes, listeners reference prosodic structure in speech processing. While Kim et al. (2018) focused on more abstract phonological inferencing effects, Mitterer et al. (2019) showed that these effects can extend to listeners’ perception of phonetic detail in speech. Kim et al. (2018) also showed that tonal cues may play a central role in listeners’ perception of prosodic structure in this domain. An emergent view that has come about based on these findings is that listeners process segmental and prosodic structures in parallel, as described by Mitterer et al. (2019): “listeners compute the prosodic structure (prosodic processing), possibly in parallel with segmental processing” (see also Cho et al., 2007; Kim et al., 2018, Nakai & Turk, 2011). The role of phrasal prosodic structure in this view is to guide interpretation of phonetic detail (Kim & Cho, 2013; Mitterer et al., 2019), and to modulate other domains of speech processing such as phonological inferencing (as in Kim et al., 2018) and word segmentation and recognition (e.g., Cho et al., 2007; Christophe et al., 2004; Salverda et al., 2003). In this vein, we can consider the set of findings presented above to evidence the involvement of prosodically guided segmental processing for domain-initial patterns (Kim & Cho, 2013; Mitterer et al., 2019), as well as highlighting the importance of pitch as a cue to prosodic domains, as shown in Kim et al. (2018).
Several past studies have also looked at right-edge positional effects on listeners’ perception of durational cues. Nooteboom and Doodeman (1980), manipulated the syntactic structure of various carrier phrases in which a Dutch target word was placed. Dutch listeners categorized it as having a phonemically short or long vowel. The prediction was that if listeners are sensitive to phrases as the domain for final lengthening (here represented by Nooteboom & Doodeman (1980) in terms of syntactic structures), they should adjust their perception of CVL. In particular, if a given target sound is perceived as being phrase, or utterance-final, listeners should expect it to have undergone lengthening. In similar fashion to VOT in American English (as in Kim & Cho, 2013) or glottalization in Maltese (as in Mitterer et al., 2019), we can therefore consider duration as serving to cue a phonemic contrast in Dutch, while also providing listeners with information about prosodic organization. If this prediction were borne out, we should expect to see listeners compensatorily adjust their categorization of vowel length, that is., a vowel in PF position will need to be longer than a vowel in phrase medial position to be perceived as phonemically long. Nooteboom and Doodeman (1980) find this result, which presents suggestive evidence that Dutch listeners may use prosody to guide their perception in this way. However, given that Nootebaum and Doodeman’s (1980) carrier phrases consisted of different words, and were produced in different utterances, their stimuli present a high degree of variability in context: the words surrounding a target, and adjacent segmental durations, vary across conditions. In light of the possible effects of durational context on listeners’ perception of durational cues (as discussed in Mitterer et al., 2016; Steffman, 2019a), these results should perhaps be interpreted cautiously. More recently, Steffman (2019b) tested how American English-speaking listeners would perceive vowel duration as a cue to coda obstruent voicing, using a “coat” to “code” continuum, where longer vowels occur before voiced obstruents (e.g., Chen, 1970) and are used as a cue to voicing (e.g., Raphael, 1972). Steffman (2019b) manipulated phrasal position simply as the presence/absence of following material in a carrier phrase. The stimuli were designed such that speech rate normalization effects should predict the opposite of a prosodically guided interpretation of duration, by making added post-target material lengthened such that it would predict contrast effects when present (see also Miller & Liberman, 1979). Steffman (2019b) found the expected prosodically-guided effect: listeners required longer vowel duration for a voiced “code” percept when the target sound was PF, suggesting a compensatory adjustment for PF lengthening. These findings together suggest that both Dutch and American English-speaking listeners exhibit a sensitivity to right-edge durational patterns in their perception of durational cues.
In the present study we present two experiments which seek to extend the research outlined above. We test how Tokyo Japanese speaking listeners’ interpretation of a target sound as (intonational) PF or phrase-medial influences their perception of CVL, in the same vein as Nootebaum and Doodeman (1980), and Steffman (2019b). This will, broadly, help better our understanding of right-edge temporal effects on listeners’ processing of durational cues cross-linguistically, and accordingly help inform recent proposals related to the parallel processing of prosodic and segmental structures. Our goal in testing Japanese is to additionally implement purely tonal cues to prosodic structure, as informed by models of Japanese intonational phonology, discussed below. To this end we manipulated only F0 in a carrier phrase, which avoids possible pitfalls of changing the durational context in which a target sound appears, as outlined above. This also presents a departure from Nootebaum and Doodeman (1980), and Steffman (2019b), who either did not control for F0 across conditions, or did not manipulate it all. This will thus allow us to assess the role that pitch-based cues play for Japanese listeners in their interpretation of prosodic structure, building on Kim et al.’s (2018) finding for AP-final tonal patterns in phonological inferencing in Korean.
1.2 The present study
Tokyo Japanese presents a valuable test case, given that it has a well-described intonational system, and CVL. The present study thus tests if listeners rely on intonation to compute prosodic boundaries, which may mediate their processing of durational cues.
In describing the intonational contexts manipulated in our experiment, we adopt the autosegmental-metrical (AM) model of Japanese intonational phonology developed by Beckman and Pierrehumbert (1986), Pierrehumbert and Beckman (1988), Venditti (1995, 2005) and Maekawa et al. (2002). In the recent versions of the AM model, there are two tonally-defined prosodic groupings above the word level: the AP; and the IP. The AP is defined as having a phrasal H tone (H-) around the second mora and a subsequent gradual fall to a low tonal target (L%) at its right edge. It is also regarded as the domain of pitch accent realization: an AP can accommodate at most one pitch accent (shown as “A” in XJ-ToBI, as proposed by Maekawa et al., 2002), which is realized as a sharp F0 fall starting near the end of the accented mora. The IP, on the other hand, consists of one or more APs, and is marked by an initial L boundary tone (%L). Each IP has its own pitch range, so the effect of downstep, by which the F0 height of a pitch accent is lowered when following another pitch accent, is reset at an IP boundary.
As in many languages, PF lengthening has been documented in Japanese (e.g., Takeda et al., 1989). On the basis of the prosodic hierarchy proposed by Beckman and Pierrehumbert (1986) and Pierrehumbert and Beckman (1988) (which includes an intermediate phrase), Ueyama (1999) conducted a production experiment, comparing vowel duration at the right edge of four different prosodic levels: AP; intermediate phrase (ip); and sentence. Ueyama’s (1999) results showed that the four prosodic levels fell into two groups such that a vowel was significantly longer in IP-final position and sentence-final position as compared to AP-final position and ip-final position. This indicates that the duration of a domain-final vowel varies with the strength of the prosodic boundary. Seo et al. (2019) additionally find that PF lengthening in Japanese appears to be mediated by the syllable structure (cf. Shepherd, 2008) by showing that the effect of PF lengthening on the rime of the consonant–vowel–noun syllable (e.g., takan “sensitive”) was comparable to that on the final vowel of the consonant–vowel syllable (e.g., taka “hawk”).
The present study addresses the perceptual relevance of PF lengthening in two experiments, testing if listeners’ perception of CVL shifts based on whether a target sound is expected to undergo PF lengthening (when cued as phrase final). In both experiments, listeners categorized a vowel from a vowel duration continuum as phonemically long or short. Target words in both experiments contrasted only in the length of the final vowel, and were both disyllabic; however, accentedness varied across experiments. In Experiment 1, the target word had an accent on the first syllable, while in Experiment 2 it did not. Testing both an accented and unaccented word pair allows for basic replication of the predicted effect, and allows us to help generalize the effect by testing different tonal environments across experiments, which vary based on the pitch accent status of the target word (described below). Findings from the present study will accordingly better our understanding of how prosody, and particularly tonal cues, shape listeners’ perception of domain-final temporal structure in speech, extending the lines of research outlined above.
2 Experiment 1
In Experiment 1 we tested how Japanese listeners were influenced by changes in contextual F0 in their perception of CVL, in this case for an accented target word minimal pair. We implemented a two-alternative forced choice task, in which listeners categorized a sound from a vowel duration continuum as phonemically long or short. F0 was manipulated in a carrier phrase to signal a target as IP-medial or IP-final.
2.1 Methods
Listeners categorized a disyllabic minimal pair with accent on the first syllable, that contrasted in terms of the phonemic length of the second syllable. The target was categorized as shi’sho “librarian” (司書) or shi’shoo 1 “master” (師匠). These words are both fairly low frequency based on word counts in Corpus of Spontaneous Japanese (CSJ) (Maekawa, 2003; Maekawa et al., 2000) (log-transformed frequency = 0.3 for shi’sho, 1.28 for shi’shoo). The carrier phrase used for both medial and final conditions is shown in (a) below, with glosses. It had a similar design to the carrier phrase from Shepherd’s (2008) production study.
(a) wata’shitachi-wa x (target) de’sukara shinraideki-ma’su we-TOP x (target) because/therefore reliable-be
(a) can be phrased in two different ways, which are represented in (b) and (c) below, where [. . .]AP represents an AP boundary and [. . .]IP represents an IP boundary. Importantly, changing the phrasing along these lines alters the implied position of the target sound, such that in (b) it is IP-medial (and AP-medial), and in (c) it is IP-final. Both phrasings are possible, by virtue of the fact that Japanese sentences can end with a noun omitting a copula. 2 A translation is given as well.
(b) One IP: “Because we are x (we are) reliable.” (x = medial) [[wata’shitachi-wa]AP [x de’sukara]AP [shinraideki-ma’su]AP]IPI (c) Two IPs: “We are x. Therefore (we are) reliable.” (x = final) [[wata’shitachi-wa]AP [x]AP]IP [[de’sukara]AP [shinraideki-ma’su]AP]IP
Visual representations of this manipulation for the stimuli used in Experiment 1 are shown in Figure 1. In creating the different possible phrasings, F0 was manipulated to differ only on the second syllable of the target word, and the following syllable (/de/ in de’sukara “because/therefore”). Manipulating F0 in this way was judged to be desirable because it entailed a minimal difference across conditions, but also produced a clear perceived change in phrasing as judged by a ToBI-trained native Japanese speaker (author Hironori Katsuda). In the medial condition (Figure 1, top panel), the target word x forms a single AP with the following conjunction de’sukara. Since an AP has at most one pitch accent, the accent on the first syllable of de’sukara is deleted (Poser, 1984) or at least phonetically reduced (Kubozono, 1993; Maekawa, 1994). This results in a gradual F0 fall over the AP containing the target word and the following conjunction. In the final condition (Figure 1, bottom panel), on the other hand, the target word does not phrase with the de’sukara, and is followed by an IP boundary. This condition is signaled by lower F0 on the target syllable due to the L% associated with the right edge of the phrase, and the realization of the pitch accent on the first syllable of de’sukara as well as the absence of downstep on that syllable, which is signaled by the F0 height of the first syllable of de’sukara being as high as that of the accented syllable of the target word.

Example stimuli in both medial (top panel) and final (bottom panel) conditions in Experiment 1. Step 4 from the continuum is shown, with a vowel duration of approximately 110 milliseconds. Spectrograms (0–5 kHz range) overlaid with pitch tracks (50–200 Hz range) are shown. The black boxed region highlights the second syllable of the target word and the post-target syllable /de/; note this is the only difference across conditions. X-JToBI labels are given for tonal events, and below, glosses are bracketed according to their phrasing, where [. . .]accentual phrase (AP) indicates an AP boundary and [. . .]intonational phrase (IP) indicates an IP boundary.
2.2 Materials
Stimuli were created by resynthesizing the speech of a ToBI-trained male speaker of the Tokyo dialect of Japanese (author Hironori Katsuda). The speaker was first recorded at 44.1 kHz in a sound-attenuated booth, using an SM10A Shure™ microphone and headset. Stimulus manipulation was carried out in Praat (Boersma & Weenik, 2019), using the Pitch Synchronous Overlap and Add method (Moulines & Charpentier, 1990).
The starting points for the manipulation were a production in which the target was phrased IP-medially as in the top panel of Figure 1, and one in which it was phrased IP-finally, as in the bottom panel. As outlined above, only the F0 on two syllables varied across conditions: the second syllable in the target word; and the following syllable /de/. The remainder of the carrier phrase was acoustically identical across conditions, including the duration of acoustic silence following the target word (likely attributable mostly to the stop closure for following /d/), as shown in Figure 1. The duration of this interval between the end of the target word and the beginning of the following syllable /de/ was approximately 50 milliseconds (ms) in duration. As it is argued in Venditti (2005), a pause is not obligatory for Japanese listeners to perceive a disjuncture equivalent to an IP boundary. In spontaneous speech, there are many cases in which a large degree of disjuncture is solely cued by a boundary tone without an intervening pause. Likewise, a pause can be present without cueing large disjuncture. Thus, this relatively short silent interval is compatible with both phrasings.
The starting point for F0 manipulations in Experiment 1 was a naturally produced IP-medial phrasing, as in (b), with a phonemically long vowel target (shi’shoo). F0 from another medial production of the second syllable in the target word, and the post-target syllable /de/ was resynthesized onto these two syllables in the carrier phrase, to create the medial condition, shown in the boxed region of the top panel of Figure 1. To create the final condition, we overlaid the second target syllable and post-target syllable /de/ with F0 values from a natural IP-final production (Figure 1, bottom panel). In this way, F0 was resynthesized on the crucially differing syllables in both conditions. These manipulations were judged to sound like a natural medial and final phrasing by a ToBI-trained native speaker of Japanese (author Hironori Katsuda). Subsequently, a vowel duration continuum was resynthesized from both medial and final conditions. The duration of the starting vowel was approximately 100 ms. It was manipulated to range from 60 to 180 ms of vowel duration. This manipulation was accomplished by linear compression and expansion of the target vowel material such that the entire vocalic portion was compressed or expanded. The continuum had eight evenly spaced steps that included these endpoint values, with a between-step durational difference of approximately 17 ms (note that all continuum steps were created via resynthesis—the unaltered original was not used). Spectrograms of continuum endpoints are shown in more detail in Figure 2. These manipulations resulted in 16 unique stimuli (two positional conditions × eight continuum steps), with the only difference across conditions being the F0 on the second syllable of the target and post-target /de/ (shown in Figure 1). By varying only F0 we ensured that changes in adjacent segmental duration are ruled out as a possible explanation, as discussed in by Mitterer et al. (2016). Additionally, psychoacoustic influences of pitch on perceived duration, whereby high F0 and more dynamic pitch contours can lead to increased perceived duration, are unlikely to play a role given that these effects have been shown to occur only in isolated monosyllables (Van Dommelen, 1993), and given that psychoacoustic effects also do not seem to occur when pitch has a possible prosodic interpretation (Steffman & Jun, 2019).

Spectrograms showing the continuum endpoints for the target word in Experiment 1 (the longest step at left, shortest at right). Time, marked by ticks on the x-axis, is indicated in 100 millisecond intervals.
2.3 Participants and procedure
Twenty-six participants (13 males and 13 females; mean age 28) were recruited for Experiment 1. All participants were native Japanese speakers from the greater Tokyo area. Participants provided informed consent to participate and were paid for their time.
During the experiment, participants were presented with audio stimuli binaurally via a PeltorTM 3MTM listen-only headset, while seated in a quiet room, in front of a laptop computer, in Tokyo. Participants were presented with orthographic representations of the target words on the laptop during audio presentation, 司書 representing shi’sho “librarian,” and 師匠 representing shi’shoo “master.” These words were each displayed centered on either half of the computer screen. The side of the screen on which each word appeared was counterbalanced across participants, that is, for 13 participants 司書 was on the left side of the screen and for 13 participants 師匠 was on the left side of the screen. Participants indicated their response via key press: an ‘f’ key press indicated the choice on the left side of the screen and a ‘j’ key press indicated a choice on the right side of the screen. Prior to the beginning of the test trials, participants completed eight practice trials, in which they heard each continuum endpoint, in each positional condition twice, in random order. During experimental trials, participants categorized 12 instances of each unique stimulus, in random order, for a total of 192 trials (eight continuum steps × two position conditions × 12 repetitions). All responses (excluding practice trials) were analyzed.
2.4 Results and discussion
Results were assessed statistically by a linear mixed-effects model with logistic linking function, implemented using the lme4 package in R (Bates et al., 2015). The model was fit to predict listeners’ response (shi’sho or shi’shoo) as a function of continuum step, position condition, and the interaction of these two fixed effects. The dependent variable was coded such that a long vowel (shi’shoo) response was mapped to 1, with a short vowel response mapped to 0. Therefore, a positive coefficient would represent an increase in long vowel responses. Position was contrast-coded, with
Following, for example, Barr et al. (2013), we specified the model random effect structure as by-participant intercepts with maximal random slopes, and no correlation between random effects. A model that included correlations did not converge, at which point the correlation parameter was removed. The converging model therefore contained de-correlated random slopes for both fixed effects and the interaction between them. Further simplification of random slopes led to decreased model fit, assessed by inspecting the Akaike information criterion and by likelihood-ratio tests (Matushek et al., 2017), suggesting that the fully specified random effect structure is justified. The model output is shown in Table 1, with results plotted in Figure 3.
Fixed effects from the model in Experiment 1. Values are rounded. Asterisks indicate p-values, where * = p < 0.05, ** = p < 0.01, and *** p < 0.001.

Categorization responses for Experiment 1, split by condition. The x-axis shows numbered continuum steps (step 1 = 60 milliseconds (ms), step 8 = 180 ms, and step intervals are approximately 17 ms). The y-axis shows the proportion of long vowel responses. Error bars around each point represent 95% confidence intervals.
Increasing vowel duration along the continuum (the “step” factor in Table 1), significantly increased long vowel responses (β = 6.08, z = 17.78). This is expected as outlined above. Position, the predictor of interest also showed a significant effect. As can be seen visually in Figure 3, the position of the target word impacted listeners’ categorization: an IP-final target showed significantly decreased shi’shoo responses (β = -0.39, z = -2.84). The interaction between the two fixed effects was not significant.
The effect of position found in the model supports our predictions: listeners required longer vowel durations for a long vowel (shi’shoo) percept when the target word was cued as PF. The results of Experiment 1 can therefore be taken to suggest that Japanese listeners rely on phrasal boundaries in their perception of CVL. Because only contextual F0 varied across conditions, we can conclude that this computation of this prosodic context is crucially informed by F0. This provides a new piece of evidence in the same vein as Steffman (2019b), and additionally shows that listeners use pitch-based cues alone to construe the prosodic boundary location relative to a target sound (as in Kim et al., 2018). Following these results for an accented target word, we can ask if an analogous pattern can be observed for a target that is unaccented. In doing so, we can see if the effect replicates, and generalizes to a different context, given that an unaccented target word engenders a different realization of contextual F0, outlined below.
3 Experiment 2
3.1 Methods
As in Experiment 1, participants categorized a sound from a vowel duration continuum as phonemically long or short. In Experiment 2, listeners categorized an unaccented disyllabic minimal pair as dookyo “housemate” (同居) or dookyoo “townmate” (同郷), which were comparable in log-transformed frequency based on word counts in the CSJ (1.62 for dookyo and 0.95 for dookyoo). Note these are both fairly low frequency as with the target words in Experiment 1. The same carrier phrase as Experiment 1 was used for both medial and final conditions.
3.2 Materials
As in Experiment 1, the stimuli for the medial and final conditions differ only in the F0 of the target syllable and that of the following syllable (i.e., /de/ in de’sukara “because/therefore”). Visual representations of both medial and final conditions for Experiment 2 are shown in Figure 4. In the medial condition (Figure 4, top panel), the target word x forms an AP with the following conjunction de’sukara. Unlike Experiment 1, however, the accent on the conjunction is realized, since the target word is unaccented. The medial condition is thus characterized by the higher F0 on the target syllable due to a lack of the L% boundary tone, as well as the absence of pitch reset on the following syllable. In the final condition (Figure 4, bottom panel), in which the target word does not phrase with the following conjunction, the F0 on the target syllable is lowered due to the presence of the L% associated with the right edge of the AP. Furthermore, the F0 on the following syllable is slightly higher in the final condition than in the medial condition, reflecting a pitch reset, which cues the presence of an IP boundary after the target word. In this way, the tonal context of the target is shaped by its accentual status, and we can therefore test of how the effect observed in Experiment 1 generalizes to unaccented words. Representative examples of the conditions in Experiment 2 are shown in Figure 4. Figure 5 shows spectrograms of continuum endpoints in more detail.

Example stimuli in both medial (top panel) and final (bottom panel) conditions in Experiment 2. Step 4 from the continuum is shown, with a vowel duration of approximately 110 milliseconds. Spectrograms (0 5 kHz range) overlaid with pitch tracks (50–200 Hz range) are shown. X-JToBI labels and bracketed glosses are shown, as in Figure 1. The boxed region highlights the second syllable of the target word, and post-target /de/.

Spectrograms showing the continuum endpoints for the target word in Experiment 2 (the longest step at left, shortest at right). Time, marked by ticks on the x-axis, is indicated in 100 milliseconds’ intervals.
The carrier phrase used for the Experiment 2 stimuli was identical to that used for Experiment 1, with one exception: the F0 on the post-target syllable /de/. Because of the different realization of accent on this word described above, the F0 on this syllable was resynthesized to match with natural productions following both an IP-final and IP-medial unaccented target word produced by the same speaker in the same carrier phase. All other parts of the carrier phrase were identical to the carrier phrase in Experiment 1. The starting point for the manipulation of the target was a phonemically long vowel target dookyoo, produced in IP-medial position, as in Experiment 1. This vowel was approximately 100 ms in duration. F0 from another medial production was resynthesized onto the second syllable of the target word, as in Experiment 1 (to ensure both conditions were equally resynthesized). This medial target word was cross-spliced into the medial frame used in Experiment 1 (with different F0 on post-target /de/ as outlined above), shown in the top panel of Figure 3. To create the final condition target, the second syllable of the naturally produced (not resynthesized) medial target was overlaid with the F0 from an IP-final production, and cross-spliced (with the first syllable) into the IP-final carrier sentence, shown in the bottom panel of Figure 3. Following this, a vowel duration continuum with the same endpoint durations and inter-step intervals as in Experiment 1 was resynthesized from both conditions, resulting in a total of 16 unique stimuli (two positional conditions × eight continuum steps). By virtue of changing F0 on post-target /de/, the new carrier phrase was now judged to sound natural for both phrasings of an unaccented target word, while remaining fairly comparable to the carrier phrase used in Experiment 1.
3.3 Participants and procedure
Twenty-six different participants (15 males and 11 females; mean age 26) were recruited for Experiment 2. Unlike Experiment 1, for logistical reasons, participants in Experiment 2 were tested in two different locations. 14 were tested in Tokyo, Japan, and 12 were tested in Los Angeles, California. Participants tested in Los Angeles have been living in the United States for an average of one year. Given that all participants were native speakers of Tokyo Japanese, we did not expect testing location to impact their performance. However, to be sure, we ran a preliminary model with the same structure as that used in Experiment 1 which additionally included location (Los Angeles/Tokyo) as a predictor, as well as its interaction with position (medial/final). The model revealed, as expected, that location did not impact responses overall, and that crucially, it did not interact with position (suggesting an analogous effect of position regardless of the country in which participants were tested). The model we report below does not include location as a predictor. The procedure was identical to that in Experiment 1.
3.4 Results and discussion
The model specifications used to assess the Experiment 2 results were the same as that in Experiment 1, with a long vowel (dookyoo) response mapped to 1. The same full random effect structure, with no correlation between random effects, was specified in the model. As in Experiment 1, this fully specified random effect structure was observed to provide the best model fit in comparison to simplified variants, and so was retained. The model output is given in Table 2. The results are plotted in Figure 6.
Fixed effects from the model in Experiment 2. Values are rounded. Asterisks indicate p-values, where * = p < 0.05, ** = p < 0.01, and *** p < 0.001.

Categorization responses for Experiment 2, split by condition. The x-axis shows numbered continuum steps (step 1 = 60 milliseconds (ms), step 8 = 180 ms, and step intervals are approximately 17 ms). The y-axis shows the proportion of long vowel responses. Error bars represent 95% confidence intervals for each point.
As would be expected, increasing vowel duration increased listeners’ dookyoo responses (β = 6.08, z = 16.49). As shown in Figure 6, position also had a significant effect: an IP-final target showed significantly decreased dookyoo responses (β = -0.66, z = -6.42). Unlike Experiment 1, a significant interaction (p = 0.045) between these two main effects was observed. This stems from the fact that the medial condition shows a sharper increase in dookyoo as a function of increasing vowel duration, as compared to the final condition, that is, the effect of continuum step is (slightly) larger in the medial condition (cf. Lunden, 2013). The main effect of position observed in the model concurs with our predictions in showing that listeners required longer vowel durations for a dookyoo response in the final condition, effectively requiring longer vowel durations to perceive a vowel as phonemically long when PF. This can therefore be taken to replicate the effect seen in Experiment 1, showing that accent-dependent pitch patterns modulate listeners’ perception of final lengthening for both accented, and unaccented words. The effect in both experiments aligns with the predictions laid out above.
One observation across Experiments 1 and 2 is that the magnitude of the effect is larger in Experiment 2 (β = -0.66, standard error (SE) = 0.10) as compared to Experiment 1 (β = -0.39, SE = 0.14), suggesting a possible sensitivity to accent in the effect of boundary, where the boundary effect is larger for an unaccented target word (though clearly robust in both cases). The present experiments only allow us to speculate in this regard, because in addition to varying in accentedness across experiments, target words also varied in moraic and segmental make-up (e.g., short /i/ precedes the target in Experiment 1 and long /oo/ precedes it in Experiment 2). Variation in acoustic context (including F0) across experiments therefore makes isolating the effect of accent on the magnitude of the positional effect impossible. Nonetheless, with the goal of motivating future research, we can consider how an accent-driven asymmetry might be explained by recent findings which show asymmetrical final lengthening in initial-accented versus unaccented words.
Seo et al. (2019) found that disyllabic words with an initial accent (e.g., ta’ka, a proper name) exhibited less lengthening on their final syllable compared to disyllabic words without an accent (e.g., taka “hawk”). Reduced final lengthening for words with a preceding accent was described as a suppression of pre-boundary lengthening by the authors and was hypothesized to serve the function of maintaining a syntagmatic contrast between the final rhyme and preceding accented syllable, consistent with previous studies which have shown an interplay between PF lengthening and prominence in other languages (cf. Turk and Shattuck-Hufnagel, 2007 for English; Katsika, 2016 for Greek; and Nakai et al., 2009 for Finnish). One hypothesized perceptual consequence of this pattern in Japanese is that listeners may show increased sensitivity to a phrasal boundary’s influence for unaccented words, which undergo more substantial lengthening (Experiment 2). In comparison, accented words may show a relatively limited effect of phrasal position, reflecting relatively limited PF lengthening (Experiment 1). If we interpret the results along these lines, they could be taken to reflect an interplay between the prominence marking and boundary marking systems of the language. 3
Given the other differences between target words across experiments, we cannot conclude anything concrete in this regard. Future research will accordingly benefit from addressing the possible involvement of pitch accent as a mediating factor for the observed boundary effect directly. Finding evidence for prominence-mediated boundary effects in perception would enrich our understanding of the amount of detail encoded in prosodic representations (e.g., reduced temporal expansion based on accentedness), and how different facets of prosodic organization interact in this domain.
4 General discussion
Taken together, Experiment 1 and Experiment 2 provide evidence that listeners compute prosodic boundary information in processing CVL and use this information to guide their interpretation of duration. Specifically, we found that a PF target required longer vowel duration to be categorized as phonemically long. We took this effect to reflect the influence of intonational structure in listeners’ perception of segmental contrasts such that PF sounds are expected to be lengthened, and longer vowel duration is therefore required for a long vowel percept.
Generally speaking, we can take this result to provide additional evidence for the relevance of prosodic factors in listeners’ perception of segmental contrasts cross-linguistically, and to highlight the continued importance of extending recent work on this topic (e.g., Kim et al., 2018; Mitterer et al., 2019).
The effect of phrasal position, though reliable in both experiments, is fairly small, particularly in Experiment 1. Shifts in categorization are such that listeners’ responses are impacted only for several steps on the continuum, and only to the extent that they increase slightly, showing a minimal rightwards shift in the categorization function. Though the effect is larger in Experiment 2, it remains fairly restricted and localized at the most ambiguous continuum steps. It can also be noted generally that prosodic effects from previous studies discussed above seem to be fairly subtle, and to engender small adjustments in categorization and online processing. In light of this, the restricted nature of the effects seen here could relate to the way in which listeners make use of prosodic information in speech processing, discussed below. The small effect size might also originate from the fact that only one cue to boundary (F0) was manipulated, and other cues, such as a pause, or post-boundary initial strengthening, are lacking. Seeing if additional boundary cues create a larger effect would be a useful further direction, especially given that we see a robust shift in categorization in this conservative test case, where temporal context is totally controlled. As an example, the presence of a post-target pause could be crossed with tonal manipulations in future studies. Though this would introduce variation in temporal context surrounding the target, it would allow us to test if a stronger boundary percept, indexed by larger shifts in categorization, obtains when boundary cues combine (cf. Mitterer et al., 2016; Nakai & Turk 2011). More generally, looking cross-linguistically to test how different boundary tones, or boundary tones in combination with other cues such as glottalization (in American English; e.g., Redi & Shattuck-Hufnagel, 2001) will help better our understanding of what informs perception of a prosodic boundary for the purposes of guiding segmental interpretation.
We can consider these results in light of two sets of previous findings outlined above. First, in regards to previous work on PF effects (Nootebaum & Doodeman, 1980; Steffman, 2019b), the present experiments offer additional evidence for the relevance of right-edge temporal patterns in perception cross-linguistically. They also offer some controls which were not present in these previous studies: by holding contextual duration constant across conditions, any possible confounds related to speech rate normalization are removed, as discussed above. With position cued only by pitch we can also complement these previous findings in highlighting the importance of intonational structure and its relation to phrasal boundaries for the purpose of informing listeners’ perception of prosody. More generally, we can take the present results to suggest that tonal cues play an important role, in line with the status of tonal events encoded in models of the intonational phonology of Japanese. These findings also suggest future work that is informed by models of intonational phonology may help shed light on the sorts of structures which listeners compute, and the cues that specify them.
We can also consider these results in terms of recent proposals for parallel processing of segmental and prosodic structures, discussed above. In this light, these results can be taken as another piece of evidence for proposed “prosodic analysis” in which a prosodic representation guides listeners’ interpretation of segmental contrasts. As discussed above, both Kim et al. (2018) and Mitterer et al. (2019) provide evidence for this sort of role for prosodic boundaries, both in perception of phonetic detail (e.g., phrase-initial glottalization) and in phonological inferencing. Time-course evidence from these studies, not discussed above, showed that these effects occur relatively late in processing, which may be taken to suggest they involve later-stage modulation of lexical competition, after segmental material activates lexical hypotheses (Cho et al., 2007). The involvement of prosodic structure in a later stage of processing might explain its relatively limited influence on categorization observed in the present experiments. That is, prosodic structure may exert a role only in the more ambiguous cases (i.e., when both possible word forms are fairly equally activated, see Newman et al., 1997 for discussion of lexical activation/competition in a forced choice perception task), and then only in a non-deterministic fashion (Cho et al., 2007). In this regard, prosodic boundary information may be integrated in perception in a secondary fashion as compared to, for example, vowel internal durational cues. The present results, being purely offline, cannot speak to this timing prediction. Accordingly, one promising extension would be to test the time-course of the observed effects with eye-tracking (using a similar paradigm as e.g., Kingston et al., 2016; Reinisch & Sjerps, 2013). Seeing if these prosodically-guided effects occur with a delayed time-course, and comparing these effects to that of vowel duration along the continuum, a segment-internal cue that is used rapidly (Reinisch & Sjerps, 2013), might offer a window into the processes that underly our observed effect. A later time-course would be a strong argument in favor of this proposed later-stage prosodic analysis. Extending the present findings in this way may thus offer useful converging evidence for this idea. These results nevertheless may tell us something informative about the role of prosodic analysis in speech perception, mainly the importance of intonational cues in guiding the perception of phonetic detail. As discussed above, Kim et al. (2018) showed that intonational structure, cued only by pitch, seems to play an important role in phonological inferencing. The present results would suggest that computation of intonational structure similarly influences processing of phonetic detail as well. Together, recent findings (Kim & Cho, 2013; Kim et al., 2018; Mitterer et al., 2016, 2019; Steffman, 2019a, 2019b) would suggest that prosodic boundaries, of various types, and cued by various means, merit further research as a mediating factor at multiple levels of speech processing.
In addition to exploring these effects with online measures, other extensions of our findings will benefit from testing our speculation related to the role of accent, outlined above. Finding evidence for a mediating effect of prominence in this sort of task would point to an interplay between these two facets of prosodic organization in perception. More generally, finding other test cases which allow for researchers to exploit intonational patterns will also enrich the present findings, and could be used as a test for the relevance of various intonational properties, or claimed functions of intonational tunes in a given language, for language perceivers. Extending the present findings along these lines will better our understanding of the role of intonation in language comprehension, and will help inform a theory of the processes which underpin it.
Footnotes
Acknowledgements
We are grateful to all of our participants for their time, and to Sun-Ah Jun, Megha Sundara, Pat Keating and three anonymous reviewers for helpful feedback and commentary. Further thanks to Shigeto Kawahara, Mami Gosyo, and Naoki Ishikawa for recruitment assistance. A previous version of this work was presented at the 10th International Conference on Speech Prosody.
Funding
This research was funded by the University of California, Los Angeles Ladefoged Scholarship, awarded to Hironori Katsuda.
