Abstract
Adults are able to use visual prosodic cues in the speaker’s face to segment speech. Furthermore, eye-tracking data suggest that learners will shift their gaze to the mouth during visual speech segmentation. Although these findings suggest that the mouth may be viewed more than the eyes or nose during visual speech segmentation, no study has examined the direct functional importance of individual features; thus, it is unclear which visual prosodic cues are important for word segmentation. In this study, we examined the impact of first removing (Experiment 1) and then isolating (Experiment 2) individual facial features on visual speech segmentation. Segmentation performance was above chance in all conditions except for when the visual display was restricted to the eye region (eyes only condition in Experiment 2). This suggests that participants were able to segment speech when they could visually access the mouth but not when the mouth was completely removed from the visual display, providing evidence that visual prosodic cues conveyed by the mouth are sufficient and likely necessary for visual speech segmentation.
Keywords
1 Introduction
A fundamental feature of the language learning environment is noise; consequently, mechanisms designed to extract structural representations from this environment must do so from a degraded signal. Fortunately, speech is an inherently multimodal process (Rosenblum, 2008) with complementary cues in the voice and face of the speaker (Massaro, 1998) that help learners overcome the degradation of speech input. Listeners routinely utilize facial cues to facilitate speech perception (e.g., Sumby & Pollack, 1954), particularly when the speech signal is obscured by noise (Grant & Seitz, 2000). Furthermore, facial cues enhance a wide array of speech processes supporting language acquisition. For example, faces can serve to help learners distinguish between speech sounds (e.g., Kuhl & Meltzoff, 1982) and refine phonemic representations (e.g., Teinonen et al., 2008). One particular speech process that is critical for early stages of language development is speech segmentation, and recent evidence has suggested that faces may play an important role in this process (e.g., Hollich et al., 2005; Mitchel & Weiss, 2010).
Speech segmentation is the process of isolating individual words or phrases from a continuous speech stream. As pauses in speech do not reliably occur at word boundaries (see Saffran, 2003), listeners must rely on a suite of distributional and acoustic cues to infer where one word ends and another begins. These segmentation strategies have been studied extensively, and one common method has been to test the ability of adults to use specific cues to segment an artificial language. In an artificial language learning paradigm, participants are passively exposed to a continuous speech stream comprising nonsense words and then tested on their ability to discriminate words from non-words. Typically, all segmentation cues other than the one(s) being tested are removed from the speech stream. Artificial language learning studies are therefore well-situated to serve as a “test-bed” for segmentation strategies by allowing researchers to isolate individual segmentation strategies (see Cutler, 2012). Using this approach, several different cues in the auditory domain have been identified, including lexical stress patterns (e.g., Tyler & Cutler, 2009), prosodic contours (Shukla et al., 2007), coarticulatory patterns (e.g., Fernandes et al., 2007), and transitional probabilities between adjacent sounds (e.g., Saffran et al., 1996). However, given that the auditory signal is often obscured by noise (Hollich et al., 2005), there has been a growing interest in how cues in the visual domain, particularly cues provided by the speaker’s face, may complement these auditory-based segmentation strategies (e.g., Mitchel & Weiss, 2010; Sell & Kaschak, 2009).
Much of the work in this field has investigated how visual cues enhance or supplement auditory cues (e.g., Cunillera et al., 2010); however, Mitchel and Weiss (2014) recently extended this line of inquiry to determine whether facial cues might provide a unique source of segmentation cues independent from the auditory signal. In the auditory signal, prosodic patterns can be used to identify word boundaries (Sanders & Neville, 2000; Smith et al., 1989). For example, Smith et al. (1989) found that adult English speakers were able to use durational cues to identify word boundaries in masked speech. Moreover, these prosodic elements (e.g., stress and pitch) are mirrored in the face of the speaker (Yehia et al., 2002) and listeners are readily able to match visual and auditory prosodic representations (Cvejic et al., 2012); thus, Mitchel and Weiss (2014) tested whether it was possible to use visual prosodic cues to segment speech. Participants in this study viewed a continuous, audiovisual speech stream that contained a miniature artificial language. Critically, participants were not able to use auditory cues to segment the speech stream; yet, when the language was paired with a synchronous talking face display produced by an actor who knew the position of word boundaries, participants were able to use subtle visual prosodic cues to identify words in the auditory stream.
Critically, it is not the case that any dynamic, synchronous visual display supports visual speech segmentation. Mitchel and Weiss (2014) created a second synchronous display with an actor who was unaware of word boundaries (having read from a list of part-words). Segmentation performance was significantly poorer in this “misinformed” condition relative to the “aware” audiovisual condition and was not significantly different from performance in the audio-only condition. Although the misinformed visual display contained phonetic detail (e.g., lip rounding), the visual prosodic cues were disrupted, suggesting that the successful segmentation in the aware condition was due to visual prosodic cues rather than enhanced attention to phonetic detail. Mitchel and Weiss (2014), therefore, provided the first demonstration that visual prosody is an independent source of segmentation cues, joining the myriad auditory cues that learners can use to overcome this early perceptual challenge.
Although Mitchel and Weiss (2014) demonstrated that it is possible to visually segment speech, it is unclear which facial cues support speech segmentation. Prior research has identified a suite of facial cues that that signal prosody in speech (Cvejic et al., 2012; Graf et al., 2002; Kim et al., 2014; Lansing & McConkie, 1999; Munhall, Jones, et al., 2004; Swerts & Krahmer, 2008; Yehia et al., 1998, 2002). These facial cues include articulatory gestures (i.e., movement of the lips and jaw), co-speech gestures such as eye and eyebrow movements, and global movements of the head (e.g., head nods). However, there is considerable disagreement across studies as to the relative prominence of these cues for conveying visual prosody, and much of that variability appears to be related to the particular task demands (Lansing & McConkie, 1999). For example, when identifying word content, observers appear to rely on oral and extraoral movements in the lower face; yet, when identifying phrase-level prosodic contours, observers’ gaze was more broadly distributed (Lansing & McConkie, 1999) and cues in the upper portion of the face play a greater role (Swerts & Krahmer, 2008). Although there are multiple markers of visual prosody in the speaker’s face, these cues provide different types of information and their utility therefore depends on the nature of the task.
To investigate which facial cues support visual speech segmentation, Lusk and Mitchel (2016) used an eye tracker to detect eye gaze patterns during online processing of the visual speech display, using the familiarization stream and test from Mitchel and Weiss (2014). Under the Gaze Direction Assumption (Lansing & McConkie, 2003), it was proposed that participants would fixate longest upon features that were most informative to the segmentation task. Based on prior research (e.g., Green et al., 2010; Yehia et al., 2002), the authors predicted that prosodic cues in the mouth region (e.g., vertical and horizontal lip aperture) might reliably signal word boundary locations and thus participants would predominantly fixate upon the mouth. Consistent with this prediction, participants spent the greatest amount of time fixating on the mouth region, compared with the nose and eye regions. The results of Lusk and Mitchel (2016), therefore, provide evidence that prosodic cues conveyed in mouth movements support the process of visual speech segmentation.
Although this study provided valuable insight into the processing strategies employed by observers, it remains to be seen what the relative contribution of each feature is to learning word boundaries, as well as the extent to which these cues are necessary or sufficient for visual speech segmentation. That is, if either the mouth or eyes (or both) were occluded, would visual speech segmentation still be possible? In this study, we investigate this question by systematically removing (Experiment 1) and then isolating (Experiment 2) the eyes, mouth, and frame of the speaker’s face. Based on the results of Lusk and Mitchel (2016), we predict that obscuring the mouth should preclude successful segmentation. Likewise, if the segmentation cues are encapsulated in the mouth region, then access to only this feature (i.e., the mouth in isolation) should still facilitate segmentation. Alternatively, if cues in the oral region are not sufficient or cues outside that region are also necessary, then reducing the facial display to any single feature should inhibit learning.
2 Experiment 1
In Experiment 1, we occlude individual features to determine if their removal affects performance on a visual speech segmentation task. Based on the results of Lusk and Mitchel (2016), we focus on the following two features: the eyes and the mouth of the speaker. We are also interested in whether cues outside of those two regions (e.g., head nods) were contributing to successful segmentation; thus, we also include a condition where we remove the rest of the face and leave only the eyes and mouth.
2.1 Method
2.1.1 Participants
In total, 141 undergraduate students at Bucknell University completed the study for course credit. Participants were monolingual English speakers. Of this original set of participants, three were removed from analysis due to technical error (equipment or software malfunction) and 10 were removed due to participant error (e.g., failure to follow directions, indicating low effort/attention on posttest questionnaire). Thus, the final number of participants included in the analysis was 128: 30 (24 female, 6 male) in the audio-only condition; 33 (24 female, 9 male) in the no mouth condition, 34 (25 female, 9 male) in the no eyes condition, and 31 (23 female, 8 male) in the no frame condition.
2.1.2 Stimuli
To create our stimuli, we modified the “new aware” familiarization stream from Mitchel and Weiss (2014), which was also used in Lusk and Mitchel (2016). Below is a brief overview of the design and creation of this base stimulus, but for additional details on stimuli creation, please see Mitchel and Weiss (2014).
The base stimulus was a continuous audio stream paired with a synchronous, informative facial display. The audio stream comprised six CV syllables that were combined into six trisyllabic words (/bo ke tɑɪ/, /pu tɑɪ bo/, /ke ɡi dɑ/, /dɑ pu ɡi/, /ɡi bo pu/, /tɑɪ dɑ ke/) concatenated into a single continuous loop. The syllables making up the words were created by recording a male speaker producing each syllable with a coda consonant in each place of articulation (e.g., “bo” was recorded as /bob/, /bod/, and /bog/) to provide coarticulatory place cues for the subsequent syllable. Each syllable was hand-edited in Praat (Boersma & Weenink, 2011) to remove the release burst of the final consonant, normalize intensity, and equate acoustic parameters (e.g., aspiration, vowel duration, pitch tier, etc.), removing any acoustic cues to word boundaries. The average duration of syllables was 210 ms, the average pitch was 132 Hz, and the average intensity was 71 dB. The syllables were then concatenated into words, which were in turn concatenated into a continuous 4-min loop. This 4-min loop was repeated three times for a total familiarization of 12 min. The words were combined in a manner to minimize statistical cues to word boundaries, such that within-word transitional probabilities were 0.33 and between-word transitional probabilities were 0.11. When presented in isolation, participants in two prior studies were unable to successfully segment this speech stream (Lusk & Mitchel, 2016; Mitchel & Weiss, 2014).
This audio stream was then paired with a synchronous visual display of a male actor lip-syncing to the audio stream. The actor was recorded lip-syncing to one loop of the audio stream (18 words), while reading from a list of words. This video loop was then manually synchronized with the audio stream using Adobe Premiere, creating a short audiovisual movie clip lasting 15 s. A fade from/to black was added to the beginning and ending of the clip to remove any movement artifacts when creating the larger loop. This clip was repeated and concatenated 16 times to produce a 4-min movie. Critically, the actor was aware of word boundary locations at the time of the recording; thus, it was inferred that he would impart visual prosodic cues to word boundaries. The presence of visual cues was previously confirmed in a visual-only test, in which participants were more likely to label visual segments as words if those segments were aligned rather than unaligned with auditory word boundaries (Mitchel & Weiss, 2014). When viewing an audiovisual stream, participants were able to detect these cues and use them to segment the auditory speech stream, resulting in above-chance identification of words during an audio-only post-test (Lusk & Mitchel, 2016; Mitchel & Weiss, 2014).
To create the stimuli for this study, we used Adobe Premiere software to occlude or remove select features in the base stimulus (taken from previous studies). We created three separate movies to produce each of the familiarization conditions: no eyes, no mouth, no frame. To create the no eyes and no mouth conditions, we created a static black rectangle (approximately 7.69° by 5.31°) as a layer in Adobe Premiere and positioned it in front of the base stimulus video. The rectangle was positioned such that both eyes and top of the nose was occluded in one video (no eyes) and the mouth, chin, and bottom of the nose were occluded in a second video (no mouth; see Figure 1). The size and shape of the occluder was chosen to ensure that, in the no mouth condition, the mouth remained fully occluded throughout the duration of the video. Because of head movement in the actor, as well as lower jaw movement while speaking, the occluder needed to cover a larger area of the video. As we wanted to keep the area occluded constant, the size and shape of the occluder was the same in both videos; only its placement changed. Thus, in the no eyes condition, the nose is also occluded while the mouth and frame are both visible. However, in Lusk and Mitchel (2016), the total gaze duration on the nose region was minimal, and the nose was viewed significantly less than both the eyes and mouth. To create the no frame video, we used the cropping tool in Premiere to remove all portions of the video other than the eyes and mouth (see Figure 1). The total area of the eyes visible was approximately 11.16° × 4.48°. The total area of the mouth visible was approximately 6.71° × 2.80°. The rest of the video outside of those regions was set to black, and the eyes and mouth regions remained in the same position as in the full video.

Still frames of the three audiovisual familiarization conditions in Experiment 1: no mouth, no eyes, no frame.
The post-familiarization test was identical to the one used in both Mitchel and Weiss (2014) and Lusk and Mitchel (2016). The stimuli for the test were audio items taken from the speech stream, and included the six words as well as six part-words created by combining the third syllable of one word with the first and second syllables from a different word. Part-words were therefore familiar items, yet inconsistent with visual prosodic cues to word boundaries. If participants were able to use these cues to segment the speech stream, then we expected them to endorse the words over the part-words.
2.1.3 Procedure
Participants were randomly assigned to be in one of the three familiarization conditions. We used E-Prime 2.0 software to present all stimuli and record participants’ responses. Each testing session consisted of a familiarization and testing phase. During familiarization, participants passively viewed the familiarization movie. Each movie consisted of three 4-min blocks separated by 1-min breaks, for a total familiarization of 12 min. Participants wore noise-canceling headphones throughout the duration of the experiment. Participants were instructed to watch the movie carefully, keep their headphones on at all times, and that they would be tested on what they learned following the movie.
Following familiarization, participants completed an audio-only 2afc test pitting words against part-word foils. During each test trial, one word and one part-word were sequentially presented (order counterbalanced across trials). The participant was then asked to indicate through keypress which item was most consistent with a word, based on the movie that they had just viewed. Each word was paired with each part-word; thus, there were 36 test trials. The total number of trials in which the word was correctly selected was recorded for each participant. The trial order was randomly generated for each participant. The same test was used for all familiarization conditions. At the end of the experiment, participants completed a demographic questionnaire and were asked to estimate their effort and attention to the video on a 10-point scale.
2.2 Results and discussion
The mean number of correct responses (with standard deviation) for each condition was: 17.83 (2.67) out of 36, or 50% in the audio-only baseline, 20.12 (3.25) or 56% in the no eyes condition; 20.36 (3.44), or 57% in the no mouth condition; and 20.61 (3.25), or 57% in the no frame condition (see Figure 2). As a point of reference, the mean number of correct response for the full audiovisual display in Lusk and Mitchel (2016) was 20.13, or 56%. We first conducted a one-way ANOVA to compare performance across conditions and found a statistically significant difference in learning, F(3, 124) = 5.31, p = .002, η2 = .114. Post hoc comparisons with a Bonferroni correction revealed that accuracy in the audio-only condition was significantly less than the three audiovisual conditions (vs. no eyes: p = .022, 95% CI = [–4.35, –0.22]; no mouth: p = .009, 95% CI = [–4.61, –0.45]; no frame: p = .003, 95% CI = [–4.59, –0.67]); however, none of the audiovisual conditions were significantly different from each other in the percentage correct at test (all p-values > .05). Next, separate one-sample t-tests revealed that performance in each audiovisual condition was significantly greater than would be expected by chance (50%): no eyes, t(33) = 4.32, p < .001, Cohen’s d = 0.74; no mouth, t(32) = 3.95, p < .001, d = 0.69; no frame, t(30) = 4.47, p < .001, d = 0.80. In contrast, performance in the audio-only condition did not exceed chance, t(29) = –0.34, p = .735, d = –0.06.

Mean accuracy on the audio-only posttest for each familiarization condition in Experiment 1.
Surprisingly, we did not find a significant effect of obscuring either the mouth or the eyes on performance in a visual speech segmentation task. In addition, restricting cues to only the eyes and mouth similarly did not affect learning. Based on the eye-tracking results from Lusk and Mitchel (2016), we had suggested that the cues in the oral region conveyed the principal visual prosodic cues supporting segmentation; thus, we predicted that occluding the mouth would impede participant’s ability to segment the speech stream. One possible explanation for why occluding the mouth did not impede performance is that visual prosodic cues are more broadly distributed across the facial regions. As summarized earlier, visual prosody is realized across an array of facial features, including features in both the eye and mouth region (see Cvejic et al., 2012); thus, it is possible that a number of different cues could contribute to visual speech segmentation. Alternatively, it is possible that participants were able to reconstruct mouth movements in all three of the familiarization conditions. As we only occluded a single feature at a time, participants had access to either the mouth or the frame in each of the conditions. A prior study tested the effect of occlusion on visual speech perception and found that synchronous facial displays enhanced identification of congruent auditory speech and impaired recognition of incongruent auditory speech (i.e., a McGurk effect) even when the entire lower half of the face was occluded (Jordan & Thomas, 2011). The authors proposed that access to visible areas of the face activated structural representations of the occluded regions. For example, as cheek and jaw movements correspond to mouth and lip movements, it is possible to infer cues in the mouth region even when it is occluded. In this study, participants may have used movements in the “frame” (i.e., cheek and lower jaw) to access visual prosodic cues conveyed by the lip gestures and mouth movements of the speaker. To test this possibility in Experiment 2, we isolate specific cues to determine whether any single cue is necessary or sufficient to enable successful speech segmentation.
3 Experiment 2
In Experiment 2, we investigate the ability of observers to extract information about word boundaries from visual displays in which a single feature (eyes, mouth, frame) has been isolated. If both the eyes and mouth independently provide sufficient information to support speech segmentation, then learning should be evident in the eyes only and mouth only condition, but not in the frame only condition. If, however, the mouth is the principle source of visual prosodic information necessary for visual speech segmentation and observers can access structural codes about the mouth from the frame, then learning should only be evident in the mouth only and the frame only conditions and performance in the eyes only condition should be disrupted.
3.1 Method
3.1.1 Participants
In total, 133 undergraduate students at Bucknell University completed the study for course credit. Participants were monolingual English speakers. Of this original set of participants, one was removed from analysis due to technical error (software malfunction) and 13 were removed due to participant error (e.g., failure to follow directions, indicating low effort/attention on posttest questionnaire). Thus, the final number of participants included in the analysis was 118:30 (19 female, 11 male) in the audio-only condition, 29 (22 female, 7 male) in the eyes only condition, 28 (21 female, 7 male) in the mouth only condition, and 31 (22 female, 9 male) in the frame only condition.
3.1.2 Stimuli and procedure
To create the familiarization movies for Experiment 2, we used the base stimulus from Lusk and Mitchel (2016; see above) and either cropped out or blurred all features except for one. In the eyes only and mouth only conditions, we used the crop tool in Adobe Premiere to isolate the eyes and mouth, respectively (see Figure 3). Each feature retained its size and position from the full video. For the frame only condition, we were unable to use the crop tool to remove the internal features of the face. Instead, we used the “gaussian blur” tool as well as the ellipse tool in Adobe Premiere. An ellipse was created to fit inside the frame of the face and then a blur was added to the ellipse. The ellipse then tracked the movements of the face using keyframes. This allowed us to mask the internal features of the face (i.e., eyes and mouth) while retaining access to the frame of the face. Participants were randomly assigned to one of these three familiarization conditions.

Screenshots of the three audiovisual familiarization conditions in Experiment 2: eyes only, mouth only, frame only.
All other aspects of the stimuli and procedure were identical to Experiment 1.
3.2 Results and discussion
The mean number of correct responses (with standard deviation) for each condition was as follows: 18.60 (2.55) out of 36, or 52% in the audio-only condition; 18.38 (2.94) out of 36, or 51% in the eyes only condition; 20.82 (3.38), or 58% in the mouth only condition; and 20.61 (3.37), or 57% in the frame only condition (see Figure 4). Performance in both the mouth only, t(27) = 4.42, p < .001, d = 0.83, and frame only, t(30) = 4.31, p < .001, d = 0.77, conditions were significantly greater than would be expected by chance (50%). However, performance in the audio only, t(29) = 1.29, p = .208, d = 0.24, and eyes only condition, t(28) = 0.69, p = .494, d = 0.13, was not significantly above chance. A one-way ANOVA found a significant main effect of condition on learning: F(3, 114) = 5.16, p = .002, η2 = .12. Bonferroni post hoc analyses confirm that performance in the eyes only condition was significantly below performance in the mouth only (mean difference = –2.44, p = .017) and frame only (mean difference = –2.23, p = .027) conditions but not different from performance in the audio-only condition. Performance in the mouth only and frame only conditions was not significantly different (mean difference = 0.21, p = 1.00), but each was significantly greater than in the audio-only condition (mean difference = 2.22, p = .025; mean difference = 2.01, p = .047, respectively).

Mean accuracy on the audio-only posttest for each familiarization condition in Experiment 2.
4 General discussion
In this study, we examined the impact of occluding various facial features on participants’ ability to segment speech from visual prosodic cues. In Experiment 1, we occluded individual features of the face (eyes, mouth, and frame). We predicted that removing access to the oral region of the face display would inhibit segmentation; however, we found that removing individual features had no discernable impact on visual speech segmentation. Performance was above chance in all three conditions and there were no differences in segmentation scores across the eyes removed, mouth removed, and frame removed (eyes and mouth only) familiarization conditions. Furthermore, the level of learning in all three conditions was consistent with full-face display conditions from previous studies (Lusk & Mitchel, 2016; Mitchel & Weiss, 2014). In Experiment 2, we tested the ability of each cue to support speech segmentation by isolating the individual cues, providing learners with an eyes only, mouth only, or frame only familiarization. We found that learning was only above chance following the mouth and frame only conditions. Performance in these conditions was again comparable to full-face display conditions in prior studies (e.g., Lusk & Mitchel, 2016). The eyes only condition resulted in at-chance segmentation performance that was significantly poorer than performance in the mouth only and frame only conditions and was not significantly different from the audio-only baseline.
The first important observation from our results is that information gleaned from the ocular region of the face is not necessary for visual speech segmentation. Participants successfully segmented speech in the absence of the eyes (i.e., mouth only and frame only conditions). Conversely, participants did not successfully segment the speech stream when restricted to the eyes (i.e., eyes only condition). Numerous studies have demonstrated that cues in the ocular region can effectively convey aspects of visual prosody (e.g., Kim et al., 2014; Swerts & Krahmer, 2008); however, the successful learning in the mouth only and frame only conditions indicates that participants need not rely on these particular cues during visual speech segmentation. Although the chance level of performance in the eye-only condition suggests that participants were unable to utilize cues in the eye region to segment speech, it is not clear from our results whether these cues were available in the visual display. Visual speech cues vary between tasks and speakers; thus, it is possible that this particular actor did not impart ocular word boundary cues when lip-syncing to the audio stream. Consequently, the at-chance performance could reflect either the absence of cues or the inability to utilize cues available in the display.
The second key observation was that information in the mouth region was sufficient to support speech segmentation. Segmentation performance in the mouth-only condition was significantly above chance and, moreover, was comparable to learning observed in previous full-face conditions. This provides evidence that cues in the oral region provide access to word boundary information that, isolated from other segmentation cues, can support speech segmentation.
Taken together, these observations are consistent with previous findings that although many facial regions provide prosodic information (see Cvejic et al., 2012), the reliance on particular facial regions appears contingent on the nature of the task and the type of information being extracted (Lansing & McConkie, 1999). Moreover, when attempting to identify word-level prosodic features, observers attend to the mouth and chin region (Lansing & McConkie, 1999), whereas phrase-level prosodic structure may be signaled to a greater extent by cues in the eye region (Swerts & Krahmer, 2008). In this study, participants segmented speech by identifying word-level rather phrase-level prosodic contours. In line with previous research (see Thomas & Jordan, 2004), therefore, we found little reliance on ocular cues; instead, participants made use of oral and extraoral cues to segment speech.
There is also some evidence that realizations of visual cues in the mouth may be particularly important for the comprehension of unfamiliar speech. Barenholtz and colleagues (2016) found that observers spent a greater amount of time fixating on the mouth region when viewing a speaker of an unfamiliar language as compared with a speaker of a familiar language, which the authors interpret as evidence the mouth provides a critical source of information for audiovisual speech encoding. This view is further supported by evidence of developmental changes in fixation preferences of infants across the first year of life, where an initial preference for viewing the eye region shifts to a preference for the mouth around the same time that infants begin to acquire language-specific representations of speech (Lewkowicz & Hansen-Tift, 2012). Once these representations stabilize and infants develop a degree of expertise in their native languages (~12 months), infants shift their gaze preferences back to the eyes for native language input but retain oral gaze preferences for novel language input (Lewkowicz & Hansen-Tift, 2012). This suggests that information cued by the mouth helps learners resolve uncertainty in the speech input. Our findings are consistent with this view, as the mouth was of critical importance for visual speech segmentation, which requires learners to deconstruct a continuous speech stream of an unfamiliar language into stable lexical units.
It is important to note that participants were also able to successfully segment the speech stream in the frame only condition. On the surface, the frame may convey an independent source of visual prosodic cues. One of the cues that tends to correspond with pitch and other prosodic elements is rotation about the x-axis, or head nodding (Graf et al., 2002; Munhall, Jones, et al., 2004). As x-axis rotations are visible in the frame of the face, it is possible that these are sufficient for speech segmentation. Our current data cannot rule out the possibility that the frame may provide an independent cue to word boundaries. However, there are a few reasons to suspect that the above-chance performance in the frame only condition was not a function of head nods. First, the videos were constructed in a way to specifically preclude head nodding as the actor was asked to keep his head affixed to a specific point on the wall behind him. Second, similar movements would also be visible in the eyes only condition (e.g., the relative height of the eyes and bridge of nose in the static video frame would signal x-axis rotation). In addition, a recent study provides evidence that head nods on their own are not sufficient for speech segmentation at the phrase level (de la Cruz-Pavía et al., 2019). In this study, monolingual and bilingual subjects were better able to segment continuous speech into phrase-level units when prosodic cues in the speech signal were accompanied by head nods from an animated face display. However, when the speech signal contained no auditory prosodic cues (as was the case in the present study), the visual display did not facilitate phrase-level segmentation. The authors of this study propose that “head nods are not interpreted as markers of prosodic prominence independent of auditory prosody” (de la Cruz-Pavía et al., 2019, p. 281). It is important to note, however, that this prior study investigated the segmentation of phrasal units, whereas this study investigated word-level segmentation; thus, it remains a possibility that head nods confer word-level segmentation cues as visual speech segmentation at different linguistic levels may rely on different facial cues.
Alternatively, participants in the frame only condition may have had direct or indirect access to the mouth that then enabled them to successfully segment the speech stream. First, the Gaussian blur tool used to create the frame only condition did not obscure low-spatial frequency information. Prior research has revealed that low-pass filtered videos can still enhance the intelligibility of speech in noise (Munhall, Kroos, et al., 2004) and quantized, visually degraded videos (similar to the blurred video in the frame only condition) can elicit a McGurk illusion (MacDonald et al., 2000). These studies indicate that audiovisual speech perception does not depend exclusively on fine details carried in high-spatial frequency bands. Information about lip aperture as well as dynamic motion of the mouth is still available in low-spatial frequency bands, and therefore, participants in the frame only condition of this study may have had sufficient access to segmentation cues in the oral region. To explore this possibility, future research could test visual segmentation with an opaque occluder over the center of the face. If learning dropped to chance in this condition, it would provide evidence that participants were relying on low-spatial frequency cues in the current frame only condition. Alternatively, if performance remained above chance, then this would lend support to the view that the frame provides an independent cue to word boundaries (e.g., head nods).
Although our results suggest that visual speech segmentation is achieved through cues in the oral region, there is some evidence that realizations of visual prosody vary across speakers (Cvejic et al., 2012; Scarborough et al., 2009). For example, Dohen and colleagues (2006) found that among five French speakers, the visual cues signaling broad or narrow focus varied across speakers, with some marking focus in eyebrow movements and others in head nodding. Although previous studies have tested the ability to visually segment speech produced by multiple actors (Mitchel & Weiss, 2014), this study utilized a single speaker across all conditions; thus, we cannot rule out the possibility that other actors may produce visual prosodic cues in other regions of the face (e.g., the eyes). Future studies on visual speech segmentation should explore this by comparing learning following familiarization to different speakers.
Our results contribute to a growing body of research demonstrating the contribution of facial cues to various aspects of language acquisition (e.g., Mani & Schneider, 2013; Teinonen et al., 2008; Weatherhead & White, 2017; Weikum et al., 2007). In particular, there is mounting evidence that visual cues, including talking face displays, aid learners during speech segmentation (Cunillera et al., 2010; Hollich et al., 2005; Mitchel & Weiss, 2010; Thiessen, 2010), suggesting that the learning mechanisms underlying this process may be tuned to operate over multimodal input (Mitchel et al., 2014; Mitchel & Weiss, 2011). Moreover, this study is yet another demonstration that adults can use visual cues, independent of the auditory input, to segment speech. Thus, visual prosody is among the suite of cues available to adult language learners to overcome this perceptual challenge. Future research will elucidate whether infants can similarly utilize these cues and at what point in the developmental trajectory these cues become available. Segmentation strategies are not static, and infants appear to prioritize different cues at various points within the first year of life (see Johnson & Jusczyk, 2001; Thiessen & Saffran, 2003). Given the ambiguity of speech and the early prominence of visible speech (Lewkowicz & Hansen-Tift, 2012), it is possible that visual prosodic cues may be a relatively early cue that could anchor later segmentation strategies.
In conclusion, this study provides evidence that visual prosodic cues provided by the speaker’s mouth are sufficient to enable successful speech segmentation. We further speculate that access to structural representations of the mouth (see Jordan & Thomas, 2011) are necessary for successful visual speech segmentation. Future work is needed to specify with greater resolution which mouth movements provide these prosodic cues (e.g., Vatikiotis-Bateson & Ostry, 1995) as well as the range of scenarios in which visual segmentation cues are utilized (e.g., do learners rely on these cues to a greater extent in noisy environments?). In addition, in line with research on auditory prosodic segmentation cues (see Tyler & Cutler, 2009), cross linguistic studies may provide insight into language-specific differences in the availability of visual prosodic cues. Although future work will endeavor to elucidate these issues, this study enhances our understanding of the contributions of visual signals in the speakers’ face to the mechanisms underlying language acquisition.
Footnotes
Acknowledgements
The authors thank Chip Gerfen and Kevin Weiss for their assistance creating the stimuli. They also wish to thank Rachel Bergin, Katie Lunceford, Chris Paine, and Sarah Shochat for aiding in data collection.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by a Swanson Fellowship to A. Mitchel.
