Abstract
Theories of speech perception agree that visual input enhances the understanding of speech but disagree on whether physically mimicking the speaker improves understanding. This study investigated whether facial motor mimicry facilitates visual speech perception by testing whether blocking facial motor action impairs speechreading performance. Thirty-five typically developing children (19 boys; 16 girls; M age = 7 years) completed the Revised Craig Lipreading Inventory under two conditions. While observing silent videos of 15 words being spoken, participants either held a tongue depressor horizontally with their teeth (blocking facial motor action) or squeezed a ball with one hand (allowing facial motor action). As hypothesized, blocking motor action resulted in fewer correctly understood words than that of the control task. The results suggest that facial mimicry or other methods of facial action support visual speech perception in children. Future studies on the impact of motor action on the typical and atypical development of speech perception are warranted.
Speech perception is a multi-modal process that involves the integration of auditory, visual, and even tactile information (e.g., Dick, Solodkin, & Small, 2010; Gick & Derrick, 2009; McGurk & MacDonald, 1976). Visual perception in particular aids in understanding speech via speechreading, or by watching a speaker’s lips, cheeks, nose, and eyes to understand what the speaker is saying (see Woodhouse, Hickson, & Dodd, 2009). Infants, young children, and adults use such visual cues (Dodd, 1987; Dodd & Burnham, 1988; Kuhl & Meltzoff, 1982; McGurk & MacDonald, 1976; Reisberg, McLean, & Goldfield, 1987; Sumby & Pollack, 1954). Despite agreement that visual input is important, the mechanisms by which it enhances speech perception remain unclear (Woodhouse et al., 2009). Some theories suggest that the perceiver’s motor processes are relatively important in speech perception (Desjardins, Rogers, & Werker, 1997; Kelly et al., 2002; Liberman, 1957; Liberman, Cooper, Shankweiler, & Studdert-Kennedy, 1967; Liberman & Mattingly, 1985; Liberman & Whalen, 2000; Sommerville & Decety, 2006), whereas other theories do not (Diehl & Kluender, 1989; Massaro, 1987; Massaro & Chen, 2008).
Theories of speech perception that deemphasize motor action instead stress the importance of the perceiver’s auditory (Diehl & Kluender, 1989) and audiovisual processing of speech (Massaro, 1987; Massaro & Chen, 2008). They also highlight justified concerns about the lack of strong evidence elucidating how and when motor input and motor output inform speech perception. There is also a concern that the recruitment of motor processes is burdensome to speech perception, rather than helpful (Massaro & Chen, 2008).
To test the proposition that motor action can support speech perception and begin to address the concerns of those who deemphasize this mode of perception, the present paper examines a specific claim made by some earlier motor-based theories as to why motor processes may be important: that facial motor action plays a role in speech perception (see Liberman, 1957; Liberman et al., 1967).
The earliest version of the motor theory of speech perception suggests that “we overtly mimic the incoming speech sounds and then respond to the proprioceptive and tactile stimuli that are produced by our own articulatory movements” (Liberman, 1957, p. 122). Liberman and colleagues theorized that these “linkages between sensory and motor components” (i.e., mimicry) mediate between what is heard and what is perceived (Liberman et al., 1967, p. 452; see Galantucci, Fowler, & Turvey, 2006 for a review). The importance of motor action in speech perception was eventually discounted, as such motor action does not appear to be necessary for either speech comprehension (e.g., MacNeilage, Rootes, & Chase, 1967) or the development of speech perception (e.g., Eimas, Siqueland, Jusczyk, & Vigorito, 1971). However, even if motor action is not necessary for speech perception, it may still aid typical speech processes and the development of speech perception.
Furthermore, a robust body of research on social mimicry demonstrates that motor mimicry itself may not always be as overt as Liberman had predicted. Rather, motor mimicry involves extremely subtle, rapid facial reactions that are often measured via small, electromyography (EMG) sensors on the skin (e.g., Dimberg, Thunberg, & Elmehed, 2000; Oberman, Winkielman, & Ramachandran, 2007). This automatic, mimetic action facilitate
Indeed, three lines of research suggest that motor action may influence speech perception. First, infants seem to mimic speech-related actions (Meltzoff & Moore, 1977, 1983). That they are capable of mimicry at this early point in development suggests that motor action has the potential to be an important component of the development of speech perception. Second, adults appear to physically mimic speech gestures. For example, listening to speech sounds and viewing speech-related lip movements elicits a covert, automatic response in the tongue and lips (Fadiga, Craighero, Buccino, & Rizzolatti, 2002; Moody & McIntosh, 2011; Watkins, Strafella, & Paus, 2003). Such motor action may provide input into speech perception processes.
The third line of research demonstrates that physically manipulating the parts of the face that are associated with speech production influences what speech sounds are heard. Specifically, Ito, Tiede, and Ostry (2009) used a robotic device to stretch participants’ skin on each side of the mouth while they listened to words. When their skin was stretched upwards, as associated with producing the word “head,” participants were more likely to hear the word “head” than “had.” Accordingly, when their skin was stretched downwards, as associated with producing the word “had”, participants were more likely to hear the word “had” than “head.” This study demonstrates that speech-related motor action plays a role in the auditory perception of speech. Because this study used auditory stimuli and focused on the effects on what word was heard, it does not provide direct evidence that somatosentory input affects visual speech perception. In addition, the input was directly manipulated, and thus the study does not address spontaneous motor action directly. Further research is needed to determine whether facial feedback generated by the perceiver influences the visual perception of speech.
These findings are consistent with the idea that motor action supports speech perception, but two issues remain. First, these findings do not demonstrate a causal link between the perceiver’s motor action and comprehension of speech. There is substantial evidence that visual input affects speech perception (Dick et al., 2010; Dodd, 1987; Dodd & Burnham, 1988; McGurk & MacDonald, 1976; Reisberg et al., 1987; Sumby & Pollack, 1954; Watkins et al., 2003; Woodhouse et al., 2009) and that the perceiver’s motor actions are involved in speech perception (Dick et al., 2010; Desjardins et al., 1997; Ito et al., 2009, Kelly et al., 2002; Liberman, 1957; Liberman et al., 1967; Liberman & Mattingly, 1985; Liberman & Whalen, 2000; Skipper, Goldin-Meadow, Nusbaum, & Small, 2007; Sommerville & Decety, 2006; van Wassenhove, Grant, & Poeppel, 2005). However, it remains unclear whether these are connected – whether the perceiver’s own facial motor action affect
Second, the role of motor action, as described by the motor theory of speech perception (Liberman, 1957; Liberman et al., 1967), has not been examined in children. Of particular interest is the point in development, primarily between ages six and eleven years old, when visual speech input begins to have more influence on audiovisual speech perception, as demonstrated by numerous applications of the McGurk paradigm (Hockley & Polka, 1994; Massaro, Thompson, Barron, & Larren, 1986; McGurk & MacDonald, 1976; Sekiyama & Burnham, 2008). The McGurk illusion refers to the effect of visual input on audiovisual speech perception tasks. For example, when an adult perceiver sees a speaker say /ga/ while hearing the speaker say /ba/, he or she is more likely to report hearing a combination of the two modalities (i.e., /da/) or what was seen (i.e., /ga/). As a group, children under eleven years old are less influenced by this illusion, and accordingly, are more likely to rely on the auditory modality (Hockley & Polka, 1994; Massaro et al., 1986; Sekiyama & Burnham, 2008).
This difference between age groups appears to be explained by developmental differences in how speech information is integrated. Research shows that adults have a more consolidated neural network during audiovisual speech perception, involving the coactiviation of multiple processing centers, compared to children between eight and eleven years of age (Dick et al., 2010). It is notable that adults’ audiovisual speech perception involves the coactivation of visual, auditory, motor and somatosensory processing centers, whereas children not only show a less consolidated pathway between visual and auditory centers, but significantly less activation of motor and somatosensory centers. These results are consistent with the notion that facial motor action plays a role in perceiving visual speech input, and that the integration of visual and motor input is particularly meaningful. Because neural activation of the motor pathways is less evident and consolidation less established, it may be the case that enhanced motor action is necessary to aid interpretation. Children of this age may be beginning to integrate the modalities, but may require more intense activation for the motor modality to have influence. Thus, if motor action plays a role in the development of speech perception, it is reasonable to assume it would be most evident when children are beginning to better integrate visual speech information, but this ability is not yet fully developed. This appears to be between the ages of six and eleven years of age.
The present study explores these questions by testing specifically whether children’s facial motor action contribute
To test this hypothesis we showed participants silent videos of a woman saying words and tested their speechreading accuracy when facial motor action was blocked and when it was not blocked. We also considered gender differences, given that girls frequently outperform boys on speechreading tasks (Bornstein, Hahn, & Hayes, 2004; Craig, 1964; Evans, 1965; Woodhouse, 2007; see Woodhouse et al., 2009).
1 Method
1.1 Participants
Thirty-five typically-developing children (19 boys; 16 girls) between the ages of 6.5–7.5 years old (M = 7; SD = .2) were recruited by phone from a pool of families who were randomly selected by birth certificate records and had agreed to be contacted about university research. Parents reported their child’s race and ethnicity and were able to mark as many categories as they wanted. The majority of the sample was identified as Caucasian (n = 34; 97%). Parents also identified their children as African American (n = 2; 5%), Hispanic (n = 2; 5%), Native American (n =1; 2%), Middle Eastern (n = 1; 2%), and Filipino (n = 1; 2%). Each family received $15 and each child selected a toy as compensation for his or her time.
1.2 Stimuli
The speechreading measure comprised 30 videos, each 3 seconds in duration. Each video pictures a woman saying a word, without sound. The videos were divided into two sets of 15 words: Words from Form A of the Revised Craig Lipreading Inventory (Updike, Rasmussen, Arndt, & German, 1992) were in the first set; words from Form B were in the second. The word lists represent common vocabulary observed in children in kindergarten and first grade. Frequencies for the words are listed in Table 1. All but five of the words were identified as high-frequency by the Corpus of Contemporary American English (see Davies, 2008-). Forms A and B are similar in difficulty. Furthermore, there is no difference in the forms based on word frequency, t(15) = .94, p = .36. Each word comprises one or two syllables.
Revised Craig Lipreading Inventory.
Note. Form A frequency and Form B frequency refer to the frequency of the spoken word between 1990–2012. Adapted from The Corpus of Contemporary American English: 450 million words, 1990–present, by M. Davies, 2008-). Available online at http://corpus.byu.edu/coca/.
1.3 Design
A within-subjects design was used to examine the effects of inhibiting motor action. During the presentation of the first set of 15 words, participants were randomly assigned one of two tasks. Half the children held a tongue depressor horizontally with their teeth, holding the lips away from the tongue depressor; this inhibits facial motor action. The other half lightly squeezed a foam ball in the palm of their hands; this allowed mouth activity while still requiring effort and concentration similar to holding the tongue depressor in the mouth (cf. Topolinski, 2012; Topolinski & Strack, 2009). During the presentation of the second set of 15 words, each participant did the opposite task. Gender and task order (tongue depressor task first or ball task first) were included as between-subjects factors. Randomizing the task order meant that the tongue depressor manipulation was not confounded with a word set (Form A was always presented first) and prevented practice effects (see Updike et al., 1992) from being confounded with manipulation.
1.4 Procedure
After the child and parent entered the lab, the experimenter explained the study and obtained parental consent and child assent. Parents then left the room. Children first viewed videos of words from Form A using a computer and engaged in either the tongue depressor or control task. The video of each word was shown twice in succession, with a blank screen displayed for two seconds between each presentation. After the second presentation of a word, the child indicated which of the four words was being spoken. The participant selected his or her answer choice by pointing to picture 1, 2, 3, or 4 displayed on the computer screen. Children then completed the other task (tongue depressor or control) with words from Form B. This procedure prevented the Form from being confounded with condition. Children did not receive any instructions beyond those involved with the task manipulation (ball, tongue depressor) and the speechreading measure. Families were then debriefed and the children received their toy.
1.5 Measures
Speechreading accuracy was assessed using word choices from the Revised Craig Lipreading Inventory (Updike et al., 1992). Possible speechreading scores for each condition ranged from 0 (none correct) to 15 (all correct). This measure was standardized with samples of children from the American mid-west, aged three to eight years, with typical hearing and age-appropriate speech and language abilities; it proves high reliability and validity (Updike et al., 1992).
As a manipulation check, participants were rated on their ability to comply with the task. Unable meant that they could not hold the tongue depressor horizontally with their teeth while keeping their lips from touching the tongue depressor for any duration of the speechreading measure. Those who were rated Inconsistent were able to hold the tongue depressor horizontally with their teeth while keeping their lips from touching the tongue depressor, but could not maintain this position for the entire duration of the speechreading measure. Participants were rated as Fully Able if they were able to hold the tongue depressor horizontally with their teeth while keeping their lips from touching to tongue depressor for the entire duration of the speechreading measure.
2 Results
2.1 Preliminary analyses
One participant scored more than 2.5 standard deviations above the mean on the speechreading measure. Results from the primary analysis were the same with and without the outlier; analyses including the outlier are reported here.
All participants were able to comply at least partially with the tongue depressor task. Nineteen were rated as Fully Able and 16 were rated as Inconsistent. Those children who were inconsistently able to comply occasionally allowed their lips to touch the tongue depressor and may have needed occasional breaks. These types of inconsistencies are not uncommon for pediatric groups. Moreover, previous research has found that allowing the lips to touch the motor-blocking device still has the intended outcome of blocking spontaneous motor action (Niedanthal et al., 2001). These compliance differences did not vary based on gender, χ2(1, N = 35) = .05, p = .83, nor was speechreading performance affected by compliance during the tongue depressor task, t(33) = .43, p = .67, indicating that there were no functional differences between the two levels of compliance (i.e., individual differences in difficulty of the task did not influence speech perception – those who did it easily did not differ from those who needed rest). Therefore, compliance was not considered further.
2.2 Primary analyses
To examine whether inhibiting motor action lowered speechreading accuracy, we conducted a condition (2; tongue depressor vs. ball, within-subjects) by task order (2; tongue depressor first vs. ball first) by gender (2; boy, girl) analysis of variance (ANOVA). As hypothesized, there was a main effect for condition, F(1, 31) = 5.17, p = .03, partial η2 = .14. As displayed in Figure 1, when children had a tongue depressor in their mouths, they showed lower speechreading scores (M =6.8, SD = 2.7) than when they were squeezing a ball, and thus free to produce facial motor activity (M =7.6, SD =2.6). There was also a trend for girls (estimated marginal mean =7.92, SE = .56) to score higher than boys (estimated marginal mean = 6.47, SE = .51), F(1, 31) = 3.65, p = .07, partial η2 = .10, and a significant three-way interaction of condition, task order, and gender, F(1, 31) = 5.27, p = .03, partial η2 = .14.

Speechreading scores during the tongue depressor task and the ball task. Bars represent plus/minus one standard error. Participants (N = 35) scored lower on the speechreading measure during the motor task that inhibited facial motor action (participants held a tongue depressor horizontally with their teeth), compared to the motor task that allowed facial motor action (participants lightly squeezed a ball with one hand).
To investigate this interaction, we first examined the means across each individual group (boys who did the tongue depressor task first, boys who did the ball task first, girls who did the tongue depressor first and girls who did the ball task first). Given the small sample size of each group (n = 6–10), interpretations should be considered cautiously. With one exception, mean speechreading performance followed the overall pattern, appearing lower during the tongue depressor task than during the ball task. The exception was for girls who performed the tongue depressor task second. However, none of these four groups showed a significant difference in scores based on condition (tongue depressor, ball). To further investigate the interaction, we conducted condition (tongue depressor vs. ball, within-subjects) by task order (tongue depressor first vs. ball first) ANOVAs separately for boys and girls. The interaction between condition and task order remained for the girls, F(1, 14) = 5.8, p = .03, partial η2 = .29, but not for boys, F(1, 17) = .40, p = .53, partial η2 = .02; girls who did the ball task first were the only group whose mean score appeared lower on the ball task than the tongue depressor task.
3 Discussion
Supporting the hypothesized influence of motor action on visual speech perception, speechreading scores were lower during the task that inhibited spontaneous mouth activity (holding tongue depressor horizontally with the teeth) compared to the task that allowed spontaneous mouth activity (squeezing a ball). When children were blocked from motor action, they perceived approximately one word less than when they were not blocked from motor action. Although this effect may amount to one missed word in the context of 15 word items, facial motor action is only one part of a complex and multi-modal process of speech perception. Certainly there are other powerful influences on speech perception. Nonetheless, our findings indicate that motor action does have an influence on perception. If there are impairments in facial action, even an effect of this size could have a compounding impact over time. For example, if children miss approximately five percent of the words presented visually as the result of impaired motor functioning, this may have an increasingly negative impact on language development and affect visual speech comprehension over time.
The results are consistent with the motor theory of speech perception, which suggests that physical mimicry contributes to speech perception (Liberman, 1957; Liberman et al., 1967). The results expand on the motor theory of speech perception by demonstrating in a randomized, controlled experimental manipulation that facial motor action appears to be one mechanism through which visual input leads to enhanced speech perception. However, the manipulation blocked facial action generally; it did not block matching action or mimicry specifically. Thus, the results do not establish that it is blocked mimicry per se that generated the deficit; it could have been the blocking of non-matching motor action that caused the lower levels of performance. Thus, this finding is generally consistent with theories of speech perception that emphasize the role of the perceiver’s motor system (Desjardins et al., 1997; Kelly et al., 2002; Liberman, 1957; Liberman et al., 1967; Liberman & Mattingly, 1985; Liberman & Whalen, 2000, Sommerville & Decety, 2006).
The three-way interaction among condition, task order, and gender was not predicted, and the cell sizes require caution in interpretation. However, this interaction is consistent with developmental interpretation
More generally, the results of the present study suggest that facial motor action contributes to children’s perception and understanding of others. There is already evidence that facial motor action, particularly mimicry, influences the perception of emotion in typically developing adults (Niedanthal et al., 2001; Oberman et al., 2007). Both emotion perception and visual speech perception appear to be enhanced by motor action.
To our knowledge, this is the first demonstration to show that blocking spontaneous mouth activity impairs visual speech perception in typically developing children. However, there are a few methodological factors to consider. Specifically, given that the manipulation blocked both matching and non-matching motor action, the study does not specifically establish that it is the blocked matching that impaired performance. Further research using EMG as a direct measure of muscle activity is needed to distinguish between the influence of facial motor action and mimetic facial action during visual speech perception. In addition, participants’ ability to comply with the control task (squeezing a ball) was not measured. Holding the tongue depressor appears to have been more difficult than squeezing the ball; this may be a reason for worse performance in the tongue depressor task. However, if difficulty influenced speechreading accuracy, we would expect that difficulty of complying with the tongue depressor task would also affect accuracy within that condition, and that was not the case. Regardless, these considerations should inform future research paradigms in that they should include tasks of similar, and measurable, difficulty.
Additional research will be needed to fully evaluate any developmental effect that facial motor action has on speech perception. That this was a cross-sectional study of children between 6.5–7.5 years old prevents us from making any strong developmental inferences. A broader developmental sample or a longitudinal study would help identify whether reliance on motor action decreases with age (i.e., older children may be less reliant on mimicry and therefore less impaired by the tongue depressor task), as theorized by Liberman and colleagues (1957, 1967). Because efforts to mimic or move the face more generally were not assessed, the present research does not provide information on potential individual differences in the role of motor action. The gender findings suggest a potential group effect. There may be other individual and group differences, including some related to atypical development. Research on the role of motor action in atypical speech development is also warranted.
These results are also consistent with findings by Ito et al. (2009) demonstrating that physical manipulation of parts of the face involved with speech production influences what speech sounds are heard. The present data go beyond their findings by supporting the first link in the chain. The present study indicates that people’s spontaneous motor action is important; the work of Ito et al. (2009) suggests that manipulated facial action (such as those blocked in the current manipulation) influences speech perception. Because selectively manipulating the face influences what words are heard (Ito et al., 2009), and blocking facial motor action impairs visual speech perception, it is reasonable to assume that mimicry is one way that what is spoken becomes what is perceived, as theorized by Liberman and colleagues (1957, 1967). An important difference in the studies is that Ito et al. (2009) implicate the auditory cortex and thus what words are heard whereas the current study did not involve auditory stimuli and simply looked at what words were perceived. Future research should further investigate this process by examining speech perception among listening, speechreading, and combined (i.e., participants can both see and hear the speaker) conditions.
Future research could extend this work by using additional manipulations of motor action. Passive means of inhibiting mimicry (e.g., Botox) have been shown to influence participants’ comprehension of emotional sentences (e.g., Havas, Glenberg, Gutowski, Lucarelli, & Davidson, 2010). If passive manipulations also influence speech perception, it would clarify whether any effect in the present study was related to the difficulty of the task or to the participant needing to take action to block natural motor action.
Furthermore, additional research will be needed to determine the scope of this effect on speech more generally. We have shown this effect in a paradigm designed to isolate the effects of facial motor action on visual speech perception. Future studies should assess the influence of facial motor action when speech includes both auditory and visual information.
4 Conclusions
Overall, the results demonstrate that facial motor action aids visual speech perception in typically developing children. These findings are consistent with the original motor theory of speech perception (Liberman, 1957; Liberman et al., 1967) and other theories that support the role of motor action in speechreading (Desjardins et al., 1997; Kelly et al., 2002; Liberman, 1957; Liberman et al., 1967; Liberman & Mattingly, 1985; Liberman & Whalen, 2000; Sommerville & Decety, 2006). They also expand the literature examining the role of motor action in the perception and understanding of others.
Footnotes
Acknowledgements
The authors gratefully acknowledge Betsy App, Annamarie Maaranen-Hincks, Taryn Brandt, Irena Pikovsky, and Wyndol Furman for their involvement with the project.
Funding
This work was funded by the University of Denver’s Partners in Scholarship Grant; the University of Denver’s Boettcher Foundation-funded Extreme Academics Grant; and the National Institutes of Mental Health (grant number NIMH F32MH81409).
