Abstract
Although prosodic deficits have been reported to occur with many different populations, little published research addresses treatment options for these deficits. This study was designed to examine one treatment’s impact, the six-step imitative approach, on the expressive affective prosody of an adolescent with autism who had average intelligence and good receptive/expressive language skills. As this approach has been successfully utilized in treating the affective prosody deficits in adults with acquired deficits, it was hypothesized that it could also be effective in treating individuals with prosody deficits associated with developmental disorders. A case study is presented to demonstrate the changes associated with the six-step imitative approach on the acoustic (fundamental frequency [F0], duration, and intensity) and the perceptual characteristics of speech. This study suggests that the six-step imitative approach may be beneficial in treating some expressive prosodic deficits in children with autism.
Introduction
It is important for clinicians to have evidence-based practices to utilize in the treatment of speech and language disorders. Prosody has been an area lacking evidentiary treatment support. As suggested by Hargrove (2013), the lack of information on the treatment of prosody may be due to the fact that individuals with prosodic deficits may also have other speech and language deficits that affect their communication to a greater extent. Therefore, treatment targeting prosody may take a backseat to the primary communication difficulties. However, there are some individuals, such as those with autism who have overall average intelligence and good receptive/expressive language skills, whose primary communication difficulties may result from their prosodic deficits. The focus of this study was to examine the efficacy of one type of treatment, the six-step imitative approach, on the affective prosody of an adolescent with autism and such characteristics.
Affective prosody allows communication partners to relate to one another, to understand changes in mood, to maintain appropriate discourse, and to share feelings (Bellon-Harn, Harn, & Watson, 2007). Affective prosody involves the vocal changes which occur based on the social situation, the communicative intention, or the emotional state that is being conveyed (Bolinger, 1989). These vocal changes are centered on the acoustic parameters of speech including fundamental frequency (F0), intensity, and duration (Kent & Read, 2002).
Hargrove (2013) evaluated four different published reports involving the treatment of prosody. One of these studies was conducted by Rosenbek et al. (2006) and involved the treatment of affective prosody using the six-step imitative approach. This imitative motor approach was designed based on the theory that expressive prosody deficits resulted from a motor disorder which included problems in planning, programming, and executing a message (van der Merwe, 1997). In its original form, the clinician followed a six-step continuum by first providing maximal cues and slowly removing the cues until the client independently carried on a conversation using varying stress patterns and emotional tones (Leon et al., 2005).
Leon et al. (2005) were among the first to suggest that prosodic deficits were amenable to behavioral treatment (i.e., the imitative approach). Their three participants were diagnosed with expressive prosodic deficits due to right hemisphere damage (RHD) from stroke. The participants’ prosody was treated using both the cognitive-linguistic and the imitative approaches. The participants’ responses (i.e., happy, sad, or angry emotionally toned sentences) were analyzed. Visual inspection and statistical analysis both showed evidence of treatment effects from each treatment for all three participants. Two participants showed greater improvement with the imitative approach, while one participant showed greater effects with the cognitive-linguistic approach. In addition, all of the participants generalized the treated emotions to untreated sentences (Leon et al., 2005).
Rosenbek et al. (2006) studied the treatment impact of the imitative approach on 14 adults with bilateral or RHD. The researchers observed improvement in six out of the seven patients initially treated using the imitative approach. Five out of five patients showed some improvement with the imitative approach when it was provided as the second treatment (Rosenbek et al., 2006).
Russell, Laures-Gore, and Patel (2010) also examined the effects of the six-step imitative therapy (Rosenbek et al., 2006) on an individual with prosodic deficits due to stroke. The individual demonstrated difficulty with the suprasegmental aspects of language needed to express his emotions. Speech samples were acoustically analyzed for changes in peak F0, peak intensity, and utterance duration. The participant showed moderate improvement in the ability to modulate prosody within utterances, but the progress was not maintained at follow-up. The imitative approach resulted in both acoustic and perceptual changes over time with this participant. Although the acoustic transformations were minimal, there were significant perceptual changes in the identification of stressed words. The researchers suggested that slight alterations in the acoustic signal resulted in improved listener perceptions (Russell et al., 2010).
The aforementioned studies on the treatment of affective prosody all focused on treating adults with acquired prosodic deficits. However, individuals with autism spectrum disorder (ASD) often experience deficits in prosody (Baltaxe & Simmons, 1985, 1992; Fay & Schuler, 1980; Paul, 1987). In examining available studies on prosody in individuals with autism, Shriberg et al. (2001) found atypical prosody in individuals in 10 out of the 10 studies surveyed. Prosodic differences in individuals with autism included longer sentence durations (Baltaxe, 1981), wide intensity variations (Pronovost, Wakstein, & Wakstein, 1966), monotone, limited or exaggerated pitch ranges (Baltaxe, Simmons, & Zee, 1984; Pronovost et al., 1966), and even a “sing-song” type of intonation (Fay & Schuler, 1980).
According to Peppé, McCann, Gibbon, O’Hare, and Rutherford (2007), children with ASD, in general, were more likely to experience deficits in social reciprocity due to vulnerability in emotional and pragmatic aspects of speech. Children with autism, with average to superior intelligence, were found to be capable of improving both their prosodic awareness and prosodic production abilities (Peppé et al., 2007). McCann and Peppé (2003) suggested that because some speakers with autism experience difficulties with social deficits due to their “odd” prosodic characteristics, prosody should be a main target for intervention. In addition, Paul et al. (2005) found that resonance and stress had some impact on listeners’ perception of both social and communicative competence with some individuals with autism. Deficits in social competence can have far-reaching impacts on individuals with autism. Many adults with autism may be either unemployed or underemployed, be unable to live independently, and lack important social relationships (Gutstein & Whitney, 2002).
Like adults with acquired prosodic deficits, children with autism have shown improvement in their expressive prosody skills with treatment. Bellon-Harn et al. (2007) clinically managed an 8-year-old child diagnosed with autism. The child demonstrated typical pitch and intensity but atypical syllable prolongations. The participant’s overall prosody was significantly affected because the syllable prolongation changed both the tempo and rhythm of speech. In addition, the participant displayed excessive and atypical pausing. The researchers employed visual and tactile feedback, modeling, discrimination, and metalinguistic strategies along with the interactive approach in treatment. Through the intervention, the participant’s prosody improved, atypical pausing declined, and the perception of “oddness” diminished (Bellon-Harn et al., 2007).
Further investigation of the association of prosody and autism was completed by Shriberg et al. (2001). These researchers used the Prosody Voice Screening Profile (PVSP; Shriberg, Kwiatkowski, & Rasmussen, 1990) as the standardized measurement of prosodic ability for 30 male participants ranging in age from 10 to 49. The researchers compared the conversational speech samples from the participants with those of age-matched controls. The recordings were analyzed based on phrasing (repetitions), rate, stress, intensity, pitch, laryngeal quality, and resonance. Forty percent of the speakers with autism displayed inappropriate or nonfluent phrasing on more than 20% of the utterances, more than half demonstrated deficits in the production of appropriate stress, and 40% were characterized as hypernasal (Shriberg et al., 2001).
Nadig and Shaw (2012) examined the speech productions of 15 children with autism. The researchers found that the children with autism had higher pitch ranges in conversational speech, but their mean pitch and rate values were not significantly different from a control group of children with typical development. Diehl and Paul (2013) attempted to identify the clinically relevant acoustic characteristics of speech in individuals with ASD. The researchers found that children with ASD had difficulties controlling the precise temporal aspects of word production. Acoustic differences included longer durations suggesting that the timing of stressed and unstressed syllables were less differentiated in children with ASD (Diehl & Paul, 2013).
The current study was designed to determine if the six-step imitative approach would lead to prosodic improvement in a child with autism who had average intelligence and overall good receptive/expressive language skills. It was hypothesized that the six-step imitative treatment would be effective based on improvements shown in adults with prosodic deficits utilizing this approach and success noted in children with autism using other behavioral treatments. Acoustic characteristics (F0, intensity, and duration) were chosen as outcome measurements, as they provide objective measures of speech performance. Acoustic analysis allowed the researchers to examine the specific elements of production, as well. Perceptual measurements were also taken to determine if changes in acoustical characteristics resulted in perceptual changes.
Method
Participants
The single participant was a 14-year-old male with autism. The diagnosis of autism was based on the Diagnostic and Statistical Manual of Mental Disorders (5th ed.; DSM-5; American Psychiatric Association, 2013) and the Gilliam Autism Rating Scale (GARS; Gilliam, 1995). The participant passed a hearing screening and oral-peripheral examination. His scores for language and cognition were within the average range based on the Clinical Evaluation of Language Fundamentals–Fourth edition (CELF-IV; Semel, Wiig, & Secord, 2003) and the Test of Nonverbal Intelligence–Third edition (TONI-3; Brown, Sherbenou, & Johnsen, 1997). Based on parent completion of the GARS (Gilliam, 1995), the participant received an autism quotient of 75 with difficulties noted in the areas of stereotyped behaviors, communication, and social interaction. Disturbances in prosody were determined based on a short speech sample scored using the PVSP (Shriberg et al., 1990). On the PVSP, the participant demonstrated difficulties with prosody in terms of phrasing, rate, and stress, but vocal quality was good.
Materials and Instrumentation
The PVSP (Shriberg et al., 1990) was completed at the time of pre- and posttesting to measure perceptual changes in the participant’s speech. As directed in the PVSP training manual (Shriberg et al., 1990), a minimum of 12 utterances had to be coded from a raw speech sample to obtain accurate screening results. Guidelines for coding were followed based on the training manual. Raw speech samples were collected at pre- and posttesting. Twenty-four utterances were coded and measured for this analysis at each testing time.
The target stimuli (based on those compiled by Leon et al., 2005, and shown in Appendix A) consisted of 24 sentences with contrastive stress on one word within the utterance that conveyed the specified emotion (i.e., happiness, sadness, or anger). Pre- and posttest speech samples were collected from a separate set of 12 stimulus sentences (see Appendix A). Each of these 12 sentences also contained a key word designating the target emotion. Changes in average F0, intensity, and duration were measured using the TF32 program software (Milenkovic, 2002) on an HP laptop computer with the Windows 7 operating system. The participant wore a head-mounted Gigaware microphone (model # 4300122) connected to the computer during data collection. The microphone was positioned approximately 2 in. from the speaker’s mouth. Speech samples were saved as .wav files on the computer with a sampling frequency 44.1 kHz.
Procedure
All pre- and postassessment measures as well as intervention procedures were conducted by the second author in a therapy room located in a university speech and hearing clinic. All intervention occurred in 1 hr weekly sessions for a total of 10 sessions using the six-step imitative method (Leon et al., 2005; Rosenbek et al., 2006; see Appendix B). The participant received one-on-one intervention based on a six-step continuum arranged from easiest (most cueing) to hardest (least cueing) as shown in Appendix B. The participant progressed through every step based on three consecutive correct responses at each step. Acoustic measurements were completed using the TF32 program (Milenkovic, 2002).
Data Analysis
Raw data were in the form of F0, durations, and intensity levels. Descriptive statistics were computed and paired-sample t tests were used to compare the data collected at pre- and postintervention. The samples were derived from the production of 12 sentences targeting the specified emotions (happiness, anger, or sadness; see Appendix A). Once each of the 12 stimulus sentences was recorded, the Pitch Trace and RMS Trace features of the TF32 program were used to generate spectrograms and mean data.
Intrarater Reliability
To determine the intrarater reliability of the measurement data, 10% of the speech samples were randomly selected and reexamined by the second author using the TF32 program (Milenkovic, 2002). F0 differed by a mean of 0.97 Hz. Full utterance durations differed by a mean of 20 milliseconds (ms). Stressed syllable duration differed by a mean of 2.67 ms, while unstressed syllable duration differed by a mean of 4.95 ms. Stressed syllable intensity differed by a mean of 0.01 V, while the unstressed syllable intensity differed by a mean of 0.02 V. All of the values calculated showed very little variation which indicated good reliability of measurements.
Results
Perceptual Analysis
The PVSP (Shriberg et al., 1990) was used to measure the perceptual prosodic characteristics of the participant’s speech samples (pre- and posttest). Pre- and postintervention results are shown in Figure 1. At pre-intervention, 12% of the utterances had inappropriate phrasing (word repetitions and/or revisions). The rate (too slow/too fast) was inappropriate for 16% of the utterances. The stress (reduced, excess/equal, multiple features) was inappropriate for 56% of the utterances, while voicing was considered appropriate for 100% of the utterances. At postintervention, the participant demonstrated no difficulties with phrasing for any of the utterances. The rate remained inappropriate for 16% of the utterances, whereas the stress was inappropriate for only 16% of the utterances. Slight improvement occurred with phrasing while dramatic improvement with stress was observed at postassessment.

PVSP pre- and posttest scores of perceptual changes.
Acoustical Analysis
F0
Paired-sample t tests showed no significant differences in the mean F0 values from pre- to postintervention for happiness, M = −25.38, SD = 28.25, t(3) = −1.80, p > .05; anger, M = 16.23, SD = 23.40, t(3) = 1.40, p > .05; or sadness, M = −7.53, SD = 25.09, t(3) = −0.60, p > .05. These results coincide with the perceptual ratings indicating no difficulties with pitch at pre- and postintervention.
Duration of full utterances
As shown in Figure 2, significant differences in the mean utterance durations for all three emotions were found using paired-sample t tests: happiness, M = −267.35, SD = 129.38, t(3) = −4.13, p = .03; anger, M = −332.85, SD = 148.80, t(3) = −4.47, p = .02; and sadness, M = −347.77, SD = 112.06, t(3) = −6.21, p = .01. All utterances were significantly longer after intervention.

Mean utterance durations pre- and posttest for all three emotions.
Duration of unstressed syllables
Paired-sample t tests showered no significant difference in the mean unstressed syllable durations for happiness, M = −55.45, SD = 114.63, t(10) = −1.60, p > .05, or anger, M = −9.52, SD = 106.45, t(7) = −0.25, p > .05 (see Figure 3). However, a significant difference in the mean unstressed syllable durations was found for sadness, M = −49.81, SD = 66.92, t(10) = −2.47, p = .03. The analysis showed no significant differences in the duration of unstressed syllables for happiness and anger after intervention. However, significant differences did occur with sadness with the unstressed syllables being significantly longer after treatment.

Mean unstressed syllable durations for pre- and postmeasures of all three emotions.
Duration of stressed syllables
Paired-sample t tests showed no significant difference in the mean stressed syllable durations for happiness, M = −23.96, SD = 70.45, t(8) = −1.02,p > .05, but significant differences were found for anger,M = 95.0, SD = 150.02, t(9) = 2.00, p > .05, and sadness,M = −92.36, SD = 79.44, t(7) = −3.29, p = .01 (see Figure 4). Significant changes were seen after treatment for anger where the participant actually shortened the durations of the stressed syllables for this emotion. Significant differences also occurred for sadness where the participant produced longer durations for stressed syllables following treatment.

Mean stressed syllable durations at pre- and posttesting for all three emotions.
Intensity of unstressed syllables
As shown in Figure 5, paired-sample t tests showed no significant differences for happiness, M = 0.08, SD = 1.01, t(10) = 0.29, p > .05. However, significant differences were found for anger, M = 2.20, SD = 1.51, t(7) = 4.11, p = .01, and sadness, M = 1.62, SD = 0.99, t(10) = 5.46, p = .000. The participant dramatically reduced the intensity of the unstressed syllables associated with anger and sadness from pre- to postintervention.

Mean unstressed syllable intensities at pre- and posttest for all emotions.
Intensity of stressed syllables
Paired-sample t tests showed no significant difference in the stressed syllable intensity values for happiness, M = −0.03, SD = 1.18, t(8) = −0.07, p > .05, but significant differences were found for both anger, M = 3.05, SD = 0.95, t(9) = 10.15, p = .000, and sadness, M = 2.20, SD = 1.32, t(7) = 4.71, p = .002 (see Figure 6). The participant also significantly reduced the intensity of stressed syllables for anger and sadness at postintervention.

Mean stressed syllable intensities at pre- and posttesting for all three emotions.
Discussion
This study was designed to determine the acoustic and perceptual changes associated with the six-step imitative approach on the affective prosody of a child with autism during structured speech tasks. Pre- and postintervention data in the form of voice recordings were collected for analysis. As hypothesized, the participant’s affective prosody improved both acoustically and perceptually through the implementation of the six-step imitative approach. Analysis of the acoustical data showed significant posttreatment differences for all variables except F0. Perceptual changes were also observed for several variables.
The changes in the mean durations of full utterances were statistically significant for all target emotions. Mean utterance durations were much longer at posttesting. The duration of unstressed syllables was statistically significant only for sadness, while the duration of stressed syllables was significant for both anger and sadness. Given that all three emotions were affected by utterance duration, it may be concluded that duration is an important component when conveying and training affective prosody. Bellon-Harn et al. (2007) also found that prosody was greatly affected by syllable duration because it changed the overall rhythm and tempo of their participant’s speech. Unlike the previous research by Bellon-Harn et al. (2007), however, the current study found that the longer the syllable duration, the better the overall perception of prosody.
For intensity, significant changes were found in the unstressed and stressed syllables for both anger and sadness. The majority of the syllables exhibited a reduction in intensity postintervention. The lack of change in intensity of syllables for happiness may indicate that the participant applied the intensities used with happiness to all emotions. Intensity appeared to be a characteristic which was over applied to emotional content. This participant utilized this aspect as the major characteristic for displaying emotion. Even though the intensities of the majority of the syllables were reduced, the overall perception of stress improved.
The changes for unstressed and stressed syllables in regard to both anger and sadness were significant. Similarly, the perception of stress improved greatly. Although there were still instances of either reduced stress, excessive stress, or equal stress on words within the utterance postintervention, the participant produced appropriate stress on the majority of utterances. He added stress when needed to convey the importance or excitement of his message and alternatively reduced stress to underscore trivial words in his utterance. Even though these perceptual changes from pre- to postintervention were not dramatic, the participant’s PVSP percentages and profile improved, coinciding with the changes in the acoustical data.
The results of this study provide possible treatment options for children with autism who have abnormal affective prosody. The participant’s ability to imitate and independently produce distinct emotional tone showed that a child with autism has the ability to receptively focus on distinct acoustic features and train vocal parameters to fit a model. In addition, the participant’s ability to express the targeted emotion improved with practice. This supports previous research by McCann, Peppé, Gibbon, O’Hare, and Rutherford (2007) who analyzed the prosodic skills of children with autism using the Profiling Elements of Prosodic Systems in Children. Their results suggested receptive/expressive skills are connected, indicating that prosody is a potential therapy target, specifically for children with higher nonverbal skills who have normal or even superior knowledge of grammar and phonology (Peppé et al., 2007).
As the prevalence of ASD continues to rise, understanding prosodic abilities and how to teach specific skills within this population could make treatment more effective and meaningful for the child. Even though children with autism have normal-to-above average intelligence, their ability to express feelings and connect with others is often impaired.
The six-step imitative approach improved the affective prosody skills of a single child with autism. Significant results from a relatively short intervention period (10 weeks) supported research by Rosenbek et al. (2006) who found that the imitative approach was easy to learn and showed fast results. Anger and sadness were the targeted emotions that showed the most change acoustically, though all emotions improved perceptually. Overall, the results were consistent with previous research by Russell et al. (2010) that suggested even minimal acoustical changes produced improved perception of the speaker’s voice, which could be enhanced with training. F0 was not found to be a significant factor in affective prosody, whereas duration and intensity, especially intensity, were found to be strong indicators of appropriate emotional tone.
Further investigation, both in children with ASD and adults who experience deficits in prosody due to acquired disorders, is recommended to identify and better understand the impact of the imitative approach on affective prosody that may then be generalized to larger populations of individuals. Future investigators should also examine speech performance in more naturalistic environments as well as carryover with the six-step imitative approach.
Given the wide variety of skills and deficits that individuals with ASD experience, it is not expected that one type of treatment will work for every individual. Instead, it is critical that evidence-based strategies be available to clinicians so that they may have options to try with their clients. The six-step imitative approach may be one of those options.
Footnotes
Appendix A
Appendix B
| Steps | Stimuli provided | Participant instructions |
|---|---|---|
| Step 1 | • Clinician reads target sentence with appropriate affective tone • Clinician and client say sentence in unison |
“First I am going to say a sentence in a (sad, happy, angry, neutral) tone of voice. Just listen the first time. Then I want you to watch and listen and try to say it right along with me. Use my same tone of voice.” |
| Step 2 | • Clinician models tone of voice and facial expression associated with the target sentence • Participant repeats sentence with the same tone of voice |
“Now I am going to say a sentence and I want you to watch me and listen carefully. After I finish, I want you to say the same sentence, using the same tone of voice I used.” |
| Step 3 | • Clinician models tone of voice but conceals facial expressions • Participant asked to repeat sentence with appropriate prosody without cue of facial expression |
“Now I am going to say a sentence, and after I finish, I want you to say the same sentence, using the same tone of voice I used.” |
| Step 4 | • Clinician produces sentence in neutral tone of voice • Participant repeats sentence using the correct target emotion |
“Now I am going to say a sentence, and I want you to say it after me, but I want you to use a (happy, sad, or angry) tone of voice.” |
| Step 5 | • Clinician asks question regarding emotion • Participant answers with target sentence and required to use specified emotion |
“Now I am going to ask you a question, and I want you to answer me using the same sentence.” If the stimulus response is happy, sad, or angry, the clinician will ask, “Why are you so (happy, sad, angry)?” |
| Step 6 | • Clinician prompts participant by asking them to explain how they feel using their tone of voice | “Now I want you to imagine I am part of your family or a really close friend, and you have to tell me exactly how you feel. Use the same sentence and same tone of voice.” |
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
