Abstract

Study Rationale
Clinicians are often asked whether a child will outgrow speech errors without intervention. There are two potential outcomes for children with speech sound disorders (SSDs): normalization or long-term, persistent SSD (Roulstone et al., 2009; Wren et al., 2016). Understanding whether children outgrow speech errors without intervention is of clinical importance because there are fewer speech-language pathologists (SLPs) than the number of children who need their services, and there are often long waiting lists. Access to speech-language pathology services can impact children, their families, and society (McGill et al., 2020). Some children who do not receive speech-language pathology services may be able to resolve their speech errors at a later time point, whereas other children continue to show a large number of errors even when they receive belated therapy. For these latter children, delays in services or insufficient intervention frequency may lead to poor speech outcomes affecting children’s education, social development, and occupational prospects (McLeod et al., 2019).
This study investigated four consistently reported risk factors of SSD related to children’s speech and language profiles (low stimulability, intelligibility, presence of atypical errors, and expressive language difficulties) and how these factors related to children’s time to normalization. This article presents the findings of a naturalistic cohort study investigating speech normalization and normalization rates in Cantonese-speaking preschool children in Hong Kong, China, who may be at risk for SSD. The objectives of this study were to quantify speech normalization rates at 2.5-year follow-up and to investigate predictors of time to normalization. The researchers examined (a) children who resolved nonadult realizations of speech sounds (i.e., had normalized production of speech sounds) and (b) those who had persisting speech sound difficulties (did not normalize) over 2.5 years.
Method
Over 800 Cantonese-speaking children were recruited to participate. Eight age groups with intervals of 6 months were included (from 2.4 to 6.9 years). In Stage 1 of the project, 845 children were assessed with a Cantonese articulation test that elicited all speech sounds in Cantonese, the Intelligibility in Context Scale-Traditional Cantonese (Kok & To, 2019); an oral-mechanism evaluation, the Hong Kong Cantonese Receptive Vocabulary Test (HKCRVT; Cheung et al., 1997); and a standardized language assessment. From Stage 1, 82 children were selected to participate in Phase 2 (follow-up longitudinal study) if their standardized scores on the HKCRVT were below -1.25 SD, or if they could not pronounce the initial consonants that were expected of their age, or if they had one or more atypical errors (e.g., dentalization of alveolar fricatives and affricates; To et al., 2013), including distortions. They also were not receiving intervention at the point of screening, they did not show any significant oral-motor deficits, and they had not been diagnosed with other neurodevelopmental disorders. Follow-up interviews and assessments were conducted 6, 12, and 18 months after the beginning of Stage 2. The assessments sought to identify the children’s speech sound production and language abilities. At the last assessment point, care-givers were asked whether the child had received intervention during the study and the content of the therapy.
Survival analysis was used to analyze time to normalization (i.e., time to event). Event time (i.e., the survival time) is the difference between the initial time, when no one showed a completed sound inventory (normalization) and all participants were considered as having a potential to normalize the speech (i.e., at risk in survival analysis); and the time at which the child showed a completed speech sound inventory (normalization; if it does occur). To complete a survival analysis, a binary outcome (target event) is required. The outcome of interest was word-initial consonant inventory (as opposed to inventory of final consonants) because word-initial consonants are more frequent in Cantonese than word-final consonants (To et al., 2013), and Cantonese-speaking children with SSD have significantly lower consonant accuracy in word-initial position than within-word or word-final positions (McLeod & Masso, 2019). Completion of the initial consonant inventory, as evaluated by the Hong Kong Chinese Articulation Test (HKCAT; Cheung et al., 2006), was the target event. To be considered normalized, children has to exhibit mastery of all Cantonese initial consonants.
All speech errors produced by each child were classified into typical and atypical error patterns. Typical error patterns refer to patterns produced by 5% or more of the children in the general population who speak Cantonese as their native language who are age 2;6 or above. Atypical error patterns are those used by fewer than 5% of the children. For example, backing was regarded as a developmental pattern because it was exhibited by approximately 5% to 10% of children when 3 years old, whereas initial consonant deletion was considered an atypical process for Cantonese-speaking children because it was rare. Distortions were also considered a type of atypical errors.
Measures and Outcomes
Stimulability
Stimulability refers to a child’s ability to modify speech production errors when stimulated by a clinician with models and cues (Glaspey & Stoel-Gammon, 2007; Lof, 1996; Powell & Miccio, 1996). The actual procedures vary across clinicians’ practice (e.g., the number of presented stimulations, stimuli in isolation vs. at word level, or in nonsense syllables). This study used nonsense syllables to evaluate stimulability, which was dichotomously defined for each child (either stimulable or nonstimulable). If children were stimulable for all misarticulated sounds in nonsense syllables (e.g., /Ca/, and /Ci/ or /Cy/), they were considered stimulable. On the contrary, if children were stimulable for part of the sounds or nonstimulable for all the sounds in the nonsense syllables, they were considered nonstimulable.
Stimulability was a predictor of speech normalization in children who had received no intervention. Stimulability was defined as being stimulable for all misarticulated consonant phonemes in at least two different consonant–vowel nonsense syllables rather than in isolation or in real words as in previous studies (Lof, 1996). Syllables are a particularly important unit for speech motor control (Tourville & Guenther, 2011). Not using real-word stimuli was to avoid children accessing their existing motor plans of the words that may have been well learned but incorrect.
Stimulability is a type of dynamic assessment that indicates children’s potential. Being stimulable requires the integrity of the sensory input and a generally intact linguistic and motor output system (Powell & Miccio, 1996). Stimulability may reflect a child’s ability to focus on speech productions and being motivated to change (Lof, 1996). If children can imitate misarticulated consonants with stimulation, those consonants are likely to be added to the phonetic inventory even without intervention (Powell & Miccio, 1996). A dichotomous measure of stimulable versus nonstimulable was adopted in this study to represent children’s global ability to imitate misarticulated consonants. With this binary differentiation, some children were stimulable for all misarticulated consonants, whereas others were nonstimulable to all consonants and some were stimulable to some consonants.
Error Atypicality/Typicality
As young children start to learn their language, predictable patterns of sound errors are observed in their production due to motor and perceptual restrictions. Atypical errors refer to substitutions, syllable structure errors, and distortions that are not generally found in typical phonological development.
It has long been assumed that children exhibiting typical developmental error patterns are those children who are “delayed,” as opposed to “disordered” (Dodd, 2014). Children with a profile of speech delay/typical error patterns have been described as being more likely to have a shorter normalization time than children exhibiting disordered/atypical error patterns. It is assumed that children with atypical speech errors are more likely to require intervention because they are less likely to self-correct their errors. Dodd et al. (2018) and Morgan et al. (2017) reported that children who receive speech-language pathology services demonstrated a more atypical error profile than those without speech-language pathology services. SLPs may have used speech error types to prioritize cases for intervention. However, no association between error type and time to normalization was found in this study. Children with atypical error patterns may not necessarily take a longer time to normalization than those children with typical developmental errors when no intervention was received, contrary to expectations from previous literature (cf. Morgan et al., 2017). Some children demonstrated atypical error patterns at Stage 1 and then resolved all errors and achieved speech normalization within the study period without any speech-language pathology intervention. This contrasts with the study by Morgan et al. (2017), who found that type of error (i.e., delay or disorder) was the only significant predictor of normalization by age 4, with children with the delayed pattern more likely to resolve than those with atypical errors. A possible explanation for the differences between these studies may be possible differences between atypical errors in English compared with atypical errors in Cantonese. For example, although many atypical errors overlap, backing is considered to be atypical in English, but typical in Cantonese. Thirty-nine children who started speech-language pathology services during the study period were shown to have significantly more atypical error patterns and a smaller consonant inventory at the initial time point.
Intelligibility
Speech intelligibility refers to the degree to which the listener understands what the speaker says, and it has been described as the most practical single index to apply in evaluating oral communication competence (Subtelny, 1977). Clinicians have used this as an important measure to determine the presence of SSD and the need for intervention (McLeod, 2020; Williams et al., 2021), implying that children with poor intelligibility are not likely to resolve their errors naturally and intervention is necessary. McLeod et al. (2012) developed the Intelligibility in Context Scale (ICS), a parent rating scale to estimate children’s intelligibility with various communication partners. The ICS has been translated into 60+ different languages and validated in 14 languages (McLeod, 2020). The traditional Chinese translation of the ICS (ICS-TC), which has been normed and validated on the Cantonese population, was used in this study (Kok & To, 2019; Ng et al., 2014).
Children who exhibited a mean ICS-TC score that was lower than the stated cutoff (less intelligible) were less likely to normalize and took longer time to normalize. The finding that parent-reported intelligibility was associated with children’s time to normalization whereas atypical errors were not adds to the worth and convenience of a quick measure of intelligibility, justifying that intelligibility data be routinely collected by SLPs (Ireland et al., 2020). Parent-reported intelligibility on the ICS has been found to be correlated with percentage of consonants correct (PCC) in many different languages and countries (McLeod, 2020).
Expressive Language Ability
Expressive language ability was not significantly associated with speech normalization. This is in accord with the study of Morgan et al. (2017), who reported that core language scores were not a significant predictor of later speech outcomes at the age of 7 years.
Conclusion
Analysis of data in this study indicated that when Cantonese-speaking children with SSD did not receive intervention, longer term (2.5 years) outcome may be predicted by stimulability and intelligibility but not by atypical errors or expressive language ability. The median time to normalization was 6.59 years of age. Children who were more stimulable and had higher intelligibility were more likely to resolve their errors naturally and took a shorter time to normalize. In this study, atypical error patterns and expressive language ability, however, did not predict speech outcomes. Children who showed atypical error patterns and/or weaker language ability did not necessarily take longer to normalize.
The authors concluded that children who present with better stimulability and higher intelligibility are more likely to represent instances of typical development rather than atypical development. Speech therapy may speed up speech normalization in these children, but it may not be necessary for children who can self-correct their errors (i.e., who are stimulable) and whose errors do not impact their intelligibility. Children with low intelligibility and poor stimulability should be prioritized for speech-language pathology services because their speech errors are less likely to resolve naturally. Stimulability testing and a speech intelligibility rating by caregivers are important components of routine clinical assessment for children with potential SSD because these measures can be used for caseload prioritization.
