Abstract
Many children with autism spectrum disorder (ASD) demonstrate poor comprehension of language at the sentence level in both the spoken modality and the graphic symbol modality. This study explored whether children with ASD are able to follow directives when presented with a graphic symbol sentence that includes an animated symbol for a verb. A total of five participants with moderate-to-severe ASD were presented with 10 graphic symbol sentences and asked to perform the directive using the provided figurines/objects. Results demonstrated that children with ASD can correctly carry out full-sentence directives to varying degrees when the directives represented graphically include an animated verb. Several observations were noted pertaining to participants’ performance and autism severity. The results of this study may have important implications for using animation as a tool to facilitate symbol syntax comprehension.
Keywords
Many children with autism spectrum disorder (ASD) have impaired receptive language ability relative to expressive abilities (Hudry et al., 2010; Kjelgaard & Tager-Flusberg, 2001). Specifically, children with ASD often demonstrate difficulty comprehending the abstract and relational components of language (e.g., verbs, prepositions, and adjectives), and the rules of syntax (Swensen et al., 2007). Although there is limited research on receptive language interventions (Dada et al., 2020; Sevcik, 2006), prior studies have found instruction with visual cues to be more successful for children with ASD than spoken-only instructions (Mechling & Gustafson, 2009; Quill, 1997). One promising approach is augmented input, which is the use of visual supports to supplement spoken language through augmentative and alternative communication (AAC) modalities (e.g., Allen et al., 2017; Beukelman & Garrett, 1988). In contrast to other approaches such as video modeling that aim to teach a specific behavior or skill (Bellini & Akullian, 2007), the goal of augmented input using the child’s AAC system is to teach the child language (Allen et al., 2017; Goossens, 1989).
Among the most commonly used visual cues as part of augmented input are so-called element cues (ECs). These are individual graphic symbols that represent key linguistic components of a spoken utterance (Allen et al., 2017; Shane et al., 2012) and are sequenced to reflect spoken word order forming a visual sentence (Boyer et al., 2012; Shane et al., 2012). For the purpose of this study, the term “ECs” refers to a sentence that is organized by ECs. Figure 1 shows an example of Subject–Verb–Object (SVO)-structured ECs, representing the sentence “Woody hug Gumby.” When ECs are used as augmented input in clinical practice, communication partners typically point to each EC while simultaneously speaking to the child; this is thought to enhance the child’s learning of language form and content (Allen et al., 2017).

Example of subject-verb-object-structured element cues representing the sentence “Woody hug Gumby.”
Although research has found ECs as augmented input to improve receptive vocabulary (Dada & Alant, 2009), there is a paucity of research on whether ECs improve language comprehension at the phrase and sentence levels (Allen et al., 2017). Moreover a recent study by Allen et al. (2018) found children with moderate-to-severe ASD to have difficulty following spoken prepositional directives when augmented with ECs. The study presented participants with three directive-following conditions: (a) spoken-only, (b) ECs plus spoken input, and (c) visual scene cues (SCs) (i.e., short video clips or photos that depict a directive or single concept) plus spoken input. Figure 2 shows an example of a static SC depicting “Woody hug Gumby.” Participants performed with greater accuracy in the SCs plus spoken input condition, followed by spoken-only, then ECs plus spoken input. The poor performance observed in the ECs condition relative to the SCs plus spoken input condition could be partially explained by the fact that ECs represent more abstract language components (Allen et al., 2018; Shane et al., 2012). Furthermore, knowing the ECs in isolation does not guarantee comprehension of the ECs in a sentence. A child must understand the semantic relations between the symbols and the rules of syntax, a language processing skill that is often difficult for those with ASD (Boyer et al., 2012; Shane et al., 2012).

Example of a static scene cue depicting “Woody hug Gumby.”
A potential strategy to improve comprehension of ECs is to animate the symbol representing the more abstract language components. Several studies involving typically developing children (Mineo et al., 2008; Schlosser et al., 2012, 2014) and children with ASD and other developmental disabilities (Fujisawa et al., 2011; Schlosser et al., 2019) have shown that animation assists with naming and/or identifying single verbs relative to static symbols. However, it is yet unknown whether children with ASD are able to follow a directive when presented with a graphic symbol sentence that contains an animated EC. In typical development, children have been found to associate words with semantic features and group them into linguistic categories based on their specific syntactic role (Owens, 2012). Verbs in particular have a relational function and have semantic features of agents, patients, recipients, and so on (Owens, 2012; Tomasello, 2003). For example, children associate semantic features of a “pusher” and “pushee” for the verb “push” (Tomasello, 2003). This process of creating and understanding these semantic features and abstract syntactic patterns facilitate early syntactic competence (Owens, 2012). Perhaps an animated verb illuminates the meaning of ECs by clarifying the semantic relations between symbols within a visual sentence (Boyer et al., 2012; Schlosser et al., 2014).
Given the communication needs of children with ASD and the frequent use of ECs as augmented input in the clinical setting, it is essential to find ways to improve the use of ECs in enhancing comprehension of language at the phrase/sentence level. Therefore, this study aims to explore whether children with moderate-to-severe ASD are able to follow directives in terms of accuracy (and speed) when presented with graphic symbol sentences that involve an animated verb.
Method
Participants, Setting, and Experimenter
A pool of nine participants were initially recruited from a single hospital-based autism center, of which five met the inclusion criteria (see Table 1). Inclusion criteria included: (a) formal primary diagnosis of ASD based on medical records; (b) moderate-to-severe autism severity as indicated by the Childhood Autism Rating Scale–Second Edition (CARS2-ST; Schopler & Van Bourgondien, 2010); (c) corrected or within-normal-limits hearing and vision, based on parent report or medical record; (d) chronological age between 4 and 10 years; (e) no prior systematic use of visual supports for instruction (i.e., as part of a formal therapy); (f) difficulty following spoken directives as indicated by parent report and/or intake questionnaire; (g) English as the primary language spoken in the home; and (h) a passing score on the static scene cue screening. A licensed speech-language pathologist served as the independent observer. The first author served as the experimenter. The study was approved by the hospital’s Institutional Review Board (IRB).
Characteristics of Study Participants.
Note. CARS2-ST = Childhood Autism Rating Scale−Second edition; ALP = Autism Language Program; CA = Chronological Age.
Based on parent interview, observation, and previous evaluation reports at the ALP.
Materials
Materials included an (a) iPad3®, (b) iPhone7®, (c) sentences, (d) props and figurines, and (e) visual cues. An iPad3 was used to present visual cues via Microsoft PowerPoint, while an iPhone7 (airplane mode) was used to video-record all sessions and measure participants’ speed of response via the Timer application.
The sentences for the static and animated ECs were presented alongside each visual cue and were based on the SVO structure. The subject and direct objects were selected based on their semantic relatedness and compatibility with the verb. The verbs were selected due to their early acquisition in typical language development (Fenson et al., 2007).
Props and figurines used to represent the various elements in the sentence included the following: Woody figurine, Gumby figurine, toy car, ball, ladder, horse, pencil, block, and apple. Each prop representing the direct object was selected based on their size in proportion to the main subject (i.e., Woody).
Visual cues included scene cues, photographs, and graphic symbols. Specifically, one static scene cue for the screening, one static EC for the pre-assessment, and 10 animated ECs for the main experimental task. The static scene cue was a photograph that illustrated the concept “Woody hug Gumby,” taken using an iPhone7 camera. Static and animated ECs were designed using photographs illustrating “Woody” (subject) and “Gumby” (object), and graphic symbols obtained from the Autism Language Program (ALP) Animated Graphic Set and Picture Communication Symbols (PCS). Animated and static symbols acquired from the ALP set were used for each verb, as previously described in Schlosser et al. (2012). Animated symbols were played automatically and looped continuously. During the allotted 10 seconds, the number of loops varied depending on the participants’ speed of response, with an upper limit of three loops when the 10 seconds elapsed. Graphic symbols from the PCS set were used to represent the direct object in the main experimental task.
Procedures
Autism severity screening
The CARS2-ST was administered to establish the participants’ symptom severity. The assessment was based on parent interview and child observation, and it addressed 15 functional areas. Participants were rated using a 4-point response scale in each of the functional areas based on frequency, intensity, peculiarity, and duration of the behavior (Schopler & Van Bourgondien, 2010).
Static scene cue screening
Participants were presented with one static SC and asked to perform the directive (“Do this”) using the provided props and figurines. They were allowed 10 seconds to complete the action with the props. The directive (“Do this”) was repeated once if the participant did not complete the action in time. Participants passed the screening and were included in the study if they correctly performed this directive with one static SC using all of the props within the allotted time and given prompts.
Static ECs preassessment
Participants were then presented with the static ECs (“Woody hug Gumby”) and asked to perform the directive using provided props and figurines. They were allowed 10 seconds to complete the action with the props, and the directive (“Do this”) was repeated once if the participant did not complete the action in time. Participants proceeded to the main experimental task regardless of their response (correct or incorrect answer).
Experimental task
Following completion of the static SC screening and preassessment, the main experimental task was administered. The experimenter sat adjacent to the participant for all three tasks. Participants were presented with 10 ECs containing an animated verb (see Table 2). Again, the experimenter pointed to the animated ECs and asked the participants to perform the directive using provided props and figurines (“Do this”). As before, the participants were given 10 seconds to respond by completing the action with the props and figurines. As previously, the directive was repeated once if the participant did not complete the action in time. Only the relevant props and figurines were placed on the table in front of the participant for each directive. A 5-s intertrial interval (ITI) was used before moving to the next directive. Participants were provided with intermittent nonspecific reinforcement (e.g., “You’re doing great,”“Good sitting”) and general prompts (e.g., “Look!”) throughout all three tasks to sustain participation and engagement.
List of the Animated ECs in the Main Experimental Task Along With Corresponding QR Code.
Note. Scanning the QR code leads to a video demonstration of the 10 graphic symbol sentences with animated verbs presented in the experimental task. For each graphic symbol sentence, a photograph of “Woody” is used as the agent, the ALP Animated Graphic symbols are used for each verb, and graphic symbols from the Picture Communication Symbol (PCS) set are used for the direct objects. Woody® is a registered trademark of Disney Enterprises Inc., Burbank, California, United States. ALP Animated Graphic Symbol Set was developed by the Autism Language Program at Boston Children’s Hospital. All rights reserved. Used with permission. PCS and Boardmaker are trademarks of Tobii Dynavox LLC. All rights reserved. Symbols used with permission.
EC = element cues. bALP = Autism Language Program.
Interrater Agreement and Procedural Integrity
An independent observer, a speech-language pathology graduate student unfamiliar with the study’s objectives, reviewed randomly selected video recordings from 40% of sessions (2/5 sessions). Interobserver agreement (IOA) was calculated for directive-following accuracy for all three tasks by the researcher who administered the tasks and the observer. IOA for directive-following accuracy was 100%. In addition, procedural integrity (Schlosser, 2002) was assessed using procedural checklists as described in Allen et al. (2018). Procedural integrity was 94% across 2/5 sessions.
Dependent Measures
Directive-following accuracy
The response was marked as correct if the participant completed the entire directive (SVO) by having the agent act upon the object as dictated by the verb within the allotted time and prompts. A response was considered incorrect if the participant gave no response, performed an action that differed from the symbol of the verb, or if the participant failed to carry out the directive within the allotted time. Accuracy was calculated as the percentage of directives followed correctly by dividing the number of correctly followed directives by the total number of directives and multiplying by 100.
Speed of response
Participants’ response was timed for each trial in the experimental task. Speed in seconds was recorded from the end of the examiner’s utterance (“Do this”) to when the participant completed the directive. Specifically, the stopwatch ended when the participant completed the action using the agent and object as dictated by the verb and ceased contact with the figurines/props. The stopwatch was restarted if participants were given the second directive (“Do this”). Participants’ speed of response was recorded under the directive (first or second prompt of “Do this”) that the correct response occurred. The speed of response mean for each participant was taken from only correct responses (i.e., dividing the total speed of correct responses over total number of correct responses).
Results
Due to small sample size, we examined the data descriptively and refrained from using statistical analyses. Table 3 summarizes the results for each of the participants. Regarding directive-following accuracy, two out of five participants correctly followed the directive in the pre-assessment. In the experimental task, all five participants correctly followed several directives, with the highest performance being 9/10 correct (i.e., Participant 5) and the lowest performance by Participant 3 with 1/10 correct, and a total mean score of 5.2/10 (52%) correct.
Directive-Following Accuracy for Static EC Preassessment and Main Experimental Task.
Note. EC = element cues; SVO = subject-verb-object; VO = verb-object; + = correct—SVO; √ = completed VO; − = did not complete/incorrect.
Post hoc examination of the results showed participants completed a portion of the directive (VO; verb-object). Their responses were considered incorrect but included in this table for discussion purposes. Please see discussion. bParticipants 3 and 5 correctly performed these directives but exceeded the allotted time limit of 10 seconds. Therefore, their responses were considered incorrect but included in this table for discussion purposes.
Post hoc examination of the results showed that there was variability in the degree to which participants followed the directives, with participants performing (a) the entire SVO correct, (b) verb-object (VO); participant acted as the agent himself), or (c) incorrect/did not complete the directive. For example, Participant 3 correctly completed one animated EC (SVO), three animated ECs as the agent (VO), and six animated ECs incorrectly.
Regarding the speed of response, as shown in Table 4, participants completed the directives in a total mean speed of 5.1 seconds with a range of 3.0 to 9.8 seconds. Participants 3 and 5 required additional time over the allotted 10 seconds ranging from 1.2 to 5.2 seconds to complete several of the directives.
Speed of Response for Correct Responses (SVO) in Main Experimental Task (Seconds).
Note. “First” and “Second” are the prompts. First refers to the initial spoken directive, “Do this.” Second refers to the additional directive (e.g., “Do this”) that was given if the participant did not respond. Please see “Main Experimental Task” under “Procedures” for more information. SVO = subject-verb-object; s = seconds; − = did not complete/incorrect.
Participants 3 and 5 correctly performed these directives but exceeded the allotted time limit of 10 seconds. Therefore, their responses were considered incorrect but included in this table for discussion purposes. Please see discussion.
In terms of severity of autistic symptoms and performance on the experimental task, several observations were noted pertaining to directive-following accuracy, speed of response, and level of prompting. First, severity of autism does not appear to be related to directive-following accuracy. Participant 1, who demonstrated the greatest autism severity (CARS2-ST of 40.5), was able to correctly perform 6/10 (60%) animated ECs. Comparatively, Participant 4, the participant with the lowest autism severity rating in this sample (CARS2-ST of 36.5), performed similarly with 7/10 (70%) animated ECs.
Second, participants who demonstrated greater severity completed the directives in approximately the same amount of time as less-severe participants. While the participant with the lowest severity rating (i.e., Participant 4; CARS2-ST of 36.5) demonstrated a speed of response mean of 3.6 seconds, a participant with one of the highest severity ratings (Participant 2; CARS2-ST of 40.5) demonstrated a comparable speed of response mean of 3.0 seconds. Finally, two of the participants that demonstrated the greatest autism severity (Participants 1 & 2) benefited the most from additional cueing. Participants 1 and 2 required a second prompt for at least one of their completed directives (see Table 4). Comparatively, the two participants with relatively less-severe autism (i.e., Participants 4 & 5) consistently produced the correct response on the first cue. Interestingly, there may be a relationship between speed of response and level of cueing.
Discussion
The results of this exploratory study suggest that the presentation of animated verbs within SVO-structured sentences led to varying degrees of successful directive-following. Three of the five participants who failed to complete the preassessment still demonstrated success in the experimental task. It is noteworthy that successful responses occurred without any instruction or prior exposure to animated ECs. Although one might speculate that the participants merely imitated the animated verb symbol, the participants followed the directives using props that were not explicitly represented in the animation. This suggests that they did not purely reenact the animated verb symbol but successfully interpreted it within the graphic symbol sentence using the provided props. Interestingly, several participants acted as the agent themselves. Perhaps the unequal proportion between the figurines/props influenced participants’ performance. It is also possible that the provided directive (“Do this”) was interpreted by the participants to act as agents themselves. This nonspecific instruction was chosen to ensure that the participants cannot follow the directive by understanding the spoken equivalents of the graphic symbols rather than having to interpret the graphic symbols themselves. Nonetheless, participants who acted as agents (VO) simultaneously completed several directives using all of the given props/figurines (SVO).
The level of severity was not strongly associated with directive-following accuracy and speed of response. In this study, participants with greater severity had similar performance to those of less-severe participants. Although this is a preliminary finding due to the small sample size, this is surprising because one might expect poor performance from participants with more severe symptoms. A study by Mayo et al. (2013), though of larger sample size and power, found lower language skills to be strongly tied to greater severity of autistic symptoms (Mayo et al., 2013). However, this study suggests that autism severity does not fully predict our participants’ language ability, but our sample size precludes a more definitive statement in this regard. The CARS2-ST rates children with ASD based on the frequency, duration, and intensity of a general target behavior and is derived from direct observations and parent interviews. Although target behaviors include verbal and nonverbal communication, the CARS2-ST does not specifically assess the child’s language abilities per se. Future investigations should strive to obtain more information about the child’s baseline language abilities by implementing more targeted formal/informal language assessments. For example, children with ASD who had higher receptive vocabulary assessment scores followed directives augmented with SCs more accurately than those with lower scores (Allen et al., 2018). Thus, obtaining more targeted information on one’s language abilities can provide insight into potential mechanisms that may make sentences involving animated verbs effective for children with ASD.
The observed differences in speed of response may be attributed to the participants’ cognitive abilities, specifically in the areas of attention, working memory, and processing speed. Prior studies have found that children with ASD demonstrate relative weaknesses in these cognitive skills (Hedvall et al., 2013; Oliveras-Rentas et al., 2012). Specifically, they have difficulty shifting their attention between visual stimuli and are unable to disengage from fixation of irrelevant stimuli to focus on visual tasks (Ciesielski et al., 1990; Elsabbagh et al., 2009; Keehn et al., 2013). Thus, this heightened nature of distractibility may increase the amount of time needed for processing incoming visual inputs (Courchesne et al., 1994; Keehn et al., 2013). In this study, Participants 3 and 5 demonstrated the most challenges related to the aforementioned skills, impacting their response time and accuracy. Several of their otherwise correct responses were not included because they exceeded the allotted time limit. For example, Participant 3 was observed to continually fixate on the physical prop (e.g., change the direction of Woody’s arm) and become distracted by his surroundings (e.g., turning lights on/off), which hindered his ability to shift his attention to the experimental task. Nonetheless, both participants responded positively to the animated ECs and this increased response time did not relate to their ability to correctly follow the directives. Therefore, deficits in these cognitive skills seen in children with ASD should be appropriately considered in language assessments/treatments to more accurately assess task performance. This can be done by allowing extra processing time for objective tasks, and minimizing visual and tactile distractions in the learning environment.
This study also found that participants with greater autism severity benefited the most from the second prompt, as shown by the performance of Participants 1 and 2. While they were able to give several correct responses to animated ECs with second prompt, they were unable to do so for the preassessment despite also receiving a second prompt. However, their difficulty in comprehending the preassessment could have been partially attributed to the specific concept presented—they may have been able to successfully comprehend other concepts represented by static ECs. Nevertheless, this is similar to the results of Allen et al. (2018), who found spoken directives augmented with static ECs to be ineffective in facilitating comprehension of spoken directives compared to other types of visual cues, further suggesting that the current clinical practice of augmented input is ineffective at facilitating comprehension of spoken language at the phrase/sentence level. Although this is an exploratory and descriptive study, our findings suggest that animation could be useful as a potential instructional tool to facilitate symbol syntax comprehension in children with ASD.
Limitations
This exploratory study is limited by its small sample size, making it difficult to extrapolate the results to a larger population of children with ASD, including those who are female. Another limitation is the absence of a condition in which participants were presented with ECs that involve static verbs. The preassessment only provided a small indication of the participants’ ability to understand static ECs. It is possible that they would have been successful following additional static ECs other than the one presented in the preassessment. Therefore, the lack of a control condition makes it difficult to ascertain the success with animated ECs solely to the type of representation (animated/static) of the verb. This study also presented with an unexpected procedural issue. As previously mentioned, Participants 3 and 5 required more than the allotted time to carry out the directives because of extreme attentional difficulties. When designing the procedures, the possibility of stopping and restarting the timer was not considered in instances where attentional limitations impacted participants’ performance. Future studies should consider increasing the time limit to 15 to 20 seconds, as response time is significantly affected by various cognitive and environmental factors. Alternatively, future studies might consider asking relevant stakeholders as part of social validation efforts (Schlosser, 1999) about the response time needed for a particular participant to be successful prior to finalizing the procedures (Schlosser et al., 1998). Another limitation is the use of written text alongside the symbol. Participants 4 and 5 were observed to read out loud the verb symbol prior to the initiation of the animation during the main experimental task. It is unknown what sight words the participants recognized, making it difficult to attribute their success with animated ECs solely to animation. That being said, ECs are typically presented with written text in the clinical setting and are therefore ecologically valid. Furthermore, the participants were reported to read sight words without comprehension and performed the directive only after watching the animation play out, suggesting that their success can be attributed to animation.
Conclusion
To our knowledge, this exploratory study is the first to investigate whether animated graphic symbols in SVO-structured visual sentences permit the completion of directives indicated by these sentences. Although it is unknown whether the animated verb was the critical element that helped participants implement several of the presented directives, the results of this study highlight the promising use of animation as a tool to facilitate comprehension of visual sentences. Future research with a control condition and a larger number of participants is warranted.
Supplemental Material
sj-tiff-1-cdq-10.1177_1525740120976332 – Supplemental material for Directive-Following Based on Graphic Symbol Sentences Involving an Animated Verb Symbol: An Exploratory Study
Supplemental material, sj-tiff-1-cdq-10.1177_1525740120976332 for Directive-Following Based on Graphic Symbol Sentences Involving an Animated Verb Symbol: An Exploratory Study by Nicole Choe, Howard Shane, Ralf W. Schlosser, Charles W. Haynes and Anna Allen in Communication Disorders Quarterly
Supplemental Material
sj-tiff-2-cdq-10.1177_1525740120976332 – Supplemental material for Directive-Following Based on Graphic Symbol Sentences Involving an Animated Verb Symbol: An Exploratory Study
Supplemental material, sj-tiff-2-cdq-10.1177_1525740120976332 for Directive-Following Based on Graphic Symbol Sentences Involving an Animated Verb Symbol: An Exploratory Study by Nicole Choe, Howard Shane, Ralf W. Schlosser, Charles W. Haynes and Anna Allen in Communication Disorders Quarterly
Supplemental Material
sj-tiff-3-cdq-10.1177_1525740120976332 – Supplemental material for Directive-Following Based on Graphic Symbol Sentences Involving an Animated Verb Symbol: An Exploratory Study
Supplemental material, sj-tiff-3-cdq-10.1177_1525740120976332 for Directive-Following Based on Graphic Symbol Sentences Involving an Animated Verb Symbol: An Exploratory Study by Nicole Choe, Howard Shane, Ralf W. Schlosser, Charles W. Haynes and Anna Allen in Communication Disorders Quarterly
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available on the Communication Disorders Quarterly website with the online version of this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
