Abstract
To date, researchers have not identified an efficient methodology for selecting items that will compete with automatically reinforced behavior. In the present study, we identified high preference, high stereotypy (HP-HS), high preference, low stereotypy (HP-LS), low preference, high stereotypy (LP-HS), and low preference, low stereotypy (LP-LS) items based on response allocation to items and engagement in stereotypy during one to three, 30-min free-operant competing stimulus assessments (CSAs). The results showed that access to HP-LS items decreased stereotypy for all four participants; however, the results for other items were only predictive for one participant. Reanalysis of the CSA results revealed that the HP-LS item was typically identified by (a) the combined results of the first 10 min of the three 30-min assessments or (b) the results of one 30-min assessment. The clinical implications for the use of this method, as well as future directions for research, are briefly discussed.
Historically, the assessment and treatment of automatically reinforced behavior has been challenging for clinicians because the consequences that maintain such behavior cannot be directly controlled (e.g., Lovaas, Newsom, & Hickman, 1987; Piazza, Adelinis, Hanley, Goh, & Delia, 2000; Rapp, 2008). Despite this general problem, noncontingent reinforcement (NCR) has been shown to be an effective intervention for stereotypy and other automatically reinforced behavior (for reviews, see LeBlanc, Patel, & Carr, 2000; Rapp & Vollmer, 2005; but see Rapp et al., 2013). Although researchers have used a variety of methods to empirically identify stimuli that are provided during treatment with NCR, an efficient method with high predictive validity has yet to be identified (DeLeon, Toole, Gutshall, & Bowman, 2005).
Competing stimulus assessments (CSAs) were developed to identify items that were both preferred and competed with automatically reinforced behavior. Most CSAs involve measuring an individual’s engagement in automatically reinforced behavior when a potentially preferred (or competing) stimulus is provided for a fixed period of time (e.g., 5-10 min; for an exception, see Derby et al., 1995). In one of the first studies on CSAs, Piazza, Fisher, Hanley, Hilker, and Derby (1996) used a single-item preference assessment (Pace, Ivancic, Edwards, Iwata, & Page, 1985) to determine the extent to which each item from an array of 12 to 16 items was preferred by two participants and competed with their automatically reinforced self-injurious behavior (SIB). Based on aggregated results for each item across ten, 30-s trials, Piazza et al. (1996) categorized items as (a) high preference, low SIB; (b) high preference, high SIB; or (c) low preference, low SIB. When the items were provided as the consequence for the omission of SIB, Piazza et al. found that neither the high preference, low SIB items nor the low preference, low SIB items decreased SIB for either participant; however, Piazza et al. did find that the high preference, high SIB items increased SIB for both participants. The authors concluded that the CSA successfully predicted the negative effects of the high preference, high SIB items, but incorrectly predicted the positive effects of other items.
The results of the Piazza et al. (1996) study are potentially limited insofar as the CSA only identified items that should not be used to treat the participants’ automatically reinforced behavior. Nevertheless, the CSA method described by Piazza et al. (1996) has been used in other studies to successfully identify items that competed with automatically reinforced behavior (e.g., Piazza et al., 1998). Consistent with the findings from Piazza et al. (1998) several subsequent studies have found that some individuals engage in higher levels of stereotypy in the presence of specific stimulation (Cook, Rapp, Gomes, Frazer, & Lindblad, 2014; Friman, 2000; Lanovaz, Rapp, & Ferguson, 2013; Rapp, 2005; Rapp et al., 2013; Van Camp et al., 2000). Thus, identifying stimuli that exert this effect on stereotypy is important for developing interventions.
Piazza et al. (2000) categorized items used in a CSA as either matched (those that produced the same putative sensory consequences as the target behavior) or unmatched (those that produced sensory consequences different from the putative products of the target behavior). This method identified items that were correlated with reduced levels of automatically reinforced behavior for each participant; however, the results were potentially limited insofar as the procedures were still very time-consuming (the assessment required approximately 60-150 min) because each item was evaluated in isolation. As an alternative to evaluating preference and competition during single-item trials, Ringdahl, Vollmer, Marcus, and Roane (1997) evaluated the predictive effects of a free-operant stimulus preference assessment (FOSPA) wherein participants were provided with multiple items. Using this approach, Ringdahl et al. (1997) hypothesized that access to preferred objects may decrease automatically reinforced problem behavior, when presented during environmental enrichment (EE) or differential reinforcement procedures, if engagement in the target behavior (as was measured by 10-s partial-interval recording [PIR]) during the FOSPA was (a) lower than in baseline and (b) lower than engagement with one or more items. Ringdahl et al. found that the results from the FOSPA predicted the effectiveness of preferred items presented during EE for all participants, but was less predictive when those items were presented during differential reinforcement treatment procedures.
The method employed by Ringdahl et al. (1997) is potentially limited in at least two ways. First, data on engagement in automatically reinforced behavior were aggregated for the entire session. Thus, as noted by Ringdahl et al., the presence of one or more items may have increased the participant’s engagement in automatically reinforced behavior and obscured the effects of EE. Second, it is not clear that converting frequency measures of problem behavior and duration measures of object engagement into a discontinuous measure (i.e., percentage of 10-s intervals) is an appropriate way to equate frequency events with duration events. Recent studies have found that measurement of duration events with 10-s PIR yielded false negatives and false positives when measuring changes in duration events with single-subject designs (Meany-Daboul, Roscoe, Bourrett, & Ahearn, 2007; Rapp, Colby-Dirksen, Michalski, Carroll, & Lindenberg, 2008; Rapp et al., 2007). Alternatively, it may be more appropriate to compare duration of engagement with items with duration of engagement in automatically reinforced behavior as a method of predicting the effects of the former on the latter during treatment.
Groskreutz, Groskreutz, and Higbee (2011) evaluated still another CSA method with one participant. Specifically, Groskreutz et al. (2011) conducted a paired stimulus preference assessment with eight items and then evaluated the effects of each item, alone, on stereotypy during 5-min sessions. Thereafter, the results of a treatment validation showed that high-preference items did not compete with the participant’s stereotypy. Instead, the items that competed most effectively with the participant’s stereotypy (hereafter denoted “high-competition” items) were either the third-most preferred or third-least preferred item from the array. This finding suggests that items need not be highly preferred to compete with stereotypy or other automatically reinforced behavior.
Experiment 1: Using a Free-Operant Stimulus Preference Assessment to Predict the Effects of Items on Automatically Reinforced Behavior
The purpose of this experiment was to extend the procedures described by Piazza et al. (1996), Ringdahl et al. (1997), and Groskreutz et al. (2011) by evaluating a new CSA method wherein engagement with items and automatically reinforced behavior is simultaneously assessed during a multiple-stimulus free-operant assessment. By comparing the percentage of time an individual engaged in the target behavior while he was simultaneously engaged with a particular item with the percentage of time he engaged in the target behavior for the whole session, it may be possible to predict the effects of the item on his stereotypy during a formal treatment evaluation. If multiple items can be evaluated in this manner, and the results have predictive validity, the time that is needed to identify high-competition items may be reduced. For example, a typical CSA with 5-min trials would require approximately 50 to 100 min of session time to evaluate 10 items once or twice, respectively. By comparison, when using a FOSPA, it may be possible to evaluate the same 10 items in approximately half the time.
Based on participants’ responding during FOSPA sessions, we categorized items by preference (high or low) and engagement in stereotypy (high or low). For example, a high preference, low stereotypy item would have been manipulated by a participant for a relatively high percentage of time with little or no simultaneous engagement in stereotypy. When presented alone, this item would be expected to decrease this participant’s engagement in stereotypy. By contrast, a high preference, high stereotypy item would have been manipulated for a high percentage of time with high levels of simultaneous engagement in stereotypy. Said item would be expected to increase this individual’s engagement in stereotypy.
Method
Participants and setting
Four individuals who were diagnosed with autism spectrum disorders participated in this study. Billy was a 5-year-old boy who was also diagnosed with Down’s syndrome. He followed simple instructions and requested preferred food items via a picture exchange communication system and manual signs. Jon was a 5-year-old boy who used one- to four-word sentences to request items. Ben was a 7-year-old boy who used short one- to three-word sentences to request items. Finally, Derek was an 8-year-old boy who requested preferred items using single words.
Sessions for all participants were conducted in their homes within rooms that were relatively devoid of alternative stimulation. Billy’s sessions were conducted in his bedroom, which contained two dressers, a bed, a bench, and a fan. Jon’s sessions were conducted in his bedroom. The bedroom contained two dressers, a bed, and two window seats. Ben’s sessions took place in the therapy room of his home. This room contained two child-sized tables, three child-sized chairs, a dresser, a white board (attached to the wall) and three bookshelves. Derek’s sessions took place in the basement of his home. The basement contained an assortment of storage bins that were placed in the middle of the room and along two exterior walls; Derek never attempted to access these bins.
Response measurement
We used the continuous duration recording (CDR) method described by Miltenberger, Rapp, and Long (1999) to score stereotypy and item manipulation during the CSA for all four participants. This same method was used to score the treatment validation sessions for Jon, Ben, and Derek; however, treatment validation sessions were scored using a 10-s PIR method for Billy. All experimental sessions were video recorded.
For Jon, Ben, and Derek, vocal stereotypy was defined as audible, acontextual vocalizations or noises including words or sentences not directed toward another person, not including breathing, coughing, or sniffing. For Jon, “mouth noises” (e.g., tongue clicking, whistling-like noises) were also included in this definition. For Jon and Ben, any three-word request directed toward the experimenter was not scored as stereotypy. Billy engaged in stereotypic chin-tapping, which was defined as contact of any part of Billy’s hand or an object held in his hand with his chin, mouth area, or jaw line.
For all four participants, item manipulation was defined as any part of the participants’ hands or feet contacting the item for more than 1 s. For Billy, music manipulation was scored any time when Billy shifted his weight back and forth from one foot to the other while music was playing. We adopted this definition because prior observations indicated that Billy always danced when music was playing. For the other participants, music manipulation was scored when any item produced music or other sounds. Each participant was provided with at least one toy that played music for a short duration (10-15 s) when a button was depressed; each participant was able to operate the buttons independently. Derek was the only participant for whom a video was included in the item array. Item manipulation for the video was scored when the auditory component of the video was present, as well any time when the participant’s head was oriented toward the laptop screen while he was seated in front of it.
Interobserver agreement (IOA)
The first author trained observers using the following procedures. First, observers were provided with the operational definitions of the target behaviors. Second, observers were required to verbally identify examples and nonexamples of the target behaviors. Third, observers were required to reach an accuracy criterion of 90% or higher when compared with an expert observer for two or more sessions (preliminary videos of the participants were used for training).
A secondary observer scored at least 33% or more of the baseline and treatment evaluation sessions for each participant. For Billy, IOA was calculated for the first 10 min of the 30-min FOSPA session. The IOA scores for data collected via 10-s PIR were calculated on an interval-by-interval basis by dividing the number of intervals with agreements by the number of intervals with agreements plus disagreements and multiplying by 100%. The IOA scores for data collected with CDR were calculated on a bin-by-bin basis for the average agreement with 10-s bin method (Mudford, Martin, Hui, & Taylor, 2009). The mean IOA scores for Derek’s stereotypy and item manipulation (all items combined) during the CSAs were 94% (range = 90.1%-97.8%) and 99.8% (range = 98.4%-100%), respectively. Derek’s mean IOA scores during the treatment validation were 95.6% (range = 94.3%-97.8%) for stereotypy and 97.12% (range = 90.3%-100%) for item manipulation. The mean IOA scores for Billy’s stereotypy and item manipulation during the CSA were 99% (one IOA session) and 99.4% (range = 97.5%-100%), respectively. The mean IOA score for Billy’s stereotypy during the treatment validation was 92.1% (range = 73%-100%). Item manipulation data were not collected during the treatment evaluation for Billy. The mean IOA scores for Jon’s stereotypy and item manipulation during the CSAs were 94.9% (one IOA session) and 99.8% (range = 98.5%-100%), respectively. Jon’s mean IOA scores during the treatment validation were 91.6% (range = 85.8%-93.8%) for stereotypy and 99.4% (range = 98.3%-100%) for item manipulation. The mean IOA scores for Ben’s stereotypy and item manipulation during the CSA were 97.2% (one session) and 99.2% (range = 97.2%-100%). Ben’s mean IOA scores during the treatment validation were 94.1% (range = 90.5%-96.2%) for stereotypy and 99.3% (range = 98%-100%) for item manipulation.
Baseline (no interaction)
For each participant, we conducted an initial no-interaction baseline phase to evaluate the persistence of his behavior in the absence of social consequences (e.g., Querim et al., 2013). Billy’s initial baseline sessions were conducted following the CSA. Baseline sessions for Jon, Ben, and Derek were conducted prior to the CSA. A minimum of three, 10-min baseline sessions were conducted with each participant. During baseline sessions, the participant was led into the room, the door was closed, the experimenter moved to the edge of the room, and the session began. No social consequences were provided for stereotypy. Participants’ attempts to leave the room, open the door, or manipulate other items in the room were blocked by the experimenter (these events rarely occurred).
CSA
We conducted one 30-min CSA for Billy, three 30-min CSAs for Jon and Ben, and six 30 min CSAs for Derek. During the first three CSA sessions (first array of items), Derek allocated nearly all of his responding to one item (a train), which was correlated with high levels of stereotypy. Thus, we conducted three additional CSA sessions with a second array of stimuli to identify items that may compete with his stereotypy. The procedures for each session were identical to those described by Roane, Vollmer, Ringdahl, and Marcus (1998) except that each session was 30 min in duration. Depending on the participant, 8 to 10 objects with varying sensory qualities were chosen based on informal observations, as well as the report of parents and trainers. For each participant, items were placed on a table; however, during the fourth, fifth, and sixth sessions (second array of stimuli) Derek’s items were placed around the outside edge of a 2.4 m × 3.0 m red area rug as many were too large to fit on a table. Each participant was prompted to manipulate each item for 3 to 5 s prior to the start of the session and thereafter he was provided with free access to the items for 30 min.
Unconditional percentages
The unconditional percentage of stereotypy was calculated by dividing the total number of seconds in which stereotypy was scored by the total number of seconds in the session (1,800 s) and multiplying by 100%. This measure indicates the percentage of time in which the participant engaged in stereotypy across a given session, regardless of what item was manipulated.
Conditional percentages
A conditional percentage of stereotypy was calculated for each item by dividing the total number of seconds in which the participant manipulated an item while he simultaneously engaged in stereotypy by the total number of seconds the participant manipulated that item (Lanovaz et al., 2013). This quotient was multiplied by 100% to produce the conditional percentage of time in which the participant manipulated a specific item while engaged in stereotypy. Conditional percentages of stereotypy for each item were compared with the unconditional percentage of stereotypy to determine if specific items increased or decreased the participant’s engagement in stereotypy (Lanovaz et al., 2013). Conditional percentages were calculated on a per-item basis. Thus, if a participant manipulated two items simultaneously during the CSA, a conditional percentage of stereotypy was calculated for each item separately but not for the combined manipulation.
Experimental design
The effects of two to four conditions (described below) on each participant’s stereotypy were evaluated in a multielement design with an initial baseline phase. All sessions were 10 min in duration. Experimental sessions were conducted up to twice a week; each condition was presented only once per day. Sessions were separated by as few as 1 day and as many as 10 weeks. The average number of days between sessions was 8.04 (range = 0-54) for Derek, 2.31 (range = 0-14) for Billy, 7.21 (range = 0-70) for Jon, and 6.24 (range = 0-75) for Ben. Session order was counterbalanced across days, and each session was separated by 15 to 30 min within a given day. Due to the need to conduct five sessions each day (four test conditions and a baseline condition) with a minimum of three sessions to detect stability for a given data path, we conducted a maximum of four sessions with each condition in the treatment validation. To aid in visual inspection of the results, data for each participant are depicted in a pair-wise fashion whereby each test condition is plotted against the baseline condition.
Procedures
Based on the results of the CSA, items were selected that were either low preference or high preference and produced either low levels of stereotypy or high levels of stereotypy. Levels of stereotypy were then measured in the treatment validation wherein each item was presented continuously and noncontingently, in isolation, to determine if results from the CSA accurately predicted the effects of each item on stereotypy.
Baseline
Baseline conditions in the treatment validation were identical to those in the initial baseline phase.
High preference, low stereotypy (HP-LS)
This condition was identical to baseline condition except that the participant was provided with noncontingent access to an item that was highly preferred and correlated with little or no stereotypy. The HP-LS item was selected based on the item (a) producing lower than average (unconditional) levels of stereotypy during the CSA (i.e., engagement in stereotypy was lower while the participant manipulated that object than the overall level of stereotypy throughout the CSA) and (b) being ranked in the top 50% of items presented during the CSA as indicated by percentage of item manipulation. For example, from an array of 10 items, if an item was ranked within the top five items in terms of time allocation, and it produced low levels of stereotypy, it may be selected as a HP-LS item. If there were three items ranking in the top 50% of the CSA, and each produced low levels of stereotypy, the item correlated with the least stereotypy was selected as the HP-LS item. We predicted that these items would decrease stereotypy. The same general rules were applied to the other items. The HP-LS item was placed on a small table in the middle of the room (or on the floor immediately next to the table for Derek); the participant was provided free access to the item. The other items were presented in the same manner within the respective condition (described below).
High preference, high stereotypy (HP-HS)
This condition was identical to the HP-LS condition except that the participant was provided with noncontingent access to the item identified as HP-HS based on the CSA. An item was categorized as an HP-HS item if it (a) produced higher than average (unconditional) levels of stereotypy during the CSA and (b) was ranked in the top 50% of items presented during the CSA.
Low preference, low stereotypy (LP-LS)
This condition was conducted if the items present during the CSA met the criteria for this category. The goal in evaluating this category was to determine whether a low preference item that produces low levels of stereotypy could potentially decrease stereotypy when presented alone, as was demonstrated by Groskreutz et al. (2011). This condition was identical to the previous two conditions except that the participant was provided with noncontingent access to the item identified as LP-LS based on the CSA. An item was categorized as a LP-LS item if it (a) produced lower than average (unconditional) levels of stereotypy during the CSA and (b) was ranked in the bottom 50% of items presented during the CSA.
Low preference, high stereotypy (LP-HS)
This condition was conducted if an item met the criteria for this category. This condition was identical to the other three test conditions except that the participant was provided with noncontingent access to the item identified as LP-HS based on the CSA. Based on the CSA, an item was categorized as a LP-HS item if it (a) produced higher than average (unconditional) levels of stereotypy and (b) ranked in the bottom 50% of items presented.
Results and Discussion
During the first three CSA sessions (Figure 1, first panel), Derek manipulated the train (M = 67.3%) the most, followed by the slinky (M = 2.1%), expanding ball (M = 1.2%), pom pom (M = 0.6%), wiggly ball (M = 0.4%), and the music toy (M = 0.02%). He did not manipulate the other items. The overall unconditional percentage of time with stereotypy during these assessments was 91%. Based on these results, the train was selected as the HP-HS item because it was the most preferred item and levels of vocal stereotypy when Derek manipulated it were higher (M = 95.77%) than the unconditional percentage of stereotypy (91%). Because only a HP-HS item could be identified, we conducted three additional CSAs with a second array containing some novel items to identify items that that may compete with stereotypy.

Percentage of time Derek engaged in item manipulation and vocal stereotypy while engaged with each item during the first array (first panel) and second array (second panel) of items in the CSA. The percentage of time Derek engaged in vocal stereotypy during the BL conditions compared with the HP-HS condition (third, left panel), the HP-LS condition (third, right panel), the LP-HS condition (fourth, left panel), and the LP-LS condition (fourth, right panel). Mean percentage of time of IM by Derek during the CSA (second array) and TV sessions (fifth panel).
Figure 1 (second panel) depicts the results from the second set of CSA sessions for Derek (Sessions 4, 5, and 6 combined). Overall, Derek engaged in vocal stereotypy for an average of 27.9% of the time. Derek allocated the most responding to the video (M = 81.8%), followed by the slinky (M = 30.6%), the train (M = 1.6%), and the bouncer (M = 1.6%). He allocated the least responding to the book (M = 0.01%), expanding ball (M = 0.3%), and scooter (M = 1.5%). The conditional percentage of time that Derek engaged in vocal stereotypy was higher than the unconditional percentage when he manipulated the bouncer (M = 65.1%), scooter (M = 70.9%), and expanding ball (M = 33.3%), and was lower than the unconditional percentage when he manipulated the video (M = 24.3%), slinky (M = 23.7%), and book (M = .01%). He did not manipulate the coloring stimuli. Based on the results of the second set of CSA sessions, the video was selected as the HP-LS item; it was the most preferred item and was correlated with low levels of stereotypy. The scooter was selected for testing as the LP-HS item as it was low preference and correlated with high levels of stereotypy in the second set of CSAs. Finally, the book was selected as the LP-LS item; it was the second-least preferred item and had the lowest level of stereotypy (0%). It should be noted; however, that the book was manipulated for only 1 s during all three 30-min assessments.
The third, left panel of Figure 1 shows the percentage of time Derek engaged in vocal stereotypy during baseline and HP-HS conditions. During the extended baseline phase, Derek’s vocal stereotypy persisted without social consequences (M = 70.1%). In the treatment validation, the percentage of time during which Derek engaged vocal stereotypy was higher in the HP-HS condition (M = 93.3%) than in the baseline (M = 82%) condition. The third, right panel shows the level of Derek’s stereotypy during HP-LS sessions, again compared with the baseline condition. Levels of his stereotypy were much lower in the HP-LS condition (M = 23.3%) than in the baseline condition. The fourth, left panel depicts the level of Derek’s stereotypy during the LP-HS condition compared with the baseline condition; there was no clear differentiation between LP-HS condition (M = 79.4%) and the baseline condition. The fourth, right panel also shows that levels of Derek’s vocal stereotypy were slightly lower in the LP-LS condition (M = 66.5%) than in the baseline condition (M = 82%). The results for Derek indicated that the HP-LS and LP-LS items decreased his stereotypy (though the former to a greater extent), the HP-HS item increased his stereotypy, and the LP-HS item did not affect his stereotypy.
Figure 1 (fifth panel) depicts the mean percentage of time Derek manipulated the HP-LS, HP-HS, LP-HS, and LP-LS items during the CSA and treatment validation. Derek manipulated each item for a higher percentage of time during the treatment validation sessions (i.e., when each item was presented alone) than during the CSA. Results show that Derek engaged with the HP-LS item for a high percentage of the treatment validation sessions. Nevertheless, as shown with the LP-HS item, high levels of item manipulation did not necessarily produce lower levels of vocal stereotypy for Derek.
Figure 2 (first panel) depicts the percentage of time that Billy manipulated each item and engaged in stereotypy (chin-tapping) while manipulating each item during the one 30-min CSA. Billy allocated the most responding to the bells (68.2%), followed by the texture ball (27.7%), the vibrating light toy (19.7%), and the music toy (6.1%). Overall, Billy engaged in chin-tapping for 4.2% of the session. The conditional percentage of time that Billy engaged in chin-tapping was higher than the unconditional percentage when he manipulated the clackers (14.3%), the push and spin toy (6%), and the bells (5.1%), and was lower than the unconditional percentage when he manipulated the music toy (2.8%), the vibrating tomato (0%), the vibrating light toy (1.7%), and the texture ball (3.2%).

Percentage of time Billy engaged in item manipulation and chin-tapping stereotypy while engaged with each item during the competing stimulus assessment (first panel). The percentage of 10-s intervals Billy engaged in chin-tapping during the BL conditions compared with the HP-HS condition (second panel) and the HP-LS condition (third panel).
Based on the results of the CSA, the texture ball and music toy were presented during the HP-LS condition for Billy because these items were his second- and fourth-most preferred, and each was correlated with low levels of chin-tapping. We used two HP-LS items with Billy (our pilot participant) because we wanted to maximize the likelihood that we would engage with the items for the duration of each session. Although the light toy yielded less stereotypy, the texture ball was used because the light toy was broken by Billy between sessions. Conversely, the bells were presented during the HP-HS condition because it was Billy’s most preferred item, and was correlated with slightly higher levels of chin-tapping. Although the bells were the highest preference item, we hypothesized that they would be less effective in reducing chin-tapping than the texture ball and music toy, due to the slightly elevated conditional percentage of stereotypy (5.1%) compared with the unconditional percentage of stereotypy.
Figure 2 shows the percentage of intervals that Billy engaged in chin-tapping during baseline and HP-HS conditions (second panel) and baseline and HP-LS conditions (third panel). Because he was our pilot participant, we did not conduct LP-HS and LP-LS conditions with Billy. During the extended baseline phase, the percentage of 10-s intervals with chin-tapping increased sharply across Sessions 5 through 8 (M = 36.4%). Data in the initial baseline phase show that Billy’s chin-tapping persisted without social consequences; the initially low levels of chin-tapping may have been a function of Billy’s prior history of receiving verbal reprimands for chin-tapping. In the treatment validation, the percentage of 10-s intervals during which Billy engaged chin-tapping was relatively high during the baseline (M = 70%) and HP-HS (M = 48.3%) conditions (second panel), but was near zero in the HP-LS (M = 4.2%) condition (third panel). The results for Billy indicated that the HP-LS item decreased his stereotypy and that the HP-HS item produced little if any change in his stereotypy.
Figure 3 (first panel) depicts the percentage of time that Jon engaged in vocal stereotypy while manipulating each item during three CSA sessions. Jon allocated the most responding to the foil blanket (M = 33.1%), followed by the music toy (M = 27.2%), the car toy (M = 10.9%), the book (M = 4.8%), and the microphone (M = 3.2%). He allocated the least responding to the texture ball (M = 0.02%), followed by the magna doodle (M = 0.1%), pom pom (M = 0.2%), wiggly ball (M = 0.3%), and the clapper (M = 2.8%). Overall, Jon engaged in vocal stereotypy for an average of 54.2% (unconditional) of the sessions. The conditional percentage of time Jon engaged in vocal stereotypy was higher than the unconditional percentage when he manipulated the foil blanket (M = 78.2%) and the book (M = 63.4%), and was lower than the unconditional percentage when he manipulated the music toy (M = 19%), car toy (M = 53.4%), microphone (M = 10%), clapper (M = 5.9%), wiggly ball (M = 11.8%), pom pom (M = 16.7%), magna doodle (M = 0%), and texture ball (M = .02%). Based on the results of the CSA, the music toy was presented during the HP-LS condition because it was the second-most preferred item and was correlated with low levels of vocal stereotypy. Conversely, the foil blanket was presented during the HP-HS condition because it was the most preferred object and was correlated with higher levels of vocal stereotypy. The texture ball was selected as the LP-LS item because it was the least preferred item, and was correlated with low levels (0%) of stereotypy. The assessment did not yield a LP-HS item for Jon.

Percentage of time Jon engaged in item manipulation and vocal stereotypy while engaged with each item during the CSA (first panel). The percentage of time Jon engaged in vocal stereotypy during the BL conditions compared with the HP-HS condition (second, left panel), the HP-LS condition (second, right panel), and the LP-LS condition (third panel). Mean percentage of time of IM by Jon during the CSA and TV sessions (fourth panel).
Figure 3 (second, left panel) shows the percentage of time that Jon engaged in vocal stereotypy during baseline and HP-HS conditions. During the extended baseline phase, Jon’s vocal stereotypy decreased across the first three sessions, so we conducted an additional four sessions during which his behavior increased (M = 46.2%). Data in this phase show that Jon’s vocal stereotypy persisted without social consequences. In the treatment validation, the percentage of time Jon engaged in vocal stereotypy was similar in the HP-HS (M = 45.2%) and baseline (M = 45.2%) conditions. The second, right panel shows that Jon engaged in less stereotypy during the HP-LS (M = 23.1%) condition than during the baseline condition (M = 45.2%); however, his stereotypy increased slightly across the HP-LS sessions. The third panel shows that levels of Jon’s stereotypy were comparable during the LP-LS (M = 48.3%) and baseline conditions. The results for Jon indicated that the HP-LS items decreased Jon’s stereotypy (at least initially), whereas the HP-HS and LP-LS items did not affect his stereotypy. It is possible that the increasing trend in the HP-LS condition was indicative of satiation for the stimulation generated by repeated access to stimulation produced by the HP-LS item.
Figure 3 (fourth panel) depicts the mean percentage of time Jon manipulated the HP-LS, HP-HS, and LP-LS items during the CSA and treatment validation. As with Derek, Jon manipulated each item for a higher percentage of time in the treatment validation than in the CSA. Also consistent with Derek, Jon engaged in high levels of item manipulation during the HP-LS sessions.
Figure 4 (first panel) depicts the percentage of time that Ben manipulated each item and engaged in vocal stereotypy while manipulating each item during the three CSA sessions. Ben allocated the most responding to the water bottle (M = 42.2%), followed by the magna doodle (M = 14.5%), the train (M = 1.6%), the worm (M = 1.2%), and the car (M = 0.7%). He allocated the least responding to the flower (M = 0.1%), followed by the duck (M = 0.1%), puzzle (M = 0.2%), and the pin toy (M = 0.3%). He did not manipulate the light wand during any of the sessions. Overall, Ben engaged in vocal stereotypy for an average of 43.8% (unconditional) of the time during the three CSA sessions. The conditional percentage of time that Ben engaged in vocal stereotypy was higher than the unconditional percentage when he manipulated the train (M = 50%), the puzzle (M = 50%), and the duck (M = 66.7%), and was lower than the unconditional percentage when he manipulated the water bottle (M = 32%), magna doodle (M = 25.7%), worm (M = 10.9%), car (M = 10.3%), pin toy (M = 0%), and flower (M = 0%). Based on the results of the CSA, the magna doodle was presented during the HP-LS condition because it was Ben’s second-most preferred item and it was correlated with less stereotypy than his most preferred item (water bottle). Conversely, the train was presented during the HP-HS condition because it was the third-most preferred object and was correlated with higher levels of vocal stereotypy. The pin toy was selected as the LP-LS item because it was one of the least preferred items, and it was correlated with low levels of stereotypy. Finally, the puzzle was selected as the LP-HS item because it was the third-least preferred item, and it was correlated with high levels of stereotypy.

Percentage of time Ben engaged in item manipulation and vocal stereotypy while engaged with each item during the CSA (first panel). The percentage of time Ben engaged in vocal stereotypy during the BL conditions compared with the HP-HS condition (second, left panel), the HP-LS condition (second, right panel), the LP-HS condition (third, left panel), and the LP-LS condition (third, right panel). Mean percentage of time of IM by Ben during the CSA and TV sessions (fourth panel).
Figure 4 (second, left panel) shows the percentage of time that Ben engaged in vocal stereotypy during baseline and HP-HS conditions. During the initial baseline phase, the percentage of time he engaged in vocal stereotypy was high and stable (M = 61.9%). Results from the initial baseline phase show that Ben’s vocal stereotypy persisted without social consequences. In the treatment validation phase, Ben engaged in comparable levels of vocal stereotypy during the HP-HS (M = 59.5%) and baseline (M = 66.2%) conditions. The second, right panel shows that Ben’s stereotypy was (with the exception of Session 14) generally lower during the HP-LS (M = 49.6%) condition than in the baseline (M = 66.2%) condition. The third, left and right panels show that levels of Ben’s stereotypy in the LP-HS (M = 72.4%) and LP-LS (M = 77.4%) conditions were not differentiated from the baseline condition. The results for Ben indicated that the HP-LS item decreased his stereotypy to some extent, whereas the HP-HS, LP-HS, and LP-LS items did not alter his engagement in stereotypy.
Figure 4 (fourth panel) depicts the mean percentage of time Ben manipulated the HP-LS, HP-HS, LP-HS, and LP-LS items during the CSA and treatment validation. As with Derek and Jon, Ben engaged with each item for a higher percentage of time during the treatment validation than during the CSA.
Although the results of this experiment indicate that this CSA method was effective for identifying items that decreased each participant’s stereotypy, the analysis required a considerable amount of time to conduct. Initially, we used multiple, extended sessions because we were interested in testing the predictive validity of each of the four stimulus categories. Nevertheless, the conditional percentage CSA method will not be a practical tool for clinicians who treat automatically reinforced behavior unless the HP-LS item can be identified using either (a) a single assessment with enduring predictive effects or (b) briefer sessions that can be updated on a regular basis. Therefore, we reanalyzed the results of the CSA for each participant to determine if HP-LS items could have been identified in Experiment 1 with briefer or fewer sessions.
Experiment 2: Identifying the HP-LS Item With Less Session Time
Rapp, Rojas, Colby-Dirksen, Swanson, and Marvin (2010) found that the results from extended 30-min FOSPAs were often predicted by the results from the first 5 min (62% of sessions), 10 min (55% of sessions), and 15 min (69% of sessions) of the respective session. In addition, Rapp et al. (2010) found that items identified as top-rated in the first 30-min assessment typically remained one of the top three preferred items across subsequent assessments. Thus, the purpose of Experiment 2 was to determine if the overall HP-LS stimulus was identified in the first 5 min, 10 min, or 15 min of each individual session, the combined data from the first 5 min, 10 min, or 15 min of all three sessions, or by the results from the first 30 min assessment. Specifically, we wanted to determine if the same HP-LS item could have been identified with multiple, brief sessions or with a single extended session.
Method
Participants
Data collected for Derek, Jon, and Ben in Experiment 1 were reanalyzed in Experiment 2 in a manner similar to that described by Rapp et al. (2010). Although Derek was exposed to two arrays of items in Experiment 1, we opted to only analyze his responding with the second array of items because it is likely that a clinician would have moved on to the second array after the first session with the original array yielded high levels of stereotypy. Data for Billy were not included in this experiment because he only participated in one CSA session.
Procedures
Each participant’s HP-LS item(s) from Experiment 1 is hereafter referred to as the “overall HP-LS” item. We conducted a within-session and a between-session analysis to determine if the each participant’s overall HP-LS could be identified with briefer or fewer sessions. First, each 30-min session was analyzed to determine if the overall HP-LS item was identified in that session. Results from this analysis indicate whether the same overall HP-LS item was predicted by the data from individual 30-min assessments. Second, we analyzed data from individual sessions to determine if the overall HP-LS item for each participant was identified during the first 5 min, 10 min, or 15 min of each session. Results from this analysis indicate whether data from multiple, brief sessions predicted the overall HP-LS item for each participant. Third, we analyzed the combined data from the first 5 min, 10 min, and 15 min of all three CSAs to determine the earliest point at which the overall HP-LS item was identified for each participant. The results of this analysis indicated whether the combined data from the three briefer (5 to 15 min) CSAs predicted the overall HP-LS items.
Results and Discussion
Table 1 provides a summary of the results from Experiment 2. The third column depicts the results from the first analysis in which each 30-min session was analyzed to determine if the overall HP-LS item would have been identified based on that session alone. Results showed that the overall HP-LS item was identified in 77.7% of all individual CSAs across participants. Specifically, Derek’s and Jon’s overall HP-LS items were identified in two of three individual 30-min CSAs, whereas Ben’s overall HP-LS items were identified in all three individual CSAs. The fourth, fifth, and sixth columns show the results of the second analysis, in which individual CSAs were analyzed to determine whether the overall HP-LS item would have been selected in the first 5, 10, and 15 min of each CSA session. Results indicated that the overall HP-LS item was identified in the first 5 min, 10 min, and 15 min of individual 30 min CSA sessions for 44.4%, 55.5%, and 66.7% of the sessions, respectively, across participants. Finally, the seventh column shows the results of the third analysis in which the combined data from three CSA sessions were inspected to determine whether the overall HP-LS item would have been identified by combining the first 5, 10, or 15 min across all three sessions (i.e., by conducting multiple, briefer CSA sessions). Results showed that each participant’s HP-LS item was consistently predicted by the combined data from first 5 min of the three CSA sessions.
Identification of the Overall HP-LS Item.
Note. HP-LS = high preference, low stereotypy.
Based on the findings above, we also reanalyzed data for the overall HP-HS from Experiment 1 items to determine whether using the combined data from the first 5 min of the three CSAs produced false positives. A false positive might occur if an item was identified as HP-LS during the combined first 5 min of three sessions but was in fact a HP-HS item (i.e., the overall HP-HS item) when data from all 30 min of these sessions were analyzed. Specifically, we evaluated whether each participant’s overall HP-HS item would have been identified as an overall HP-LS item based on the combined data from the first 5 min of the three CSAs. Indeed, false positives were produced with the HP-HS items for two of the three participants (Derek and Ben) when only the first 5 min of the three CSA were analyzed; however, the false positives were eliminated when we expanded the analysis to include data from the first 10 min of each session. Finally, we evaluated whether false positives were produced when only the first 30-min session was considered for each participant. The results indicated that a false positive was produced for one participant. Specifically, the train, which was ultimately identified as the overall HP-HS item for Ben, qualified as a HP-LS stimulus when only the first 30 min CSA session was considered.
Taken together, the results from Experiment 2 suggest that each participant’s overall HP-LS items was typically identified with (a) one 30-min CSA session or (b) the combined data from three 10-min CSA sessions. Thus, the results indicate that high-competition items were identified using approximately 30 min of assessment time. Although these are preliminary findings, this outcome suggests that the conditional duration CSA method has the potential to be a useful assessment tool for clinicians.
General Discussion
We used conditional percentages derived from a CSA to identify HP-HS, HP-LS, LP-HS, and LP-LS items. Subsequently, we evaluated the effects of these items on vocal stereotypy (Jon, Ben, and Derek) and stereotypic chin-tapping (Billy) and found that the CSA accurately predicted the effects of HP-LS items for all four participants. The CSA also correctly predicted that HP-HS items would increase Derek’s stereotypy; however, the effects of the HP-HS items were not correctly predicted for the other three participants. Similarly, the results from the CSA did not predict the effects of the LP-HS items for either Ben or Derek and the effects of the LP-LS items were accurately predicted for only one (Derek) of the three participants. Taken together, conditional percentages from CSAs correctly predicted that items from the HP-LS category would decrease stereotypy for all four participants; however, the results for the other three categories were less predictive.
The results from this study are consistent with those from the Piazza et al. (1996) study insofar as the HP-HS increased Derek’s engagement in automatically reinforced behavior. Based on the results in the Groskreutz et al. (2011) study, we expected that the LP-LS item in the present study would have more effectively reduced stereotypy when compared with the LP-HS item (it produced a modest reduction in stereotypy for only Derek). This finding may not be surprising, given that each of the low preference items (across all three participants) was manipulated for less than 3% of CSA sessions. This low percentage of item manipulation may not provide enough information to accurately predict the effects on stereotypy when the item is presented alone. In future research, it may be useful to stipulate a minimum level of item manipulation for inclusion in a LP condition.
The results of this study extend the literature by showing that conditional percentages from a CSA can be used to identify items that compete with automatically reinforced behavior. This study also contributes to the literature by showing that the highest preference items may not necessarily compete with automatically reinforced behavior. Finally, results from this study extended previous research by showing that combined data from multiple, brief (i.e., 10 min) CSAs may produce results similar to 30-min assessments, particularly when data from multiple brief assessments are aggregated.
Some limitations of this study should be noted. First, we did not compare the results of the conditional duration CSA method with those of a conventional CSA method (e.g., Piazza et al., 1996) to determine whether the two methods produce similar or disparate results with the same stimuli for the same participants. In addition, our criteria for categorizing items differed from those employed by Piazza et al. (1996). Nevertheless, the conditional duration method correctly predicted that the HP-LS items would decrease each participant’s stereotypy. Interestingly, by using the Ringdahl et al. (1997) CSA method of comparing the level of stereotypy during the no-interaction baseline with the level of stereotypy in the FOSPA (first array of items) for Derek, we confirmed Ringdahl et al.’s speculation about the possibility of an item increasing stereotypy and, thereby, potentially obscuring the effects of other items.
Another potential limitation is that the amount of time between sessions was not specifically controlled. Although this too was a departure from procedures described in other studies, it may actually provide stronger evidence for the generality of the conditional duration CSA method by showing that the effects of the HP-LS maintained over sizable periods of time (but see results for Jon in Figure 3). As a related limitation, we needed to use a second array of items (i.e., another 90 min after the first 90 min) to identify an item that competed with Derek’s stereotypy. We opted to continue with the first array for two additional CSA sessions because we expected his engagement with items to change across sessions. Rapp et al. (2010) found that participants’ top-ranked item from an initial 30-min preference assessment predicted that the item would remain the top-ranked item in only 39% of subsequent sessions. In practice, a clinician would have likely shifted to the second array after one CSA session with the first array.
Our response definitions for items producing auditory and audiovisual (video) stimulation were somewhat unconventional and, therefore, some discussion is warranted. These definitions were in part a function of simultaneously evaluating multiple events; however, informal observations of the participants’ engagement with the items also supported the use of such definitions. For example, Derek immediately clicked on a new video if a video stopped while he was looking away from the screen, suggesting that he was consuming the auditory component of the video even when was not touching it or looking directly at it.
From a clinical standpoint, two final limitations should be noted: First, the decrease in Ben’s vocal stereotypy during the HP-LS condition may not be viewed as clinically significant because his stereotypy continued to occur for 50% of the sessions. It is possible that greater reductions in Ben’s stereotypy could have been produced with an additional HP-LS item, a consequent intervention (e.g., response interruption and redirection; Ahearn, Clark, MacDonald, & Chung, 2007), or both. Nevertheless, the purpose of the study was to determine if conditional percentages predicted the subsequent effects of the items on stereotypy; to this end, the purpose was addressed. Second, in the present study, simultaneous item manipulation during the CSA was not analyzed. That is, conditional percentages of stereotypy may have been lower while a participant simultaneously manipulated two items than when he manipulated any single item. This information may be clinically helpful, particularly in cases where stereotypy continues to occur at moderate levels, as presenting these items may further reduce levels of stereotypy.
Future research should replicate the procedures described in this study and also evaluate the minimum CSA duration required to accurately predict the effects of items on automatically reinforced behavior. Specifically, future research should evaluate the generality of the procedures proposed by DeLeon et al. (2005) for determining the optimum session length for the conditional percentage method of CSA. Specific attention should be given to determining whether optimum session lengths differ for lower frequency behaviors.
Footnotes
Acknowledgements
The authors wish to thank Christine Eadon for her assistance in collecting data for this project.
Authors’ Note
Tyla M. Frewing and Sarah J. Pastrana are currently doctoral students at the University of British Columbia, Canada.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
