Abstract
We conducted a meta-analysis of single-case research design (SCRD) studies on functional communication training (FCT). First, we used the What Works Clearinghouse (WWC) Standards to evaluate each study. Next, we calculated effect sizes using Tau-U. Then, we aggregated the effect sizes across the studies to produce an omnibus effect size. Results indicate that more than half of the SCRD studies met the WWC Standards and that FCT was effective in decreasing the level of problem behavior and in increasing the level of the alternative communicative response (ACR), but effectiveness varied according to such factors as type of disability and age. Furthermore, the results of visual analysis corresponded with Tau-U effect sizes in more than half of the cases. Implications for researchers and practitioners are discussed.
Individuals with disabilities engage very often in problem behavior, such as self-injurious behavior, aggression, or property destruction (Kurtz, Boelter, Jarmolowicz, Chin, & Hagopian, 2011). The occurrence of problem behavior in this population has been attributed to various factors, including the level of cognitive functioning, difficulties in acquiring conventional language, and the need for long-term levels of support (e.g., Drasgow, Lowrey, Turan, Halle, & Meadan, 2008). Regardless of the origin, problem behavior has a negative impact on an individual’s quality of life by limiting his or her participation in the natural environment, by restricting the development of social relationships, and by decreasing the level of effective independent functioning (Fox, Vaughn, Wyatte, & Dunlap, 2002; Machalicek, O’Reilly, Beretvas, Sigafoos, & Lancioni, 2007).
Functional communication training (FCT) is one intervention widely used to address the problem behavior displayed by individuals with disabilities (Carr et al., 1999). FCT begins with an assessment to determine the environmental variables that maintain the problem behavior (i.e., the function of problem behavior). Next, a differential reinforcement procedure is implemented in which an individual is taught to use a socially acceptable, alternative communicative response (ACR) that has the same function (i.e., produces the same outcome) as the problem behavior (Carr & Durand, 1985).
Numerous researchers have demonstrated that FCT can be effective for decreasing the problem behavior in individuals with various disabilities, including developmental disabilities (DD; e.g., Carr & Durand, 1985; Carr et al., 1999; Volkert, Lerman, Call, & Trosclair-Lasserre, 2009), intellectual disabilities (ID; e.g., Tiger, Hanley, & Bruzek, 2008), and autism spectrum disorder (ASD; e.g., Grow, Kelley, Roane, & Shillingsburg, 2008; Suess et al., 2013). Despite the large number of studies examining the effects of FCT on decreasing problem behavior for individuals with disabilities, questions remain about the extent to which the findings can be extended beyond the specific conditions of a single study. It can be difficult to predict how effective FCT examined in one study may be for other individuals or across different settings.
One approach to determine the overall effectiveness of FCT for other individuals and under different conditions than the ones examined in a single study is meta-analysis. Meta-analysis is a procedure consisting of aggregating and quantifying results from multiple studies related to an intervention to determine its overall effectiveness and the contribution of moderators (Richman, Barnard-Brak, Grubb, Bosch, & Abby, 2015). A moderator is a variable, such as age or disability, that influences the strength of the relation between an independent and a dependent variable (Kazdin, 2011). For example, a meta-analysis may reveal that FCT produces better outcomes for males than females and, thus, suggests that gender is a moderator.
We were able to locate only one meta-analysis that examined the effectiveness of FCT on problem behavior for individuals with disabilities. Heath, Ganz, Parker, Burke, and Ninci (2015) conducted a meta-analysis of 36 single-case research design (SCRD) studies published between 1980 and 2011. The authors calculated an effect size (i.e., robust improvement rate difference [IRD]; Parker, Vannest, & Brown, 2009) to determine the overall magnitude of FCT and the influence of three potential moderators: mode of communication, age, and disability. Results suggested that FCT was highly effective for individuals with disabilities (IRD = 0.86). Data also revealed that FCT was more effective (a) when vocal responses rather than aided alternative and augmentative communication (A-AAC; e.g., picture exchange) were taught; (b) when A-AAC rather than unaided augmentative and alternative communication (U-AAC; e.g., sign language) was taught; (c) for primary-aged individuals than for elementary-aged individuals; (d) for elementary-aged individuals than for adults; and (e) for individuals with ASD than for individuals with ID. Heath et al.’s (2015) study makes a valuable contribution to the literature by providing important information about for whom and under what conditions FCT is likely to be effective.
However, several aspects of the FCT literature and the Heath et al. (2015) study warrant further exploration. First, FCT typically addresses both problem behavior and an ACR. In their meta-analysis, Heath et al. reported a very large overall effect size, but did not report separate effect sizes for problem behavior and the ACR. The intervention may be differentially effective for problem behavior and for the ACR, and analyzing the impact of FCT on each may offer additional details related to the specific outcomes that can be anticipated from FCT.
Second, Heath et al. (2015) did not examine the methodological quality of studies on FCT, which may compromise the validity of their finding because some of the studies analyzed may have had poor internal validity (e.g., low interobserver agreement [IOA], poorly described procedures). Several tools have been developed to evaluate the methodological quality of research studies (e.g., Kratochwill et al., 2010; Wendt & Miller, 2012). For example, the What Works Clearinghouse (WWC) uses a gating procedure, whereby the methodological quality of each study is evaluated, and only those studies that meet quality standards are included in subsequent analyses (Kratochwill et al., 2010; Maggin, Briesch, & Chafouleas, 2013). The WWC provides a set of standards that can be used to determine the methodological quality of SCRD studies. To apply the WWC Standards, each study’s design is evaluated with respect to four design standards and classified as follows: “meets design standards without reservation,” “meets design standards with reservation,” or “does not meet design standards.” Next, for the studies that meet standards without or with reservation, the strength of evidence of a functional relation is evaluated as “strong evidence,” “moderate evidence,” or “no evidence” (Kratochwill et al., 2010).
We decided to use the WWC Standards to evaluate the methodological quality of studies on FCT for three reasons. First, the WWC Standards have been developed specifically to evaluate the empirical evidence collected through SCRD, the methodology employed to examine the effects of FCT studies. Consequently, they do not require adaptations of existing evaluation criteria applicable to group design studies that were previously used to evaluate the empirical base of FCT studies (see Kurtz et al., 2011 for more details). Second, the Standards provide objective and systematic guidelines that can be implemented with fidelity by researchers who have training in or are familiar with their implementation. Third, the Standards allow researchers to identify those studies that have strong internal validity and, thus, increase the credibility of the empirical evidence documented in a study.
Third, it is possible that publication bias influenced the results of the previous meta-analysis on FCT. Heath et al. (2015) analyzed only articles published in peer-reviewed journals. Negative or less successful outcomes would be less likely to be published and, consequently, are excluded from the review. The exclusion of unpublished dissertations may have influenced the findings of the meta-analysis by underrepresenting the outcomes of FCT. Unpublished dissertations that meet the WWC Standards would warrant inclusion in a meta-analysis.
Our purpose in this study was to extend the meta-analytic literature on FCT in several ways. First, we included peer-reviewed studies and unpublished dissertations to reduce the likelihood of publication bias. Second, we used the WWC Standards to determine the methodological rigor of the studies. Third, we used an effect size estimation method (i.e., Tau-U) that controls for positive baseline trends and has strong statistical power, thus increasing the probability of accurate conclusions related to the effectiveness of FCT. Fourth, we added new levels within the age moderator and new moderators (i.e., functional assessment method, function of problem behavior, and agent implementing the intervention). Fifth, we examined the overall magnitude of the effect of FCT and the influence of potential moderators on the problem behavior and on the ACR. Finally, we calculated the extent to which the Tau-U corresponded to the results of visual analysis. Our research questions were as follows:
To what extent do SCRD studies on FCT meet WWC Standards?
What is the overall effect of FCT on both problem behavior and ACRs?
What is the effect of the potential moderators (i.e., age, disability, topography of ACR, functional assessment method, function of problem behavior, and agent implementing the intervention) on the effectiveness of FCT?
To what extent does Tau-U correspond to visual analysis using WWC Standards?
Method
Literature Search and Initial Screening
To identity relevant studies, we first conducted an electronic search using ERIC, EBSCO, Education Full Text, JSTOR, PsycINFO, Education Research Complete, and Dissertations and Theses. We used the following search terms: functional communication training, FCT, functional communication, differential reinforcement of alternative behavior, DRA, and mand training. The literature search resulted in 177 studies (156 articles and 21 unpublished dissertations).
Second, we screened the 177 studies to determine whether they met the following seven inclusion criteria. First, the study had to be published in a peer-reviewed journal or be an unpublished dissertation conducted between 1985 and 2016. Second, the authors had to use a SCRD. At this stage, we included studies that did not have three attempts to demonstrate an effect at three different points in time (e.g., a reversal design with three phases (ABA) or a multiple baseline with two legs).
Third, the FCT intervention had to (a) be based on an analogue functional analysis (Iwata, Dorser, Slifer, Bauman, & Richman, 1982; Iwata et al., 1994), a trial-based functional analysis (e.g., Chezan, Drasgow, & Martin, 2014), or a descriptive functional assessment consisting of interviews and direct observation, conducted during the study or cited from a previous study and (b) teach an ACR. Fourth, the authors had to repeatedly measure the ACR in all phases of the study. Fifth, at least one participant in the study had to receive the FCT intervention. Sixth, the authors had to display the results for each participant receiving the FCT intervention. Seventh, the study had to be published in English.
We excluded conceptual or position papers, literature reviews, meta-analyses, group designs, nonexperimental designs (e.g., AB design), and duplicate studies. Seventy six studies, consisting of 65 published articles and 11 dissertations, met the inclusion criteria and were included in the next phase of the screening process. We completed an ancestral search of the 76 studies and meta-analyses and literature reviews on FCT to ensure that we identified all potentially eligible studies. We did not identify any additional studies through this search.
Next, we reviewed the 76 studies and excluded studies in which authors used a SCRD other than multiple-baseline design, multiple-probe design, and reversal or withdrawal design. We included only these designs because the WWC Standards provide clear and specific guidelines regarding the analysis of these designs. Furthermore, the WWC Standards lack specificity for some SCRDs (e.g., alternating treatments designs) and are nonexistent for others (e.g., changing criterion designs). We also excluded studies in which the authors’ main purpose was not to examine the basic effects of FCT (e.g., studies examining different schedules of reinforcement). Our review resulted in 44 studies consisting of 38 articles and six dissertations.
Application of the WWC Standards
We used the WWC Standards to evaluate the design quality for each participant included in the 44 studies. After analyzing the quality of the research design for each dependent variable for each participant in each study using the four design evaluation standards, we determined whether it met standards or not (Kratochwill et al., 2010). Then, we evaluated the strength of evidence for all cases that met standards (with or without reservations).
Design standards
The first WWC design standard related to number of phases and data points per phase. A multiple-baseline design had to consist of at least six phases and five data points per phase to meet standards without reservations and of at least six phases and three or four data points per phase to meet standards with reservations. A multiple-probe design had to meet additional criteria consisting of overlapping of initial baseline sessions, collection of data points prior to introducing the intervention, and probe sessions in baselines not receiving the intervention. A reversal design had to consist of at least four phases and five data points per phase to meet standards without reservation and of at least four phases and three or four data points per phase to meet standards with reservation.
The second standard was the systematic manipulation of the independent variable. To meet this standard, the study had to provide evidence that the researcher determined when and how the intervention was introduced to participants. The third standard pertains to the reliability of the dependent variable and IOA. To meet this standard, each dependent variable had to be measured by at least two independent observers, in each phase of the study and on at least 20% of the data points in each condition, and agreement had to be at least 80% or 0.60 using Cohen’s kappa. The fourth standard pertains to the number of attempts to demonstrate an effect of the independent variable on the dependent variable at different points in time. The researcher had to systematically manipulate the independent variable at least 3 times at different points in time to attempt to demonstrate experimental control (i.e., a functional relation). If this standard was not met, then the study did not meet standards.
We evaluated the design for 310 cases (i.e., 154 problem behaviors and 156 ACRs) for 116 participants in the 44 studies. A case was defined as one opportunity to demonstrate a functional relation between the independent and the dependent variable. Our evaluation indicated that, across 26 of the 44 studies, the design for 73 cases (i.e., 40 problem behaviors and 33 ACRs) met standards without reservations and the design for 71 cases (i.e., 38 problem behaviors and 33 ACRs) met standards with reservations. The design for the remaining 166 cases did not meet standards, and therefore were not included in the remainder of our analyses. The 26 studies (21 articles and five dissertations) are marked with an asterisk in the reference list.
Evidence standards
In the second phase, we evaluated the evidence of intervention effect using visual analysis for 144 cases (i.e., 78 problem behaviors and 66 ACRs) for 65 participants included in 26 studies. Each case was classified as “strong evidence” if it demonstrated three basic effects and no noneffects, “moderate evidence” if it demonstrated three basic effects and at least one noneffect, or “no evidence” if it demonstrated fewer than three basic effects. To determine the existence of a basic effect, we used visual analysis to determine the level, trend, variability, overlap, immediacy, and consistency of data within and across phases (Kazdin, 2011; Kratochwill et al., 2010).
Coding and Potential Moderators
We coded 16 variables for each study, including participant age, gender, and disability; topography of problem behavior; topography of ACR; recording system used to measure the dependent variables; type of functional assessment; function of problem behavior; type of intervention; intervention setting; agent implementing the intervention; social validity; type of SCRD; treatment integrity; maintenance; and generalization. From the 16 variables, we selected six potential moderator variables: age, disability, topography of ACR, functional assessment, function of problem behavior, and agent implementing the intervention.
Age
We classified each participant’s age as preschool (ages 0–5 years old), child (6–12 years old), adolescent (13–18 years old), and adult (ages 19 and older).
Disability
We coded each participant’s disability as ASD, ID, DD (developmental delay or developmental disability), emotional or behavioral disorder (EBD), or “Other.” We assigned a code based on the primary disability reported by the authors of the study.
Topography of ACR
We coded the ACR as A-AAC if it consisted of a communication device, picture, or card; U-AAC if it consisted of gestures or sign language; vocal if it consisted of vocalizations, words, or phrases; and multiple if it consisted of any combination of these.
Functional assessment
We coded the method used to identify the function of the problem behavior as functional analysis when the authors used an analogue functional analysis conducted in a controlled setting; as trial-based functional analysis when the authors conducted the analysis in the participants’ natural environment using a predetermined number of trials; and as descriptive when the authors conducted interviews and direct observation.
Function of problem behavior
We coded the function of problem behavior as access to tangibles, access to attention, escape from academic demands, escape from social interactions, or automatic reinforcement. When more than one contingency maintained the problem behavior, we used the code “multiple,” and when the authors could not determine the function of the problem behavior, we used the code “inconclusive.”
Agent implementing the intervention
We classified the agent implementing FCT in four categories: researcher (author or a graduate student), practitioner (teacher or staff member), parent (parent or caregiver), and therapist. The therapist code was used when the authors did not specify whether the intervention was implemented by a researcher or practitioner.
Effect Size Estimation
Phase contrasts and effect size calculation
We used Tau-U, which is an effect size estimation measure that combines nonoverlapping data points between A and B phases, with trend from within an intervention (or B) phase while controlling for the baseline (or A) phase trend (Parker, Vannest, & Davis, 2011). Tau-U controls for baseline trend, is distribution free, can be used with ordinal and interval data, is only partially influenced by autocorrelation, avoids ceiling effects, and has strong statistical power (Parker, Vannest, Davis, & Sauber, 2011).
We used GraphClick (Arizona Software, 2008) to extract the raw numerical data for each study before conducting a three-step process to estimate the effect size. First, we calculated Tau-U for each AB phase contrast for each dependent variable using the online Tau-U calculator (Vannest, Parker, & Gonan, 2011). Prior to calculating Tau-U for each contrast, we determined whether the baseline trend needed to be corrected by calculating Tau for the baseline phase. If the Tau value for baseline was greater than .20 (for the ACR) or less than –.20 (for problem behavior), we corrected for baseline trend in the calculation of Tau-U and Var-Tau for that phase contrast. Second, we obtained the standard error (SETau) for each Tau-U. Third, we calculated the overall effect size with standard error and confidence interval (CI) for each study by entering the Tau-U and SETau values in WINPEPI (Abramson, 2011). An effect size was considered “small” if it was lower than 0.20, “moderate” if it was 0.20 to 0.60, “large” if it was 0.60 to 0.80, and “very large” if the value was larger than 0.80 (Vannest & Ninci, 2015). Fourth, we used WINPEPI to calculate Tau-U, SETau, and CI for each level of each potential moderator.
Statistical Significance
We used a 95% CI (p = .05) to determine the statistical significance for Tau-U values. We also calculated an 84.3% CI for each level within a potential moderator variable and compared the CIs to determine if the difference between levels was statistically significant. Results were significant if the CI for each level did not overlap at the upper or lower limits when visually compared. The visual comparison of the two effect sizes with 84.3% CI is equivalent to the 95% CI (or p = .05) between two scores (Payton, Greenstone, & Schenker, 2003). A statistically significant difference between the Tau-U values for levels of a moderator (e.g., age) would indicate that FCT was differentially effective for one of the levels (e.g., preschool vs. elementary).
Correspondence of Tau-U Effect Sizes and Visual Analysis
To evaluate the extent to which Tau-U and visual analysis corresponded for each case in the meta-analysis, we assigned categorical scores of 1, 2, and 3 to the results according to Tau-U and to visual analysis to produce a common metric. Absolute Tau-U scores that were smaller than .20 were assigned a 1, between .20 and .60 a 2, and above .60 a 3, approximately corresponding to the small, moderate, and large effects, respectively (Vannest & Ninci, 2015). Visual analysis decisions of no evidence were assigned a 1, moderate evidence a 2, and strong evidence a 3.
We calculated the agreement between the two methods by subtracting the categorical score for Tau-U from the categorical score for visual analysis for each case for problem behavior and ACR to produce a difference score. The minimum absolute value of this score is 0; the maximum is 3. We calculated descriptive statistics for difference scores across the cases and the percent of cases on which the two methods had exact agreement and agreement within one, two, or three.
Interrater Agreement
Across all phases of the study, we used a point-by-point method (Kazdin, 2011) to calculate the percentage agreement scores by dividing the total number of agreements by the total number of agreements plus disagreements and multiplying the quotient by 100. In the screening phase, the first author served as the primary coder and the third author served as the secondary coder. An agreement was scored if both coders recorded either “yes” or “no” for each of the seven inclusion criteria. In the evaluation of the design quality phase, the first author served as the primary coder and the second author served as the secondary coder. An agreement was scored if both coders recorded the same information for each of the four design evaluation standards.
In the evaluation of the evidence phase, the coders were the same as in the previous phase. We compared the coders’ ratings related to the existence of an effect for each phase contrast within a specific SCRD. In the coding phase, the first and the second authors served as primary and secondary coders. An agreement was scored if both coders recorded the same information across the 16 variables.
In the Tau-U phase contrasts and effect size calculation, the first author served as the primary coder and a graduate student served as the secondary coder. Both coders independently determined whether baseline trend needed to be corrected, and they calculated Tau-U and Var-Tau for randomly selected phase contrasts. Agreements were scored (a) for baseline trend if both coders made the same determination about correcting baseline trend, (b) for Tau-U if both coders obtained the same value, and (c) for Var-Tau if both coders obtained the same value.
We collected interrater agreement for 23% (40 of 177) of the studies in the screening phase, for 46% (142 of 310) of the cases in the evaluation of the design quality phase, for 35% (50 of 144) of the cases in the evaluation of the evidence phase, for 35% (nine of the 26) of the studies in the coding phase, and for 35% (69 of 196) of the phase contrasts in the Tau-U calculation. Agreement was 96.6% (range = 85.5%–100%) for the screening phase. Point-by-point agreement was 99.7% (range = 90%–100%) for the evaluation of the design quality, 100% for the evaluation of evidence, and 100% for the coding. For the Tau-U calculations, agreement was 100% for baseline trend correction, 93% for Tau-U values, and 97% for Var-Tau values.
Results
WWC Standards
Of the 310 cases included in the initial pool, 73 (23.5%; 40 problem behavior, 33 ACRs) met WWC design standards without reservation. Another 71 cases (22.9%; 38 problem behavior, 33 ACRs) met the WWC design standards with reservations, and 166 cases (53.5%) did not meet standards. These 166 cases were not included in subsequent analyses. Of the 144 cases that met standards with or without reservations that were evaluated for the evidence of a functional relation, 79 cases (54.9%; 43 problem behaviors, 36 ACRs) provided strong evidence of the effects of FCT. Four cases examining the ACR provided moderate evidence, and 55 cases (38.2%; 32 problem behaviors, 23 ACRs) provided no evidence of the effects of FCT.
Overall Effect
The overall Tau-U effect size for FCT on problem behavior was .68 (SE = .02, 95% CI [0.64, 0.71], p = .05) and the overall Tau-U effect size for FCT on ACR was .65 (SE = .02, 95% CI [0.60, 0.69], p = .05), both of which are considered large effects (Vannest & Ninci, 2015). Tau-Us and the CIs were distributed between the upper and lower values without any outliers.
Potential Moderators
Age
For problem behavior, a large effect size was obtained for participants at the preschool, child, and adult levels, whereas the effect size for participants at the adolescent level was moderate (see Table 1 for effect sizes and CIs across all potential moderators). Data indicate a statistically significant difference between the preschool and adolescent age levels suggesting that FCT is more effective at reducing problem behavior for participants at the preschool level than for participants at the adolescent level. FCT seems to be equally effective in reducing problem behavior for participants at the preschool, child, and adult levels. For ACR, the effect size at all age levels was large. The difference between the preschool and child levels was statistically significant, suggesting that FCT is more effective for increasing ACR when implemented with children than with preschoolers.
Potential Moderator Variables.
Note. CI = confidence interval; LL = lower limit; ES = effect size (Tau U); SE = standard error; UL = upper limit.
Disability
A very large effect size was obtained for participants with other diagnoses, whereas for the participants with ASD, ID, and DD, the effect size was large. Data indicate that FCT was more effective in reducing the problem behavior for individuals with other diagnoses versus individuals with ASD. FCT seems to be equally effective for decreasing problem behavior when applied to individuals with ASD, ID, and DD. For ACR, the effect size was very large for participants with DD and other diagnoses, large for participants with ID, and moderate for participants with ASD. Results suggest that FCT is more effective for increasing the ACR of individuals with DD and other diagnoses than for individuals with ASD.
Topography of ACR
Only 54 of the 65 participants across the 26 studies were included in this analysis. We excluded 11 participants because the ACR targeted by the authors of the studies analyzed did not meet the criteria for inclusion in the final phase of the analysis. Forty-one percent (n = 22) were classified as A-AAC, 22% (n = 12) as U-AAC, 30% (n = 16) as vocal, and 7% (n = 4) as multiple. We obtained a moderate effect size for the ACR at the A-AAC level, a large effect size for both U-AAC and vocal levels, and a very large effect size for the ACR at the multiple levels. There was a statistically significant difference between the A-AAC and vocal and multiple levels suggesting that FCT is more effective in producing an increase in the ACR when the topography of the response is vocal or when multiple topographies are taught than when teaching an A-AAC response. Results indicate the FCT is equally effective in producing an increase in the ACR when teaching U-AAC, vocal, or multiple response topographies.
Functional assessment
A very large effect size was obtained for both functional analysis and trial-based functional analysis. For descriptive functional assessment, the effect size was moderate. Results suggest that FCT is more effective when based on an experimental method, such as analogue functional analysis or trial-based functional analysis.
Function of problem behavior
A large effect size was obtained for participants whose problem behavior was maintained by access to attention, access to tangibles, and by escape from academic demands or social interactions. For participants whose function of the problem behavior was multiple contingencies, the effect size was very large.
Agent implementing the intervention
A large effect size was obtained for problem behavior regardless of the agent implementing FCT with no statistically significant differences. For the ACR, data indicate a moderate effect size when the researcher implemented FCT, a large effect size when the practitioner or the parent implemented FCT, and a very large effect size when the therapist implemented FCT. Data suggest that FCT was more effective when the intervention was implemented by a therapist or practitioner.
Correspondence of Tau-U Effect Sizes and Visual Analysis
Of the 44 cases examining the relation between FCT and problem behavior, Tau-U and visual analysis produced the same categorical score on 61% (n = 27) of cases. The mean difference between the categorical scores for problem behavior was 0.48 (SD = 0.75). A disagreement occurred in 39% (n = 17) of the cases; however, in the majority of these cases, the categorical scores were within 1. In cases of disagreement, the visual analysis evidence rating was smaller than the Tau-U rating on 94% (16 of 17) of the cases. Of the 35 cases examining the relation between FCT and ACR, Tau-U and visual analysis produced the same categorical score on 69% (24 of 35) of the cases with eight cases having absolute difference scores of 1. The mean difference between the categorical scores for the ACR was 0.23 (SD = 0.72). Similar to problem behavior, the visual analysis rating was smaller than the Tau-U rating for most of the cases (73%; eight of 11) when the two methods did not have an exact agreement.
Discussion
We had several purposes for conducting this meta-analysis: (a) to analyze the methodological rigor of the identified studies using the WWC Standards, (b) to determine the overall effectiveness of FCT, (c) to evaluate the effect of six potential moderators on problem behavior and ACRs, and (d) to calculate the correspondence between visual analysis and Tau-U effect sizes. Results indicate that slightly more than half (26 of 44) of the studies on FCT met the rigor of scientific research according to the WWC Standards. However, the remaining studies (18 of 44) did not meet the requirements for design quality. This may be because some of these studies were conducted prior to the publication of the WWC Standards. In addition, different rubrics for evaluating methodological quality of SCRDs exist, and it is possible that the WWC Standards are more rigorous than other tools that we could have used (Wendt & Miller, 2012). Overall, our results support and extend previous findings, indicating that FCT produces large effects on problem behavior (Tau-U = .68) and the ACR (Tau-U = .65). Heath et al. (2015) reported a very large overall effect (IRD = .86). It is possible that we obtained a slightly smaller overall effect size because we eliminated studies that did not meet WWC Standards.
Our results suggest that FCT was effective in decreasing problem behavior and increasing ACR across ages. For problem behavior, the results for preschool-aged participants were statistically significant when compared with adolescents. It is possible that the small sample size of adolescents (n = 6) may have influenced the results by producing a smaller effect size compared with the other levels within this moderator that consisted of a larger sample of participants. Follow-up analyses are necessary to determine the influence of this moderator as a larger sample of this population becomes available. For the ACR, data reveal that FCT was highly effective across all ages, with a statistically significant difference between the preschool and the child levels. One potential explanation for this finding is the emphasis on early intervention services provided to children with disabilities. Specifically, implementing intensive communication interventions at a young age increases the probability of skill acquisition compared with when these interventions are implemented at an older age (Ganz et al., 2012; Heath et al., 2015).
An interesting finding relates to the fact that FCT appears to be more effective for participants with other diagnoses compared with participants with ASD. Specific participant characteristics (e.g., level of cognitive functioning) may explain this finding. This category included participants with various diagnoses, such as disruptive behavior, oppositional defiant disorder, or stereotypic movement disorder. Thus, it is possible that the participants with other diagnoses were functioning at a higher cognitive level compared with individuals with ASD, and this difference was reflected in the results. For example, a child with severe ASD and language delays may be less likely to acquire the ACR than a child with a milder disability. Another potential explanation may be the frequency and severity of the problem behavior. It is possible that a participant with oppositional defiant disorder may engage in more frequent and severe problem behavior than a participant with mild ASD, and thus, the intervention has a more dramatic effect on problem behavior for the former participant than the latter.
Data reveal that the effectiveness of FCT in teaching various topographies of ACRs ranged from highly effective (i.e., multiple), to very effective (i.e., U-AAC and vocal), and to moderately effective (i.e., A-AAC). Our findings are consistent with the previous results reported in the literature regarding the high effectiveness of FCT on producing an increase in the vocal topography of communicative responses compared with other topographies (Heath et al., 2015). In regard to the effectiveness of FCT when multiple topographies of ACRs are taught, the increased effectiveness of FCT may be the result of the combined effect of these forms. For example, researchers have argued that when an A-AAC response is taught during an intervention, there is a collateral increase in a vocal response for some individuals (Ganz et al., 2012; Millar, Light, & Schlosser, 2006). Future research is needed to compare the effectiveness of FCT in increasing the level of various topographies of the ACRs and to provide guidelines for practitioners on the selection of a response when multiple options are available.
Our findings also show that FCT is more effective when the function of the problem behavior is identified via experimental methods (i.e., analogue functional analysis, trial-based functional analysis). Because a small descriptive functional assessment sample size was included in this analysis, we cannot make definitive statements regarding the relative effectiveness of FCT when an experimental analysis versus a descriptive assessment is used. Additional research is needed to (a) further compare the effects of descriptive assessment versus experimental analysis on FCT as the body of literature grows and (b) examine the influence of the type of descriptive assessment method (e.g., interview, rating scale) on the effectiveness of FCT.
Although the results suggest that the intervention may be more effective when an experimental method is used to determine the function of the problem behavior, this finding may not easily be translated into practice. For example, functional analysis is labor intensive and time consuming, and requires extensive training to acquire the skills necessary to implement it with fidelity. Therefore, experimental methods are less likely to be adopted by practitioners in applied settings (Carr et al., 1999). Additional research is needed to examine critically the contextual fit of using experimental methods to identify the function of problem behavior in applied settings in which human and material resources may be limited. In addition, researchers should develop systematic guidelines for practitioners to promote the implementation of experimental methods, such as trial-based functional analysis, to be used in applied settings instead of a variety of descriptive assessment methods.
Data suggest that FCT was highly effective in reducing the problem behavior maintained by various contingencies, which supports the findings of previous research (e.g., Falcomata, Wacker, Ringdahl, Vinquist, & Dutt, 2013; Fisher, Greer, Fuhrman, & Querim, 2015). The results also suggest that FCT was more effective if the problem behavior was maintained by multiple contingencies than if the problem behavior was maintained by access to attention. Readers should be cautious when interpreting this finding due to the small sample of participants whose problem behavior was maintained by multiple contingencies.
Based on our results, FCT appeared to be highly effective in increasing the ACR when the intervention was implemented by the therapist, practitioner, or parent compared with when the researcher implemented FCT. This finding is difficult to interpret because the term “therapist” was often used without specifying that individual’s prior history with the participant. However, one potential explanation may be that social partners such as teachers, staff members, and parents are present in an individual’s environment throughout the day and have the option to provide multiple opportunities for practice as opposed to a researcher who spends a limited amount of time with an individual during an intervention session.
One of the novel contributions of this meta-analysis is the examination of the extent to which visual analysis corresponds to Tau-U. Our data indicate that in most cases, the two methods produced the same categorical rating. When a disagreement occurred, the rating for visual analysis was usually smaller than the rating for Tau-U. In other words, the visual analysts were more conservative than Tau-U in rating the strength of the evidence, particularly with regard to problem behavior. As a mathematical approach, Tau-U may be more sensitive to small changes than visual analysis. In general, the correspondence between the methods is promising and lends support to the use of Tau-U for quantifying SCRD data.
The results of this meta-analysis should be interpreted with caution due to several limitations. One potential limitation is the inclusion of unpublished dissertations. We included unpublished dissertations to avoid biasing the outcomes in favor of FCT; however, it may be argued that unpublished dissertations may not meet methodological standards of peer review for publication purposes. To mitigate this concern, we applied the WWC Standards and only included the studies that met these rigorous standards. A second potential limitation is the exclusion of cases that did not meet the WWC Standards for design quality. It is possible that excluding the findings of these studies distorted the obtained effect sizes and, consequently, influenced the validity of our conclusions. However, we eliminated these cases because we wanted to include only studies with strong internal validity based on a systematic evaluation of the study design. A third limitation is the exclusion of studies that used a SCRD other than multiple-baseline, multiple-probe, reversal, or withdrawal design. A fourth limitation is the small sample size for several levels within some moderators (i.e., adolescents, multiple topographies of ACRs, descriptive functional assessment, problem behavior maintained by multiple contingencies). A small sample size may skew the results of the analysis, and thus lead to inaccurate conclusions. Additional analyses are needed to replicate the procedures described in this meta-analysis as the body of the literature on FCT continues to grow.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
