Abstract

A meta-analysis by Tran, Sanchez, Arellano, and Swanson (2011) of the published RTI literature found that the magnitude of effect size (ES) between responders and low responders at posttest was significantly moderated by the pretest ES and the type of dependent measure administered, whereas no significant moderating effects were found in the mixed regression analyses for number of weeks of intervention, length of sessions, number of sessions, type of intervention (one-to-one vs. small-group instruction), and criteria for defining responders (cutoff, scores, discrepancy, benchmark). Overall, the synthesis questions whether the published evidence on RTI related to classifying responders and nonresponders at posttest has shown to be adequately separated from pretest learner characteristics. Stuebing et al. (2012) provided an excellent critique of this metaanalysis and raised at least four major issues related to interpreting the outcomes. We appreciate the opportunity to respond to these concerns.
Concern 1: Is the pre-post effect size analytic framework appropriate for a meta-analysis of intervention response?
An important issue that Stuebing, Fletcher, and Hughes (2012) raised was questioning our assumptions about the direction of effect sizes (ESs) related to differences between responders and nonresponders as a function of the treatment at pretest and posttest. As stated by Stuebing et al.,
Tran et al. (2011) used hierarchical linear modeling (HLM) to predict the posttest effect size (ES) from the pretest (ES), finding that the magnitude of the ESs increased in some cases from pretest to posttest. Thus, the data do not support the notion that posttest scores as a function of RTI provide outcomes independent of pretest scores.
In addition, Stuebing et al. commented,
It is difficult to see why effective interventions in any service delivery model, including RTI, would lead to predictions of ESs that are not larger at posttest than at pretest. In fact, the larger difference between adequate and inadequate responders at posttest seems consistent with typical RTI studies where students are rather homogenous at pretest (i.e., meet criteria for risk) but heterogeneous after intervention, with some responding adequately and others inadequately.
However, we feel that Stuebing et al.’s observation may have overlooked the context of our study in the introduction of the article. We asked whether children identified as responders and nonresponders at posttest were related to differences (the gap or ES) at pretest. The key ES we are referring to is the relationship of ES between responders/nonresponders at posttest and pretest within an experimental condition, and not the ES gain that occurs when a treatment condition is administered (although this was reported). More important, the focus of our synthesis was on the changes in the “variance” between posttest and pretest within the experimental condition and not the ES per se. We do not argue that mean levels of performance fail to change between responders and nonresponders from pretest to posttest, but rather we argue that the variance between the pretest and posttest should be reduced. Clearly, as we have stated elsewhere, “any reasonable treatment improves post-test scores” (Swanson & Lussier, 2001, p. 323) and within design studies have an upward bias because posttest standard deviations in some cases are inflated. However, we expected that if RTI procedures are controlling for the over-identification of reading disabilities (RD), there would be a reduction in the variance from posttest when compared to pretest. As we stated,
Under most intervention circumstances where there is no ceiling or floor effects, pre-test and post-test variable standard deviations are expected to be similar (see Hunter & Schmidt, 1990, pp. 250–252; also see Carlson & Schmidt, 1999, p. 853; for a review). However, one would expect that if the intensity of instruction identifies true responders from those with LD, then a significant reduction in standard deviations would occur at posttest. (Tran, Sanchez, Arrelano, & Swanson, 2011, pp. 284–285)
That is, if one of the key goals of RTI is to “reduce” the overidentification of children at risk of having a learning disability (LD), then the correlation between the two groups (ES) at pretest and posttest would be weak because the heterogeneity in the sample that exists at pretest would be reduced. This is based on the assumption that evidence-based treatments have been reliably administered and the majority of children with teaching deficits rather than disabilities are now “theoretically” responsive, thereby reducing the variance (overclassification) related to previous teaching outcomes. Thus, we assumed that greater variance should occur at pretest for the high-risk sample than at posttest. Stuebing et al. argue the opposite: Groups are more homogenous at pretest then posttest.
In our sampling of the published research, we did not find evidence for greater heterogeneity at posttest than pretest. In our study, the overall standard deviations were 1.04 at pretest and 1.11 at posttest, with an overall correlation between ES at pretest and ES posttest of .75 (p. 290). In fact, the standard errors across all measures at pretest and posttest were identical (SE = 0.03; see Tables 2 and 3 of Tran et al., 2011). This is not a strong argument for the increase of heterogeneity from pretest to posttest, as indicated by Stuebing et al. Given Stuebing et al.’s concern about including three studies they deemed as reflecting a bias at pretest, we did a follow-up by dropping these studies from the correlational analysis. After taking those studies out of the analysis, the mean ES and SD at pretest were 0.60 and 0.95, and the mean ES and SD at posttest were 0.76 and 1.14, and the correlation was r = .78. Again, this is not a strong argument that ES differences between the two groups at pretest/posttest are independent, nor does it fit RTI predictions related to the overidentification notion that the variance between the two groups is reduced.
We might also add that in contrast to Stuebing et al.’s observation, our overall findings of RTI studies suggest nothing out of the ordinary from what has been found in the treatment literature. For example, Carlson and Schmidt (1999), in their review of experimental designs, stated,
Whether treatment by subject interactions [italics added] (Cronbach & Snow, 1977) or different exposure to the treatment actually occur has been difficult to demonstrate empirically. Consequently, under most circumstances (i.e., in the absence of ceiling or floor effects), pre-training and post training dependent variable standard deviations are expected to be similar. (p. 853)
Thus, if this is the general pattern related to experimental interventions, one needs to question whether RTI studies are that unique in dealing with individual differences. The practical validity of RTI (evidence that something different is really happening) would be supported if the standard deviations related to pretest and posttest individual differences varied substantially from what generally occurs in the intervention literature. For example, if one argued that an intervention substantially increased individual differences in performance (as suggested by Stuebing et al.), then this would be reflected in systematically larger posttest standard deviations when compared to pretest standard deviations. This does not appear to be the general pattern found in the results of currently published RTI studies.
I believe the essence of our difference here is that Stuebing et al. presuppose that we selected articles showing differences at pretest, when in fact we selected articles reporting responders and nonresponders at posttest, then required in our selection of articles that pretest scores also be reported (see p. 285, Selection Criterion 5). No stipulations were made on what the pretest scores should look like. We deleted from our synthesis only studies that did not report any pretest scores. Although we recognized potential biases in published literature, it would be our contention that studies showing no differences between responders and nonresponders at pretest would also be just as likely to be in our synthesis as not (Stuebing et al. attest to this fact). Therefore, the sample of studies selected and outcomes related to author-identified responders and nonresponders within an experimental condition reflect the nature of the primary studies. Although Stuebing et al. suggest that these results should have been accompanied by simulations, this is appropriate only when we can make sense of the individual values composing the “input” to the analysis (e.g., corrected covariance matrices). Although the combination simulations and meta-analysis are more common in some domains (neuroscience; e.g., Ramsey, Spirites, & Glymour, 2011; also see Hadf & Willams, 2009, for biases in simulation studies with correlations) than, say, RTI studies, the reader needs to be aware that simulations are subject to several artifacts (e.g., pooling data across individuals, when those samples are from different sample distributions, etc.), and given the results of how our meta-analysis turned out, we would have some grave concerns at how the covariance structure would be established.
Concern 2: Can Cohen’s heuristic interpretation of d be used for dichotomized outcomes?
As stated by Stuebing et al.,
Effect size d . . . assumes that both groups are selected from the same population of individuals and that prior to intervention, these groups have the same population mean and standard deviation (SD). With randomization [italics added], the expected difference between the two means is 0 in the absence of a treatment effect.
Stuebing et al. then describe biases when selection occurs for variables established at various points on the normal curve at pretest. The reader needs to recall that the ESs between responders and nonresponders within a treatment condition were computed at two points (pretest and posttest) and not between control and experimental conditions (no doubt a problem with some of the primary studies in general). Forgetting for a moment that our comparisons were made not between a treatment and control condition (what they refer to as the treatment effect) but rather between groups in the same treatment, there are several ways to respond to the issues raised by Stuebing et al.
First, their assumption that subgroups within an at-risk sample, via a randomized design, should or could start at 0 at pretest (i.e., ES at zero) does not fit the data. On this issue, Shadish and Ragsdale (1996) make an interesting observation (in response to the findings on treatment outcomes of Lipsey & Wilson, 1993) related to random assignment versus nonrandomized studies and state that “on an average, however, randomized experiments yield an average standardized means difference statistic of d = .46 (SD = .28), trivially higher than the nonrandomized studies d = .41 (SD = .36), that is the difference is near zero” (p. 1291). More important, when they examined pretest differences, the correlation between pretest and posttest ESs was significant, being .53 for the studies with pretests (N = 81) and .39 for 54 randomized studies and .84 for the nonequivalent control group designs, that is, larger pretest ESs are associated with larger posttest effect sizes, in both designs (p. 1294). Thus, Stuebing et al.’s assumption is in conflict with some of the literature. Clearly, our data suggest that any classification of responder–nonresponder differences at posttest calls for an adjustment related to already-existing pretest differences. Simply, we would argue that posttest scores between responder–nonresponders within a treatment condition are in part a function of pretest effects (i.e., pretest sensitization as well as learner characteristics), and therefore an adjustment for pretest effects is critical when defining groups as responders at posttest (see Kim & Willson, 2004, for a related discussion).
Second, although the point made by Stuebing et al. is accurate related to the preferred research designs in the general scheme of treatments, these design characteristics were not apparent in the RTI studies that we reviewed. Although we accept the notion that a stronger case can be made for classification at posttest with extensive intervention, when we worked backward from these studies (i.e., studies were selected when responders/nonresponders were identified at posttest, and then we analyzed pretest performance), we found that this classification could be tied to pretest performance. A casual analysis of Table 1 in the Tran et al. (2011) article clearly shows that low responders performed well below responders on several normed-referenced measures at pretest. Although we agree with the assumption that children with RD are more accurately identified at risk after intense treatment, we found, however, that differences at pretest were fairly accurate at predicting differences at posttest.
Finally, the designation of responders versus nonresponders clearly implies a dichotomization or a categorical variable. Stuebing et al. critiqued our use of the d’ index on dichotomized data. No doubt Cohen’s d can be easily converted to a correlation coefficient (see What Works Clearing House [WWC] for other formulas as well as the limitations of correlations), and there are corrections when continuous data are dichotomized (Sánchez-Meca, Marín-Martínez, & Chacó-Moscoso, 2003), but the data reported in the RTI studies were in terms of separate means and SDs for responders and nonresponders at posttest. Some studies provided further gradations (mild responders, partial responders), but again there was a focus on the categorical variable of responders and nonresponders. We think their argument about whether Cohen’s d “as intended” and how we calculated Cohen’s d is convoluted. They make the curious argument that calculating group differences within a treatment conditions (as we did) is not as Cohen intended because the participants were not randomized at the beginning (or they did not start with treatment effects at zero). There is no discussion in Cohen (1988) that ESs can be interpreted as related to treatment only if randomization occurs. Obviously, the point of reference for our analysis on group differences within a treatment and their focus on “treatment effect” when compared to a control condition are different. Either calculation is testing a departure from the null hypothesis. As stated by Cohen,
Whether expressed as a difference between two population parameters or the departure of a population parameter from a constant or in an any other suitable way, the ES can itself be treated as a parameter which takes the value of zero when the null hypothesis is true and some other specific nonzero value when the null hypothesis is false. (p. 10)
Thus, we are testing whether two groups given the same intervention at pretest and posttest in the continuum is zero. Either way, the ES serves as an index for departure from the null hypothesis.
In general, Stuebing et al., as well as our research team, would agree that responsiveness is a continuous variable. In fact, we would argue that dichotomizing the data loses information and reduces statistical power, and potentially biases the estimates. Furthermore, in the grand scheme of things it jeopardizes the validity and efficiency of our meta-analysis because of the single cutoff point and/or inconsistent cutoff points of studies included in the synthesis. That being said, this was the primary data we had to work with.
Concern 3: Is phonological awareness a good predictor of response to intervention?
As stated by Stuebing et al.,
In a model where posttest effect sizes were predicted from pretest effect sizes as well as dummy coded vectors representing both the category of posttest variable and methodological variables (such as method used to determine intervention response), Tran et al. reported that the beta weight for the PA coded vector was close to 0 and was non-significant. Does this mean that PA is not important for reading acquisition?
We are not sure why this was the focus. It may be because other published syntheses have found this variable as critical, whereas other studies view its importance as overstated when other variables are included in the analysis. The critique by Stuebing et al. provides an interpretation of the outcomes related to the centering of our variables, especially dummy variables (the reader should also see Bryk & Raudenbush, 2002, p. 34, for a more in-depth rationale for centering binary variables). We do not disagree with the analysis by Stuebing et al. here. They provide the typical meta-analysis argument of mixing apples and oranges (that is why homogeneity was reported for the reader in Tables 2 and 3 of Tran et al., 2011), but their concern about centering seems odd to us. Centering, in all discussions we are familiar with on HLM analyses, calls for such procedures. Regardless, clearly not finding a significant parameter estimate for phonological awareness in predicting posttest outcomes in the full conditional model has to do with the magnitude of the standard error, as well as real word reading, word attack, passage comprehension, and rapid naming speed superseding (partialing out) the contribution of phonological awareness in predicting posttest.
Stuebing et al. do suggest that we made the analysis more complicated than necessary. We did provide zero-order correlations between pretest and posttest ESs, which we assumed told a great deal of the story. The difficulty was that significant error (random effects) existed among the studies, even when pretest and classification variables were entered into the analysis. Thus, some model testing was necessary to ensure that the relationships of pretest and other moderating variables to posttest ESs were not merely related to random effects. Even at that, we could reduce this random effect by only 76% in the full model. Thus, approximately 25% of the explainable variance was left unaccounted for.
Concern 4: Does this meta-analysis support the conclusion that RTI is not effective?
Stuebing et al. conclude that our “methods are not appropriate for determining the effectiveness of interventions based on RTI approaches, which is more appropriately determined via randomized control trials and syntheses of these studies.” We would not disagree with that statement if such studies existed at the time we did our meta-analysis. Stuebing et al. also state, “Tran et al. (2011) conclude that response to intervention (RTI) conditions were not effective at mitigating learner characteristics related to pretest conditions. The evidence presented in support of this assertion was that pre and posttest effect sizes were substantially correlated.”
They also focus on one of our sentences: “[U]nfortunately, the validity of RTI procedures, particularly in comparison to other assessment approaches, has not been adequately established in the present synthesis of the literature” (p. 293). They indicate that we went way beyond our data in this case. It is important for the reader to note that this sentence follows several paragraphs, and we were asked to provide a broader context to the findings by the anonymous reviewers. However, we would not necessarily back off from these general observations because we were unable to find any studies that made systematic comparisons with RTI to other classification procedures using randomized control conditions.
Stuebing et al. also state, “What is not clear is why Tran et al. (2011) did not conduct a much simpler meta-analysis of the correlations among pretests, posttests, and other individual characteristics measured at baseline [italics added]” (p. 9). Stuebing et al. are familiar with the studies we reviewed, and few actually provided a condition separate from the pretest conditions that captured “baseline.” A point made by Stuebing et al. is whether we took into consideration the timing of the classification (i.e., an assessment of the categorical label at different points along the intervention continuum). They are correct; we did not include this information in the analysis because it was redundant with classification criteria and length of treatment. Some studies (as indicated) may have defined children as responders and nonresponders somewhere between pretest and posttest, and we did not code this (as there were not enough studies). We did code intervention time and found that its effect (at least within this data set) was not reliable (nonsignificant) in its prediction of the magnitude of the posttest ES between responder and nonresponders. It seems to us, however, that the closer the time interval for classification of responder/nonresponder is to posttest, the stronger the correlation with posttest, which, again, is not a strong argument for increasing heterogeneity.
Final observation
A final point alluded to by Stuebing et al. was “what” the ES should be in defining the success of RTI in identifying responders/nonresponders. We might rephrase the question as follows: “What would be a reasonable ES benchmark for a valid separation of the groups (responders/nonresponders) when pretest and other methodological variables are partialed in the analysis?” We recognize some caution is necessary against applying ES with the same rigidity that one would typically use in a statistical significance testing. Because responders/nonresponders are usually determined within an experimental treatment, it would be necessary to take into consideration the design artifact for ES (this would be subtracted from the outcomes) as well as retest sensitivity (see Swanson & Lussier, 2001, p. 323, for discussion). Furthermore, Cohen (1988) intended the magnitude of ESs to serve as only a broad, general guideline, not to be used blindly. Unfortunately, the field of LD has not provided, to date, a consensus on the “benchmarks” for the overall, experimentally based ESs that includes the performance of responders and nonresponders in context. The best the field has to offer, to date, are comparisons between children with LD within an experimental condition and those with LD in a control condition. In the area of reading, the magnitude of the ES has not been impressive for children with RD. For example, the National Reading Panel (2000) reported that the evidence between reading (phonics) instruction and control conditions for students with RD varied on reading measures from .24 to .52, with a mean of approximately .33 (see Appendix E, 2-159). In short, there is a conundrum. The conundrum we confront (as quoted in Light, Singer, and Willet’s (1990) book titled By Design, is that “meta-analyses often reveal a sobering fact: effect sizes are not nearly as large as we might hope” (p. 195).
Regardless of these issues, we think one potential benchmark was established in an earlier meta-analysis of experimental studies (Swanson, 1999), which found that the overall ES for measures of word recognition for children with RD, partialed for methodological variations within pretest/posttest control group designs and type of treatment, was about .57 when comparing children with RD in the treatment and the RD control group. This is not to argue that this number serves as a benchmark for determining treatment effectiveness, but it does provide a rough approximation when interpreting the practical significance of the ESs.
Stuebing et al. also provide a reasonable critique, stating that we did not define what an adequate design might entail within an RTI framework. That was not the purpose of our synthesis. We would suggest, however, that a reasonable design to identify responders/nonresponder, as typified in the medical literature (e.g., Hewitt et al., 2011), may involve the following sequence: prescreening (finding children at risk), implementing an intervention that has an extensive evidence base, placing children in a maintenance condition to establish baseline, then randomly assigning participants to treatment conditions (conduct pretest here to ensure an adequate distribution or stratification of responders and nonresponders), and then a withdrawal period (posttesting) to determine stable responders and nonresponders. The design is more realistic in assuming that little variation in responders would occur in the “baseline” period but acknowledges extensive heterogeneity exists at the pretest conditions.
Conclusion
Clearly, Stuebing et al. have provided an excellent review of our meta-analysis. Although we disagree with them on several points, we think an appropriate conclusion is that more work is needed in this area. Their ideal study (randomization and no biases between responders/nonresponders prior to RTI) would have perhaps changed our outcomes and conclusions. But in meta-analysis there must be a reliance on the best evidence from the primary studies available. It is evident that as we look at the data, current interventions do not appear powerful enough to completely eliminate pretest differences for children at risk for RD. Perhaps more recent studies have addressed some of the design issues raised by Stuebing et al., as well as ourselves. Hopefully, as we update our analysis, this will be the case.
Footnotes
Acknowledgements
The author thanks Danielle Stomel for her comments on a draft of the article.
Author’s Note
Janette Klingner served as action editor on this interchange of the two articles.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
This article was supported by an Institute of Education Sciences (IES) Grants R324B080002 and R324A090002. The opinions expressed in this article do not necessarily reflect the opinion or policies of IES.
