Abstract
Poor program implementation constitutes one explanation for null results in trials of educational interventions. For this reason, researchers often collect data about implementation fidelity when conducting such trials. In this article, we document whether and how researchers report and measure program fidelity in recent cluster-randomized trials. We then create two measures—one describing the level of fidelity reported by authors and another describing whether the study reports null results—and examine the correspondence between the two. We also explore whether fidelity is influenced by study size, type of fidelity measured and reported, and features of the intervention. We find that as expected, fidelity level relates to student outcomes; we also find that the presence of new curriculum materials positively predicts fidelity level.
An examination of program fidelity, or “how well an intervention is implemented in comparison with the original program design” (O’Donnell, 2008, p. 22), is considered critical to modern program evaluation. Estimates of fidelity can help confirm that changes in outcomes are in fact attributable to the program, increasing the internal validity of experiments and bolstering claims made about program efficacy (Dane & Schneider, 1998; Mowbray et al., 2003). Beyond providing methodological support, reports on program fidelity can provide substantive assistance to designers and practitioners in human service sectors, especially when scholars subject these reports to systematic review. Reviews that examine common reasons for implementation failure, for instance, can help program designers strengthen their product (e.g., Durlak & DuPre, 2008).
Estimates of implementation fidelity also can help explain null results, in particular distinguishing between the possibility that the program was not delivered as designed and other sources of failure, such as methodological problems, flaws in program theory, or lack of fit to local contexts (Dane & Schneider, 1998; Hohmann & Shear, 2002; Jacob et al., this issue, pp. 580–589; Mowbray et al., 2003). Yet whereas conventional wisdom in policy analysis often locates null results in implementation failure, we have no estimates of the extent to which this is true, particularly in recent, rigorous trials of educational interventions. To this end, we review the evidence regarding program fidelity in modern education research by analyzing classroom-level intervention projects funded by seven Institute of Education Sciences (IES) programs and a second set of studies identified during a meta-analysis of science, technology, engineering, and math (STEM) curriculum and professional development programs, projects funded primarily by the National Science Foundation. Specifically, we ask the following:
How often is program fidelity reported, and how is it defined and measured in recent educational program evaluations? What proportion of evaluations report low, moderate, and high fidelity?
To what extent is implementation fidelity predictive of program success?
To what degree is the level of implementation fidelity related to study size, the type of fidelity measured and reported, and features of the intervention?
We also qualitatively explore how often authors connect null results to poor fidelity and the explanations authors offer for a lack of fidelity. We describe our methods and results below.
Methods
Our analysis combines data from two samples of studies. The first involves IES-funded studies intended to change or improve K–12 classroom instruction. IES Requests for Applications (e.g., U.S. Department of Education, 2009, p. 60) require that awardees collect implementation data, and thus we searched each major IES program (effective teachers and effective teaching, mathematics and science education, reading and writing, social and behavioral contexts for academic learning, social and character development, teacher quality in math and science, teacher quality in reading and writing) for grants awarded from 2002 to 2011. We chose these dates because we thought it unlikely that projects funded after 2011 would consistently have publicly available evidence on implementation fidelity and project outcomes at the time the search was originally conducted, in 2016. We restricted our search to efficacy and replication (Goal 3) and scale-up (Goal 4) studies because of our interest in fidelity under realistic school and classroom conditions. Because too much variation in program design and clientele would lead to difficult-to-interpret results, we excluded studies that were not based in K–12 classrooms (e.g., tutoring or online learning, preschool) and studies focused on special populations (e.g., English language learners). This screen reduced the number of eligible projects to 42.
We located as many publications as we could find from each project, as authors often distributed student outcomes and implementation data over several papers. We then contacted principal investigators from grants with no publications to learn about study results and implementation metrics. In two cases these investigators were able to provide information about main impacts but had not completed implementation analyses. In seven cases, we either could not reach the principal investigator after repeated attempts or student impact results were not ready for release. Most reports focused on one intervention/program, but one described multiple treatment arms (Penuel et al., 2011). To accommodate this, we coded study design features (e.g., type of implementation reported) for each intervention but coded null results and fidelity separately for each treatment arm. In total, IES-funded studies contributed 35 reports containing 37 treatments.
Our second sample arises from a meta-analysis of preK–12 STEM curriculum and professional development interventions (see Lynch et al., 2019). In the initial round of screening, we downloaded 1,698 studies and examined their abstracts. Of these, 477 papers and reports met basic criteria for relevance and were advanced to the next round of screening. In this round, two raters examined each paper to determine whether it met our inclusion criteria. 1 After applying these inclusion/exclusion criteria, 42 papers and reports remained. As above, we identified all available reports from a project, then reviewed those reports for evidence about implementation. To prevent double-reporting, we excluded papers already included in the IES pool. One study (Heller et al., 2012) included three treatment arms. The STEM meta-analysis sample therefore contributes information about 37 reports and 39 treatments in total.
Our achieved sample of studies is thus a convenience sample, in some senses, derived from easily accessible IES reports and an existing pool of studies located for a STEM meta-analysis. Because the IES and STEM samples differed in the requirements regarding measuring implementation fidelity, we answer the first research question, on frequency of reporting, separately. Because many studies did report on implementation, we are able to answer Research Questions 2 and 3 with a moderate-sized data set.
Scoring and Analysis
Our coding system was simple, designed mainly to categorize IES and STEM study results for descriptive analysis. Our first and simplest code was whether fidelity was reported at all. Second, we recorded the method(s) used to assess fidelity (e.g., teacher surveys, classroom observations) and whether project researchers designed their own fidelity measures or relied upon those designed by other research teams. Third, we assessed the type(s) of fidelity measured. How to do so was not immediately obvious; scholars have generated many ways to conceptualize fidelity and an equally large number of ways to measure it, with some offering as many as five categories for reporting (Dane & Schneider, 1998; Dusenbury et al., 2003). Several scholars, however, distinguish between what we will call structural fidelity (adherence to program design regarding staffing levels, case load size, budget, procedures, frequency and intensity of contacts) and process fidelity (style, client-staff interactions, client-client interactions, individualization of treatment, climate) (Century et al., 2010; Mowbray et al., 2003; O’Donnell, 2008). These scholars also call out dosage fidelity (Dane & Schneider, 1998; Dusenbury et al., 2003), which records whether a program was actually accessible to those meant to implement it. We adopted these categories and modified them to fit educational interventions. We coded positively for structural fidelity when the authors provided evidence on classroom-level compliance with program-specific elements, such as the use of project curriculum units, adherence to project-supplied lesson plans, and the deployment of program-specific instructional behaviors (e.g., worked examples featuring a particular sequence of teacher questions) with no attendant focus on quality. Typically, authors collected these data only from treatment-group classrooms. We coded positively for process fidelity when authors measured more complex and general classroom-level outcomes, such as teacher sensitivity to student learning needs, mathematical discussions, classroom climate and student behavior, and the cognitive challenge of student tasks. Typically, authors collected these data from both treatment and control classrooms. Finally, we coded yes for dosage fidelity when authors provided evidence on the extent to which teachers received a treatment (e.g., descriptions of attendance at professional development, whether curriculum materials were delivered to teachers in a timely manner). Structural fidelity typically measures adherence to program elements, process fidelity is a form of intermediate impact, and dosage fidelity measures teacher opportunity to learn or to use program materials.
We measured the extent of implementation fidelity using two purely quantitative indicators as well as a holistic, more qualitative metric. Our quantitative measures consisted of the following:
The proportion of positive structural fidelity results reported by authors. We considered structural fidelity metrics positive when authors observed 50% compliance with project activities.
The proportion of positive process fidelity results reported by authors. Because process fidelity was typically reported as a treatment-control contrast, we considered process fidelity metrics positive when authors observed positive and significant results of this treatment-control test.
Our holistic measure took into account outcomes from these quantitative fidelity metrics but also relied upon other sources of information about overall fidelity. Specifically, we considered both dosage fidelity, typically reported descriptively, and authors’ comments about fidelity of implementation. We also weighed process fidelity results over structural fidelity results when the two conflicted. We assigned a score of “low fidelity” when less than half of the structural and process fidelity codes were positive, a score of “high fidelity” when more than 80% of the structural and process fidelity codes were positive, and a score of “medium fidelity” for those in between. Both authors coded each report included in the review then met to reconcile disagreements. We recognize that this coding system requires a fair amount of judgment, but a more deterministic coding system was impossible in light of the different fidelity measures used by study authors.
To allow us to qualitatively understand the relationship between implementation fidelity and null results, we developed a rudimentary metric for assessing whether results from a study were null. For each project we calculated the fraction of total impact estimates, aggregated across all available papers and reports, that were positive at significance levels of at least p < .05. We used this estimate of the proportion of positive impacts in some of our models. We also created a binary measure by categorizing studies with less than 50% positive effects as null results studies—an arbitrary threshold but one reflective of current hopes for consistent and positive results in the field. Authors and two research assistants double-coded each study, discussing and resolving discrepancies where they arose. One potential issue with this method for determining null results is that it does not distinguish between more and less central program outcomes—for instance, when a program expects a strong impact on executive function and weaker impacts on student achievement. In practice, however, few studies prioritized outcomes in this way, with many reports containing two outcomes (e.g., a researcher-developed and a standardized measure) without information about which researchers valued more.
We also coded for a number of program and study design features that might affect fidelity. Our first study design feature was sample size; because we expected fidelity may be lower in studies with larger sample sizes, we recorded the number of teachers in the treatment and control groups combined. As noted above, we coded the method for collecting fidelity evidence, structural versus process fidelity, and whether the researchers designed their own fidelity metric. Among program characteristics, we coded for whether the program featured curriculum materials, professional development, and/or coaching and noted the maximum number of hours teachers could have experienced the coaching and professional development.
To answer our first research question, regarding how often program fidelity is measured and reported, we calculated simple descriptives and crosstabs. To answer our second research question, regarding the relationship between implementation fidelity and program success, we generated both crosstabs and a regression of the fraction of null results over fidelity level, using study characteristics as controls. Finally, to answer Research Question 3, we generated two regression models linking study and program characteristics to fidelity levels. To link study characteristics to fidelity level, we used a multilevel model; the multilevel model accounts for the nesting of multiple fidelity assessments within treatments. Because the model regressing our holistic variable as an outcome did not converge in this multilevel model, we used the ratio of positive fidelity outcomes in its place. In our second model linking fidelity to program features, we used ordinary least squares regression because program features were not nested within treatments.
Results
Program Fidelity Measurement, Reporting and Results
Program fidelity was reported in 97% of projects arising from IES grants and 74% of projects in the STEM pool (see Table 1). One project in the IES pool suggested that implementation data were collected but did not provide results. Across both study pools, structural fidelity was measured in 54% of projects, and process fidelity was measured in 50% of projects. Eighteen projects (24%) across both sources measured both process and structural fidelity. Dosage fidelity was reported in 26% of projects. For studies reporting fidelity, the most frequently used method for gauging fidelity was classroom observation, with 46% of projects reporting this data collection technique; teacher self-reports (typically logs and surveys) followed behind, with 29% of projects using this technique; 18% of studies used both (see Table 2). Our read of the studies also suggested scattered use of other methods, such as teacher interviews, student surveys, and periodic check-ins between teachers and study staff. Most projects evaluated interventions against original program design; in fact, we found no project that directly measured users’ adaptations of the program, despite scholarly interest in this topic (Blakely et al., 1987; McMaster et al., 2014; O’Donnell, 2008; Quinn & Kim, 2017).
Types of Fidelity Reported by Source
Note. IES = Institute of Education Sciences; STEM = science, technology, engineering, and math. Percentages reported are of the total set of studies and do not sum to 100 because many studies reported more than one type of fidelity. All overlapping studies are included in the IES-funded count here because of the IES requirement to report fidelity.
Method of Gauging Fidelity
Note. N = 65. Of the 65 studies that measured fidelity, 4 measured only dosage fidelity, so percentages do not sum to 100. Classroom observations include in-person, videotaped, and one instance of audio-recorded observations. Teacher self-report includes daily activity logs or postintervention surveys.
Of 65 projects that presented quantitative fidelity data, 26 (40%) included evidence of strong implementation (see Table 3, Row 3). For instance, Lara-Alecio et al. (2012) reported that treatment teachers earned 108 of 124 points on a structural fidelity metric, and Schwartz-Bloom and Halpin (2003) reported that most teachers used at least three of the four curriculum modules developed as part of their project. Twenty-five (38.4%) projects included evidence of moderate implementation. For instance, Star et al. (2015) reported that many treatment-group teachers used the newly developed curriculum materials with some aspects of structural fidelity (e.g., when using the curriculum, teacher displayed learning objective; teacher summarized major points from student discussion) but also reported that roughly one-fifth of teachers did not use those materials at all. Finally, 14 (21.5%) projects included evidence of weak or nonexistent implementation. For example, Lang et al. (2014) report no treatment-control contrast in the use of formative assessment practice as recorded in classroom observations.
Fidelity Ratings by Type of Fidelity Measured
Note. N = 65. Some studies reported more than one type of fidelity, so row frequencies do not sum to the total N. A study was scored “low fidelity” when less than half of the structural and process fidelity codes were positive, “high fidelity” when more than 80% of the structural and process fidelity codes were positive, and “medium fidelity” for those in between. Ratings also factored in dosage fidelity, when reported, and any additional author comments about fidelity.
Does Fidelity Predict Program Success?
Table 4 shows results from our holistic fidelity–level metric by whether we categorized the study as having null results. Overall, we classified 27 (35.5%) treatments as producing null results, a figure much lower than the 91% null-result rate found by the Coalition for Evidence-Based Policy (2013). Studies coded as moderate or high fidelity had more than double the chance of yielding positive results than null results; for studies coded as low fidelity, the odds of a positive and null categorization were about even. There appeared little difference between moderate and high fidelity ratings, in terms of the likelihood of positive results. Table 4 also shows that fidelity, at least as we have defined it, is not deterministic of program outcomes. Six studies had majority-positive impacts yet low fidelity ratings, and another eight had strong fidelity ratings but majority-null impacts.
Fidelity Rating by Student Impacts
Note. N = 76. Studies were coded null if they had less than 50% student impacts that were positive with significance levels of p < .05.
To further understand the relationship between fidelity and null results, we regressed the fraction of positive student impacts over both our holistic fidelity rating and controls, including the teacher sample size and the type of assessment used to measure student learning. Because neither Table 4 nor exploratory regressions revealed a difference between moderate- and high-fidelity studies in terms of the likelihood of positive outcomes, we simplified our fidelity measure to a dummy variable representing low fidelity (see Table 5). We found a significant relationship between the dummy variable representing low-fidelity implementation and student outcomes; treatments with low fidelity averaged 24% fewer positive outcomes than those with moderate or strong fidelity. We also observed that number of teachers in the study, included in our model as the teacher sample size divided by 100 to enhance intepretability of the cofficient, had a small but statistically significant negative relationship to the fraction of positive results; a treatment one standard deviation above average in teacher sample size (442 teachers) had, on average, 6.6% fewer positive impacts than a program with an average-sized sample (167 teachers). In line with C. J. Hill et al. (2008), researcher-designed assessments were also more likely to post positive impacts as compared to standardized assessments (shown) and studies that used both standardized and researcher-designed measures (the referent variable). Separately, we also examined the likelihood of null results by content area, including STEM, reading/writing, and social/behavioral interventions, and found no relationship (not shown).
Results of Ordinary Least Squares Regression Analysis for Variables Predicting Percentage of Positive Student Achievement Outcomes (N = 76)
Note. Standard errors are in parentheses.
p < .05. ***p < .001.
To complement our quantitative analysis, we examined reports from null-report studies to see the extent to which authors indicate that implementation may have contributed to a lack of impacts. We saw that 5 of the 27 null-result treatments (Gersten et al., 2010; Jacob et al., 2017; Matsumura et al., 2012; Santagata et al., 2011; Schneider & Meyer, 2012) identify dosage fidelity—teachers’ receipt of an appropriate amount of professional development—as problematic, although one of those studies (Gersten et al., 2010) reports generally strong dosage and classroom fidelity. Only 7 of 27 null-result treatments (Borman et al., 2008; Cavalluzzo et al., 2014; Dominguez et al., 2006; Hurtig, 2009; Santagata et al., 2011; Star et al., 2015; Thompson et al., 2012) contain evidence suggesting that lack of structural or process implementation fidelity may have led to an absence of impacts on student outcomes. One of those (Cavalluzzo et al., 2014) reported moderate fidelity based on teacher and student surveys but commented on a lack of fidelity in its discussion, whereas another (Grigg et al., 2013) noted that although treatment teachers were nearly twice as likely to use an inquiry science teaching method than were control teachers, the quality of those instructional elements was questionable.
Next we turn to models that predict implementation fidelity by study characteristics and program features. In Table 6, we used a two-level model (fidelity measures nested within studies) to regress the fraction of positive fidelity impacts on study design characteristics. We find that using a process (vs. structural) fidelity measure is associated with a lower fidelity rating; unexpectedly, we find the use of classroom observations (vs. teacher self-reports) associated with stronger fidelity in our final model. Researcher-designed (vs. third-party designed) fidelity measures have a positive relationship when entered into the models alone but no relationship in the final model. Finally, sample size was not related to the fraction of positive fidelity impacts.
Results of Multilevel Regression Analysis for Variables Predicting Percentage of Positive Fidelity Impacts (N = 61)
Note. Table displays results from a multilevel model in which impact estimates are nested within programs. Of the 65 studies that measured fidelity, 4 measured only dosage fidelity. Standard errors are in parentheses.
p < .1. **p < .05. ***p < .001.
In Table 7, we regress our holistic measure of fidelity level on program features. We find that when entered singly, the program’s provision of professional development and curriculum materials is positively associated with implementation fidelity; the number of hours of professional development has a slight negative relationship with fidelity level. However, all but two programs provided some professional development, leading us both to be concerned about making strong inferences from this variable and also to omit this variable from the model with multiple predictors. In that model, curriculum materials remained a positive and significant predictor of fidelity level, but professional development hours did not.
Results of Ordinary Least Squares Regression Analysis for Variables Predicting Fidelity Rating (N = 61)
Note. Of the 65 studies that report fidelity, 4 did not include professional development or were missing a duration. Standard errors are in parentheses.
p < .1. **p < .05. ***p < .001.
To again complement our quantitative analysis, we examined project reports for factors linked to fidelity. Some projects measured factors thought to affect implementation fidelity and formally tested them as part of their analyses. For instance, Matsumura et al. (2010) evaluated the extent to which coach background, coach orientation toward their role, school-level professional community, teacher experience, and principal support explained teacher take-up of coaching (dosage fidelity). Wanless and colleagues (2013) conducted both qualitative and quantitative analyses and identified principal’s buy-in, coach’s attributes, and teachers’ perceptions of validation for their efforts as critical to teacher take-up and classroom implementation.
In addition to formal testing of the factors linked to implementation fidelity, other projects offered more impressionist accounts of such factors; these accounts are especially common among studies that found a lack of fidelity. Authors of one study that had low fidelity as reported on a process metric (Murray et al., 2014) commented that its measurement of fidelity may have been problematic—that the observational metric used to capture classroom processes (the Classroom Assessment Scoring System) was not sufficiently aligned to the intervention’s outcomes. Thus, fidelity itself may not have been an issue. Santagata, Givvin, and their colleagues (Givvin & Santagata, 2011; Santagata et al., 2011) discussed a wide array of reasons for lack of implementation, from inadequate principal support for the program, competing programs that absorbed teacher time, and for some teachers, insufficient content knowledge to fully understand and implement the program. H. C. Hill et al. (2018) identify a similar set of reasons for a teacher professional development program’s failure to affect practice. Santagata and colleagues (2011) also noted that teachers often came unprepared to meetings, an observation echoed by Gersten and colleagues (2010). In some cases, the difficulty teachers experienced when implementing novel instructional practices seemed at issue. For instance, Cavalluzzo and colleagues (2014) surmise that although their teachers did engage in routines around data use, the focus of the intervention, they were not able to translate what they learned from data into classroom practice. Borman et al. (2008) surmise an “implementation dip,” in which instructional quality declines as teachers encounter new curricula and instructional routines. Three other authors (Hill, Santagata, and Hurtig) also speculate that their intervention may have not been sufficiently strong to overcome obstacles to implementation.
Conclusion
The field of implementation fidelity research has come far from the days when interventions were black boxes converting inputs to outputs. Much of the promise of implementation fidelity noted by scholars—validating experimental designs, understanding mechanisms, and explaining null results—has been realized in recent studies. Four-fifths of studies reported on some measure of fidelity. Many included measures of more than one type of fidelity—structural, process, or dosage—and many used classroom observations to examine impacts on practice. Low fidelity increases the likelihood of weak student outcomes, and fidelity itself is predicted by both study design characteristics and program features. We offer several interpretations of our findings and suggestions in this conclusion.
On average, better fidelity correlated with better program outcomes, confirming an assumption made by many scholars and aligning with empirical evidence from studies in which implementation fidelity observably mediates program impact (e.g., Penuel et al., 2011; Rimm-Kaufman et al., 2014). Interesting to note, our results suggest that moderate and strong fidelity yield the same likelihood of on-average positive impacts on student outcomes, leading to the intriguing hypothesis that moderate fidelity may be enough to yield positive program outcomes. Understanding better which level of fidelity is “enough” is a key task for future researchers. Nevertheless, this evidence suggests that intervenors should continue to place bets, as they have done, on improving teacher take-up of key program practices.
Our results also imply that implementation fidelity is a partial but not complete explanation for null-result studies. Eight studies had null results but high fidelity, suggesting that the program design and contextual factors outlined by Jacob et al. (this issue, pp. 580–589) and Kim (this issue, pp. 599–607) may play a role in producing program outcomes. We also found six studies with low fidelity but with positive program impacts, suggesting that authors’ fidelity measures may not have been sensitive to key changes in classroom practice or that teachers may not have implemented the intervention “by the book” yet nevertheless saw positive results.
Implementation fidelity appeared shaped by several factors, including how scholars measured fidelity. Structural fidelity measures—often checklists but almost always surface-level indicators of implementation—tended to show stronger fidelity than did process metrics, which often required more substantial changes in classroom climate or teacher practices. Classroom observations tended to see more positive fidelity outcomes than teacher self-reports, perhaps because of issues with response bias in the latter; it is not unusual for treated teachers to report enacting fewer practices once they gain more specific information about what the survey items intend to capture (see Jacob et al., 2017). Program characteristics, including the presence of professional development and curriculum materials, also positively predicted fidelity outcomes. Against expectation, fidelity was not influenced by the size of the teacher sample, suggesting that high-quality implementation can occur at scale. We also found that the maximum hours of professional development were either negatively related (when considered alone) or unrelated (in models with multiple predictors) to fidelity. These results suggest specific pathways through which fidelity can be intentionally supported by intervenors and suggest that strong fidelity is not out of reach for large programs and/or programs with limited resources for teacher professional development.
Descriptive evidence from our reading of these studies highlights other themes, themes that align with Kennedy’s (2005) analysis of teaching and efforts to improve teaching. Support from principals and peers is critical to implementation success, and programs placed in complex environments are likely to have more difficulty seeing their key components carried out. Programs that ask teachers to complete more “difficult” tasks—for instance, introducing higher cognitive demand tasks into classrooms or using data to inform instruction—may simply be more difficult for teachers to enact and less likely to be implemented with fidelity.
In reading project reports, we noticed the systematic absence of information we argue should be collected to advance our knowledge of implementation. To start, projects that studied teacher adaptation of interventions were rare (see Durlak & DuPre, 2008). Given the long history of debates about viewing implementation from a fidelity versus adaptation perspective (O’Donnell, 2008), and the notion that adaptation is likely and even desirable in some settings, the field should do much more to both qualitatively understand adaptation and perhaps systematically test whether planned teacher adaptation can lead to improved program outcomes (see, e.g., DeBarger et al., 2017; Kim et al., 2017; McMaster et al., 2014). Second and relatedly, few reports presented teachers’ perspectives on program implementation. For instance, we rarely found teachers’ insights into typical barriers to implementation, typical difficulties working with ideas from professional development or instructional materials, and typical reasons in which implementation varied, qualitatively, from what the authors of the interventions intended. Improving program implementation at scale cannot occur without a more nuanced understanding of these issues. Finally, the reports we reviewed often contained basic information about district contexts (size, student demographics) but rarely contained insights into other factors that might condition implementation fidelity, such as the number of competing programs, alternative sources of instructional guidance, teacher capacity, and school and district organizational characteristics (see Lynch et al., 2019). Without such information, it is difficult to develop a fieldwide sense for what level of implementation is realistic in a given context and which contextual factors need to be recognized and navigated by program developers.
Finally, for the field to continue to grow, we will need more rigorous studies of implementation as well as better understanding of how program features lead to or help mitigate against implementation challenges. A good start appears in the small number of studies that predict implementation fidelity from teacher and school characteristics (e.g., Matsumura et al., 2010). Advancing the field likely means encouraging such studies, perhaps using the framework described by Durlak and DuPre (2008), to structure the systematic measurement and testing of factors related to classroom implementation.
