Abstract
Teachers are responsible for using evidence-based practices to improve students’ academic and behavioral outcomes. Although teachers have access to a variety of resources on evidence-based practices, poor implementation can adversely affect their effectiveness. However, an inadequate student response to intervention may also be the result of a mismatch between the practice and the student’s needs. As a result, it is important for teachers to determine the degree to which they implement evidence-based practices as intended to determine if an inadequate student response is due to poor implementation or inappropriate selection of intervention. The authors discuss the importance of fidelity of implementation. Methods teachers can use to measure implementation fidelity are reported. Suggested methods are discussed and examples are provided.
According to contemporary educational policy (e.g., Individuals with Disabilities Education Improvement Act, 2004; Race to the Top Act, 2011), teachers are responsible for improving the academic performance of all students while using evidence-based teaching practices (EBP). However, access to EBPs alone does not ensure that students will benefit from these practices and improve their academic, behavioral, or social outcomes. Variability across teachers in the implementation of these practices may limit or negate the potential benefits of EBPs.
Factors that can adversely affect teachers’ use of EBPs include (a) the complexity of the intervention or practice (e.g., a lengthy multicomponent writing intervention), (b) access to materials and resources necessary to implement the program, (c) differences between the way practitioners perceive the intervention to be effective and its actual effectiveness, and (d) characteristics of the person delivering the intervention, such as motivation and skill (Johnson, Mellard, Fuchs, & McKnight, 2006). Research also suggests that given the extensive demands on educators’ time and other school-based issues, many teachers require ongoing support to deliver interventions as they were intended, also known as implementing with fidelity (O’Donnell, 2008; Schulte, Easton, & Parker, 2009). Fidelity is defined as the degree to which an intervention or practice is provided to students as intended (Keller-Margulis, 2012) and is an important consideration for school professionals. In consideration of these factors, it is important for educators to determine the degree to which EBPs are implemented with fidelity and to take action to improve classroom practices when necessary.
Online resources, peer-reviewed education journals, and professional organizations make EBPs for academics and behavior accessible to educators. In this article, we describe a framework to help teachers take instructional practices to the next level, by carefully considering and evaluating the degree to which teaching practices are being implemented as intended. We describe strategies educators can use to improve the effectiveness of instructional and intervention practices.
Benefits of High Levels of Fidelity
It has been demonstrated that school reforms and practices can be ineffective because of poor implementation (Kovaleski, Gickling, & Marrow, 1999; Levin, Catlin, & Elson, 2005) and improved when delivered with high levels of fidelity (Foorman & Moats, 2004). High levels of implementation fidelity can result in greater gains for students (Noell, Gresham, & Gansle, 2002; Songer & Gotwals, 2005; Ysseldyke & Bolt, 2007) and is also an integral part of response to intervention practices (Johnson et al., 2006). Taking these factors into consideration, it is imperative that practitioners measure fidelity to maximize instructional effectiveness and determine whether teacher practices are influencing student outcomes or if changes are needed.
When fidelity is measured, practitioners can use this information to determine if a practice was ineffective because of poor implementation or a failure to match the intervention to students’ needs (Johnson et al., 2006). When measuring fidelity, practitioners can consider whether a lack of student response is due to the intervention’s not being delivered appropriately or the intervention’s being inadequate. Without determining the degree to which students are provided instruction as it was intended, it is difficult to determine the appropriateness of the current programming a student receives (Keller-Margulis, 2012).
Although parents, teachers, principals, and other stakeholders can potentially identify when teaching is effective through observations, assessment results, or even informal methods, research provides evidence to support the necessity of high fidelity when working with students who have disabilities. For example, in a program that provided early reading intervention, researchers found that providing ongoing instructional support to teachers, and the degree to which teachers implemented the reading program with fidelity, influenced the degree to which students improved their reading skills (Stein et al., 2008). Because using proper methods of instruction for research-based practices is important for helping students improve, the field of special education is now paying more attention to fidelity of implementation for interventions (Swanson, Wanzek, Haring, Ciullo, & McCulley, 2012).
Components of Fidelity
Fidelity has two components: (a) fidelity to structure and (b) fidelity to processes (Mowbray, Holter, Teague, & Bybee, 2003). Fidelity to structure often refers to adherence to a behavioral or academic intervention’s component parts. For example, a reading intervention for students with learning disabilities could consist of teacher modeling of fluent reading, followed by teacher-guided student practice of repeated readings, followed by independent practice in which students engage in repeated readings. The structure of this intervention includes each of these activities delivered in the prescribed sequence. Skipping a part of the procedure, for example, going straight from modeling to independent practice, could limit the effectiveness of this reading strategy for the students.
When providing instruction or an intervention, it is possible to follow procedures (adherence) but do so poorly (Gresham, 2009). Fidelity to processes refers to the quality of instruction (Furtak et al., 2008). In the previous example, the quality with which the teacher delivered each component of the reading fluency intervention represents fidelity to the process. For example, high quality would be indicated when the teacher models fluent reading by using vocal expression to convey meaning as well as simply reading with automaticity. Conversely, low quality would be indicated if the teacher simply read the words in the text with 100% accuracy.
Methods for Measuring Fidelity
Fortunately for educators, fidelity can be measured using direct and indirect methods (Keller-Margulis, 2012). When making decisions regarding which types of data are feasible to collect, educators should consider factors such as time and should match the collection method to the types of data they wish to evaluate. In general, it is beneficial to use multiple methods to assess fidelity. Methods practitioners should consider are observation, self-assessment, and analysis of permanent products. Each of these techniques and examples are provided.
Observations
Observation of teacher behavior is one direct method of assessing fidelity. A person who is knowledgeable in the intervention or curriculum serves as an observer to determine the degree to which teachers adhere to core procedures and elements (Crawford, Carpenter, Wilson, Schmeister, & McDonald, 2012). Prior to observing teacher behavior, checklists that contain the core components of an intervention are created. When creating checklists, it is important to identify the critical components of the intervention or instruction that are thought to promote student performance so that the observation is focused (Ruiz-Primo, 2005). This step can be completed through task analysis. Core components may include specific activities and time spent completing specific activities. In addition to considering teachers’ implementation of core components, measuring fidelity may also include a determination of whether all activities of a particular curriculum or intervention were completed within the allotted time (Durlak & DuPre, 2008; Gersten et al., 2005; Power et al., 2005) and in the proper sequence.
Example 1: Good Behavior Game
Figure 1 is an example of a fidelity checklist for the Good Behavior Game (GBG; Barrish, Saunders, & Wolf, 1969), an interdependent group contingency with a strong evidence base supporting its effectiveness at addressing challenging behavior (Embry, 2002; Tingstrom, Sterling-Turner, & Wilczynski, 2006). It is considered a classwide intervention, meaning that the intervention is used with an entire class. This checklist contains the key components of the intervention (structure), as well as a method for measuring the degree to which teachers implement certain essential practices (processes).

Good Behavior Game fidelity checklist.
According to the initial empirical implementation of the GBG (Barrish et al., 1969), the GBG has specific, sequential procedures. The features of the initial version of the GBG are presented here. Prior to beginning academic instruction, students are divided into teams, and the teacher posts a recording sheet or designates a place that is visible to the students for recording rule violations by each team. The teacher then tells the class that the game is beginning and reviews the class rules or expectations. Next, the teacher reminds students not to go over an unknown criterion of fouls so that they can win the game. The teacher then provides academic instruction, using specific error correction procedures (i.e., stating what students did wrong when they violate a class rule or expectation and stating what they should do instead) and recording fouls on the recording sheet as they occur in the classroom. At the end of class, the teacher stops instruction and tells the class that he or she will now count the number of fouls earned by each team. After counting fouls, the teacher tells the class the maximum number of fouls allowed to earn a reward (i.e., the criterion for winning). The teacher then identifies which teams met the criterion for success and provides a predetermined reward and a token for a cumulative reward that the entire class is working toward.
When completing the checklist during the lesson, the observer marks whether components were present or absent from instruction. For example, if the teacher explicitly reviewed classroom rules with students, the observer would place a check next to that component. However, not all aspects of the intervention are rated in this manner (i.e., presence or absence). Teacher identification and recording of disruptive behaviors is rated along a continuum, ranging from always performing a behavior to a behavior’s not being observed. This is an example of measuring fidelity to process. When calculating a total fidelity score, items that are evaluated in terms of their presence or absence are awarded zero points when absent and one point when present. Items that are evaluated along a continuum are given zero points when not observed, one point when sometimes observed, and two points when always observed. Total points awarded are then divided by the total number of possible points to determine the percentage of the intervention that was delivered as intended. Low overall scores and low ratings on specific items would indicate areas in need of additional support, such as professional development and coaching.
More recent modifications to the GBG have been made since its initial implementation (Barrish et al., 1969). These modifications rely largely on awarding points to teams that demonstrate rule-following behaviors rather than posting fouls for behaviors that do not conform to rules or expectations (e.g., Babyak, Luze, & Kamps, 2000). Other research has used a combination of reinforcement focused on rule-following behavior and response-cost (i.e., fouls) for misbehavior (e.g., McGoey, Schneider, Rezzetano, Prodan & Tankersley, 2010).
Example 2: individualized intervention
Individualized interventions can be complex, making it important to monitor the degree to which they are provided as intended to maximize student growth. Individualized interventions may include interventions based on functional behavioral assessment data. For interventions such as these to be effective, teachers must deliver them with precision. Figure 2 is a fidelity checklist for an individualized, function-based intervention provided to a fourth grade student with challenging behavior. This checklist includes the intervention’s key components and was used to determine the degree to which it was provided in actual practice. As part of the intervention, the teacher used a vibrating timer as a prompt to provide the student with specific praise and explicit feedback every 5 to 7 minutes. The timer was also used to prompt the teacher to actively supervise the student to make sure that he was on task. The student also had a “check-in” sheet, which was used to award the student points for demonstrating specific target behaviors. Points were awarded during a brief check-in with the student at the end of each class and were redeemed by the student for privileges. Last, the teacher was asked to monitor the student’s transition to seat work, as this was a time when many off-task and problem behaviors occurred.

Individualized intervention.
As in the previous example, some aspects of the intervention are rated in terms of the presence or absence of the instructional event actually occurring. These include the presence of the check-in sheet on the student’s desk and filling in the point sheet with the student at the end of class. Other intervention components are rated along a continuum from never to always, such as the teacher’s delivery of specific praise and explicit feedback approximately every 5 to 7 minutes. After completing the observation checklist, a fidelity score is calculated. For items that are rated in terms of their presence or absence, items that are present are given one point, and those that are absent are given zero points. For items rated along a continuum, items that were never observed are given zero points, items that were sometimes observed are given one point, items that were observed most of the time are given two points, and items that were always observed are given three points. Total points awarded are then counted and divided by the total number of possible points to obtain a percentage, which is the overall fidelity score.
Conducting Observations
Checklists can be completed by school administrators, school psychologists, and teachers who are knowledgeable in the specific intervention or practice being conducted. When measuring fidelity, administrators can schedule a series of observations and emphasize that they are being conducted to improve teacher practices and student performance (Johnson et al., 2006). It may also be beneficial to observe less experienced staff members early in the school year so that they can receive feedback that can be integrated into practice. However, observations may not always give a clear indication of the quality of instruction, because people can act differently when being observed (Sheridan, Swanger-Gagne, Welch, Kwon, & Garbacz, 2009). To account for this possibility, some researchers have recommended conducting observations at unscheduled times in addition to times that are scheduled (Keller-Margulis, 2012). For example, fidelity data could be collected during informal walks through (Sanetti & Kratochwill, 2009). It also may be beneficial to have multiple people measure fidelity to get different perspectives on teacher implementation (Keller-Margulis, 2012). In consideration that fidelity may decrease over time (Gresham, 2009), it is also important to periodically monitor teacher delivery.
Self-Assessment of Fidelity
Checklists can also be used by teachers to self-assess their practice to improve intervention fidelity (Keller-Margulis, 2012). After delivering an intervention or curriculum, teachers can complete the checklist while being self-reflective, identifying aspects of the intervention they find difficult. Teachers can then seek out additional support, such as retraining or coaching, to improve their practice. For example, a fourth grade special educator seeking to teach her students a new reading comprehension strategy for creating story summaries may attempt this strategy, conduct a self-assessment, and then ask a colleague who is proficient in this lesson to observe the following day to offer specific feedback.
Although self-assessment can be advantageous because it can be performed without the use of additional staff members, practitioners should be aware that sometimes self-reports are inaccurate (Jobe, 2003), highlighting the need to use multiple methods to measure and enhance fidelity. For example, observations can be used to confirm self-reports (McKenna, Rosenfield, & Gravois, 2009) and can be used to compare and contrast teacher and observer perspectives when providing consultation services. Videotaping intervention delivery may also be used to supplement self-assessment data (Schulte et al., 2009). For example, a person knowledgeable in the intervention or curriculum can view the videotape and complete a fidelity checklist and discuss these results with the teacher during a consultation meeting to help the teacher improve instruction.
Permanent Products
Permanent products can also be used to measure fidelity. Examples of permanent products that can be used to measure fidelity to a behavioral strategy are student self-monitoring sheets, student point sheets, charts, and tokens (Sheridan et al., 2009). Permanent products may be particularly helpful when measuring the fidelity of interventions teachers use over the course of the school day (Noell et al., 2005). When using permanent products to determine fidelity, teachers look for evidence that specific parts of the intervention or curriculum were followed. For example, teachers could look to see if all members of a student group completed all sections of their reading comprehension learning log worksheet for the collaborative strategic reading intervention (Vaughn et al., 2011). This would inform the teacher about the degree to which students had learned the process of monitoring their comprehension while reading an expository text passage about dolphins. If the teacher finds confusion about the intervention itself, or that students are struggling with a certain element of reading comprehension, future lessons could be easily taught on the basis of this information to help the students improve their skills.
Teachers and other personnel could also look to see if a student’s point sheet was completed accurately. For example, with a daily behavior report card, teachers should be reviewing expectations and providing students with an indication of progress periodically throughout the day. Figure 3 is an example of a daily behavior report card. In this example, the student is working on three expectations or target behaviors: stay safe, follow directions, and try your best. At the end of each class period, the teacher briefly provides feedback and awards points to the student on his or her performance. If ratings are circled with a vertical circle down the report card rather than individually, this might be an indication of poor fidelity. The absence of ratings on the daily report card is also an indicator of poor fidelity. The use of permanent products may not always be an appropriate method for measuring fidelity, such as in instances when a subjective measure of implementation quality is required (Sheridan et al., 2009). Another consideration is that permanent products can inaccurately reflect fidelity of implementation, such as in instances when they are completed prior to or upon completion of the intervention (Noell et al., 2005). Practitioners can address this potential concern by supplementing data from permanent products with observations.

Daily report card.
Summary
Measuring intervention fidelity, and taking steps to improve procedures of an academic or behavior strategy, can contribute to improved student outcomes. When fidelity is low, teachers can be provided additional coaching and feedback on intervention delivery. When fidelity is adequate but student performance is lagging, teachers can reconceive and adjust the interventions to more adequately meet student needs. In this article, we have described various methods for measuring fidelity that can be used by school professionals and that consider school realities such as available time. As schools strategically plan to incorporate fidelity assessment into their typical practices, they should keep in mind that using more than one method of data collection is advantageous compared with relying on a single method. In addition, collecting data strategically, such as at the beginning of the year as well as over time, is beneficial. Collecting fidelity data in this manner provides an opportunity to quickly identify teachers in need of support as well as those whose practices may drift from established procedures.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
