Abstract
A systematic analysis was conducted of measurement and reporting practices related to procedural fidelity in single-case research for the past 30 years. Previous reviews of fidelity primarily reported whether fidelity data were collected by authors; these reviews reported that collection was variable, but low across journals and over time. Results of this review indicate that fidelity data collection was variable across journals, but increasing over time. However, despite previous recommendations for doing so, authors of many studies failed to report when data were collected, for what behaviors, and for which participants. Recommendations include continued fidelity measurement, increased breadth of measurement, increased precision of measurement, and explicit reporting of fidelity data and measurement procedures by authors.
Researchers often collect data to decrease the likelihood that human error accounts for changes in the dependent variable of interest (e.g., interobserver agreement [IOA] data). High levels of agreement may increase confidence that observers are accurately measuring dependent variables of interest (Gast, 2010), and IOA data are widely reported and expected in educational and behavioral research (e.g., Snell et al., 2010). However, even when experimenters provide evidence that dependent variable data were reliably recorded, there exists the possibility of human error in the implementation of experimental procedures. There is a long-standing acknowledgment in behavioral research with human implementers that procedural infidelity is possible and perhaps likely (e.g., Baer, Wolf, & Risley, 1968; Billingsley, White, & Munson, 1980; LeLaurin & Wolery, 1992). In early intervention research, infidelity may be even more likely, with two “levels” of fidelity measurement: (a) whether researchers implement training procedures correctly and (b) whether indigenous implementers (e.g., early childhood special education teachers, parents) can (and do) implement interventions successfully after training.
Several reasons exist for measuring fidelity data, and for reporting the degree to which implementers adhere to the procedures described (Wolery, 2011). One reason is to use the data formatively: To determine the ongoing training needs of implementers or to allow researchers to provide additional support for difficult-to-implement procedural components. For example, after training early interventionists to implement a multicomponent intervention, researchers should measure their fidelity of implementation in practice. If the early interventionists implement one component of the intervention with low fidelity, “booster” training sessions can be provided that are specifically targeted toward improving performance related to the single component that is problematic.
The other reasons for measuring fidelity are summative, and are related to the arguments presented by Baer et al. (1968), Billingsley et al. (1980), and others: Confirming adherence to a described protocol increases the internal validity of the study by demonstrating that the procedures were implemented accurately and thus are the likely reason for changes in dependent variables. Although a change in behavior concurrent with a change in experimental condition might convince a reader that the independent variable has been applied, an alternate explanation still exists: The implementer may have made unintended changes to the environment, outside the procedural plan (Gresham, Gansle, & Noelle, 1993).
Fidelity data may also explain the effects when inconsistent results exist—it is possible that variations in fidelity might be related to outcome variability, answering the question of whom and under what conditions a specific intervention is likely to be effective and increasing the specificity of conclusions that can be drawn regarding generalization (Strain et al., 1992). In a single-case study in which two participants made substantial gains, but no behavior change was evident for a third participant, one potential explanation is that the interventions were differentially implemented for that participant. For example, an early childhood special education teacher might be asked to implement an intervention that required her to respond to a participant’s nonverbal requests. The difficulty of responding to nonverbal requests might vary by participant, as some children might engage in more obvious requesting behavior (e.g., pointing while pulling the teacher toward an item), while another child might engage in more subtle nonverbal behavior (e.g., pointing without gaining teacher attention). Thus, the teacher might differentially reinforce requests across students, resulting in potential intervention failure for the student with subtle requests. This information could be used both formatively (e.g., by alerting the teacher to missed opportunities) and summatively (e.g., to conclude that the intervention may be most effective for children with pronounced rather than subtle nonverbal requesting behaviors).
Procedural fidelity data also can answer the question of whether indigenous implementers—the people who are typically present in a child’s usual environments, such as parents and teachers—can implement a set of procedures accurately (Halle, 1998; C. A. Peterson & McConnell, 1993). If the answer to this question is “no,” researchers can provide more intensive implementation training or change the intervention so that implementers can conduct the procedures with fidelity. For instance, component analysis research for an intervention might allow researchers to delineate practitioner steps that are necessarily completed with high fidelity and those that may be implemented with lower fidelity. This may be important, as some interventions have been applied by indigenous implementers with very high fidelity (e.g., response prompting procedures in small group direct instruction; for a review, see Ledford, Lane, Elam, & Wolery, 2012), while others have been implemented with variable or low fidelity (e.g., Wood, Umbreit, Liaupsin, & Gresham, 2007). In addition, some components of complex interventions may be entirely excluded, resulting in a more parsimonious practice (Cook, 1985). For example, an early interventionist might train the mother of a young child with disabilities to perform a two-part intervention: (a) ignoring inappropriate behavior designed to get her attention (e.g., hitting) and (b) reinforcing appropriate attempts to get attention (e.g., touching her arm). But, many mothers might find ignoring inappropriate behaviors difficult or unacceptable. A component analysis of this intervention might show that mothers more accurately implement an intervention with only the (b) component, and that child outcomes were not affected by the removal of the (a) component. This research, while rare, may be increasingly necessary as researchers try to support indigenous implementers in typical environments.
One conceptualization of procedural fidelity common in behavioral and educational literature involves categorizing the types of behaviors to be measured as independent or control variables. Independent variables are those that are experimentally manipulated and are considered essential components of the intervention being evaluated. Control variables are those that should remain the same (constant) across conditions. Although measurement of only independent variables is common, measurement of control variables allows the researcher to assert the independent variable occurred only during the intervention condition and that all other important variables remained the same. For example, when assessing the effectiveness of a differential reinforcement procedure to increase correct responding, measured independent variables should include the consequences for correct and incorrect responding. In this example, an important control variable might be the number of opportunities to respond across baseline and intervention conditions.
Five common variations occur when measuring procedural fidelity in single-case experiments:
Measurement of independent and control variables, across conditions (baseline and intervention);
Measurement of only independent variables, across conditions;
Measurement of only control variables, across conditions;
Measurement of independent and control variables, in the intervention condition only;
Measurement of only independent variables, in the intervention condition only.
Conclusions that can be drawn from each variation are shown in Table 1. When measuring variables across conditions (numbers 1-3 above), statements can be made about the extent to which a particular variable was implemented during a specific condition. These data show evidence of adherence—a variable was present or not present to the extent planned during a single condition. In addition, statements can be made about changes across conditions. These data show evidence of differentiation—the degree to which a procedural step (implementer behavior) changed or did not change across two experimental conditions. When occurrences of variables are assessed during only the intervention condition, it is possible only to assess adherence; it is not possible to also assess the degree to which there was a change in the levels of independent and control variables across conditions (differentiation). It is possible, for example, that an independent variable was implemented with equal fidelity in baseline and intervention, and that unmeasured variables are responsible for behavior change.
Common Variations in Fidelity Measurement
Note. IV = independent variables; CV = Control variables; BL = Baseline; Int = Intervention. Similar variations are possible for comparison designs, with measurement contexts of Independent Variable 1 and Independent Variable 2, rather than Baseline and Intervention.
Assuming acceptable adherence to described procedures and expected behavior change.
When defining the measurement of fidelity, an important distinction should be made between the terms procedural and treatment when each is combined with the terms integrity, fidelity, or reliability. Although not always used precisely, these two groups of terms have different meanings. Treatment fidelity refers to the measurement of the independent variable, or (more rarely) independent and control variables, during only the treatment condition. Thus, treatment fidelity allows only for analysis of adherence. Procedural fidelity refers to the measurement of independent and control variables during both baseline and treatment conditions (or across different treatment conditions), and allows for analysis of both adherence and differentiation for both types of variables. No common terminology has been used to describe (a) collection of data on the independent variable only, across conditions, and (b) collection of data on control variables only, across conditions.
Since the seminal article by Billingsley and colleagues (1980), numerous reviews have been published regarding the use and reporting of fidelity data in specific journals or for specific populations or interventions (e.g., Dane & Schneider, 1998; Dusenbury, Brannigan, Falco, & Hansen, 2003; Gresham et al., 1993; McIntyre, Gresham, DiGennaro, & Reed, 2007; Moncher & Prinz, 1991; L. Peterson, Horner, & Wonderlich, 1982). Few studies have analyzed what factors may impact fidelity (e.g., C. A. Peterson & McConnell, 1996), although presence of acceptable fidelity data is often described as one indication of the quality and generality of research (e.g., Ledford et al., 2012; Roberts & Kaiser, 2011). Recently, in a special issue devoted to fidelity measurement, articles in School Psychology Review pointed to the importance of fidelity measurement in increasing the applicability of evidence-based practices by indigenous implementers and in determining what “dosage” is necessary for an intervention to be effective (Greenwood, 2009).
Several reviews exist regarding the extent to which fidelity data were reported in articles published in the Journal of Applied Behavior Analysis (JABA). The first review (L. Peterson et al., 1982) found that 16% of studies in JABA reported fidelity data between 1968 and 1980, and that the percentage of studies reporting fidelity did not systematically increase over time. A second review (Gresham et al., 1993) showed again that no increase was evident, with 16% of studies with children reporting fidelity data between 1980 and 1990. The most recent review of articles in JABA (1991-2005; McIntyre et al., 2007) found a substantial increase, with 30% of studies reporting fidelity data and an additional 8% reporting fidelity was measured, with no data being reported. However, because each of these reviews targeted different subsets of articles, with narrower inclusion criteria for the latter studies, direct comparisons are not possible.
Other reviews have been published analyzing fidelity data for specific journals related to special education or for participants with specific disabilities (e.g., Armstrong, Ehrhardt, Cool, & Poling, 1997; Griffith, Hurley, & Hagaman, 2009; Snell et al., 2010; Wheeler, Baggett, Fox, & Blevins, 2006), as well as in areas outside special education (e.g., school psychology; Sanetti, Gritter, & Dobey, 2011). These reviews have included relatively few studies reporting fidelity (fewer than 50) and report variable results. For a 5-year period (1991-1995) in Journal of Developmental and Physical Disabilities, 23% of studies reported fidelity data (Armstrong et al., 1997). In studies targeting communication skills for children with severe disabilities, only 32% reported fidelity data (Snell et al., 2010). Fewer studies (18%) reported fidelity data in intervention studies for children with autism (Wheeler et al., 2006), while more (52%) reported fidelity in literacy interventions for children with behavior disorders (Griffith et al., 2009). The variability across reviews may be related to small sample sizes (e.g., a single study had a relatively large impact), time ranges reviewed, and journals included. Although no review specific to early intervention has been reported with an emphasis on fidelity, a review published to determine the strength of evidence in early childhood special education (Odom & Strain, 2002) found that about half of the studies in the area measured and reported fidelity data. In addition to these data-based reviews, discussion articles have been published consistently over the last 30 years, calling for an increase in fidelity measurement and suggesting that this measurement should be done following specific guidelines (e.g., should not rely on self-report). The consensus from these discussion and review articles is that the historical emphasis and recent interest in fidelity have not resulted in consistent increases in the measurement or reporting of fidelity in any discipline, though a comprehensive review has not been conducted.
This review addresses the extent to which recommendations made in previous articles have been implemented in studies published in 14 journals across 31 years. In previous articles, recommendations for fidelity measurement and reporting include (a) defining critical features for which evidence of implementation is necessary and reporting separately for each implementer (Gresham et al., 1993; Moncher & Prinz, 1991; Perepletchikova, 2011; Schlosser, 2002), (b) reporting the measurement procedures used (e.g., direct counts, checklists; Moncher & Prinz; Perepletchikova) and using direct observation rather than self-report (Schlosser), (c) measuring and reporting fidelity for each participant (Moncher & Prinz), (d) measuring fidelity across conditions (Wolery, 2011), and (e) measuring and reporting data for independent and control variables (Wolery).
The primary research question was as follows:
For studies that measured fidelity, we asked to what extent previous recommendations have been followed: What percentage of studies (a) named procedural variables for which data were measured and reported? (b) used direct counts, checklists, or self-reports? (c) measured and reported fidelity for each participant? (d) measured and reported fidelity across conditions? and (e) measured independent variables, control variables, or both?
Method
Journal and Article Identification and Selection
Fourteen journals were identified because of (a) perceived importance to the field of special education (e.g., high readership) or (b) publication of a large number of studies using single-case design. Journals were selected that included specific discipline subgroups (e.g., early childhood special education, developmental disabilities, high-incidence disabilities). Journals selected for analysis were as follows: The American Journal on Intellectual and Developmental Disabilities (AJIDD), Behavioral Disorders (BD), Education and Training in Autism and Developmental Disabilities (ETDD), Education and Treatment of Children (ETC), Exceptionality (EXC), Exceptional Children (EC), Focus on Autism and Developmental Disabilities (FOCUS), JABA, Journal of Autism and Developmental Disorders (JADD), Research and Practice for Persons With Severe Disabilities (RPPSD), Journal of Behavioral Education (JBE), Journal of Early Intervention (JEI), Journal of Special Education (JSE), and Topics in Early Childhood Special Education (TESCE).
Because coding all single-case research across a wide range of journals would have been prohibitively resource-intensive, we randomly selected articles from each journal during each year from 1980 to 2010 to attain a representative sample of studies from which to make generalizations. For 13 journals, 2 issues per year were randomly selected using a random number generator, with a minimum value of 1 and a maximum value equal to the number of issues in that year (e.g., in a year with 4 issues, two numbers between 1 and 4 were selected) from 1980 to 2010. For the 14th journal (JABA), only 1 issue per year was randomly selected. This was done because the average number of single-case research articles per issue in JABA was at least 4 times the average number from any other journal; selecting fewer issues prevented overrepresentation of a single journal. In some cases, 2 issues were not selected because (a) publication started after 1980 or (b) fewer than 2 issues were published in a given year. Articles were previously identified for 1 issue per year for 25 years (1983-2007) for 8 of the 14 journals, as reported by Hammond and Gast (2010; see the appendix). For those 8 journals, we used the already-identified articles with single-case studies for 25 years and independently identified remaining articles. For the other 6 journals, we identified all articles across all years. Across 31 years and 14 journals, a total of 743 issues were randomly selected for review.
Each article from every selected issue was reviewed to determine whether the article contained a study conducted using single-case research methods. Studies were reviewed by doctoral students in an introduction to single-case methodology course, with limited or no incoming experience with single-case research and procedural fidelity. For each identified article using a single-case design, data collectors coded (a) journal name, (b) year, (c) the number of studies in the article, and (d) whether fidelity data were collected. Multiple studies were coded if (a) the author reported studies separately (e.g., Study 1 and Study 2) or (b) two or more designs were used to answer separate research questions (e.g., a functional analysis was conducted using an alternating treatments design, and then an intervention was implemented using an A-B-A-B design). Some articles included multiple levels of fidelity measurement, in a cascading logic model. When using this model, one implementer (often a researcher) trains a second implementer (often an indigenous adult, such as a parent or a teacher) to implement an independent variable. For these studies, fidelity data can be collected for both the first and second implementer; data from each implementer were coded separately, if such data existed. Only formative evaluation of fidelity was coded; articles including pre-intervention training with fidelity assessment were only included if fidelity data were also collected in the context of ongoing study activities. If fidelity measurement was reported by authors, articles were selected for further coding.
Five additional questions were asked for studies measuring fidelity data. The first was whether authors reported the specific behaviors assessed and whether data were reported by variable. Thus, authors might name implementer behaviors by step, but report fidelity as a single percentage of correct implementation; they might also report implementer behaviors and report fidelity results for each behavior separately.
The second question was related to the type of measurement used (checklist, direct counts, self-report). A checklist was defined as a measurement not including specific counts of implementer behavior that may have occurred multiple times per session (e.g., praising correct responses). Direct count was coded if specific counts were recorded using direct observation (e.g., number of praise statements), if behaviors were “checked” on a trial-by-trial basis, or if all behaviors were planned to occur only once per session, and this occurrence was recorded. Direct counts or checklists completed by the implementer were always coded only as “self-report.”
The third question was whether authors collected data across participants and conditions, and whether reporting was done by participant and condition. Studies were coded as “yes” for measuring across and reporting by condition if data collection or reporting was done for at least two conditions that were compared with determine whether a functional relation existed (e.g., studies did not need to collect data in generalization and maintenance conditions if the functional relation was assessed in the context of baseline/intervention sessions; no reporting was necessary for baseline conditions for comparison designs unless the baseline condition was related to determination of a functional relation).
The fourth question was related to variable type (independent, control, or both). Characterization of variable type was done by the first author, using the description of conditions provided in the article. Using these variable type codes, and the code for measurement across conditions, the percentage of studies using analyses of adherence, differentiation, or both was determined for each variable type.
Results
A total of 481 articles with 545 different studies were identified as those that measured fidelity in the context of single-case designs. In addition, 16 studies reported fidelity for multiple levels, using a cascading logic model; each of these levels was coded as a separate study. Thus, the total number of studies related to fidelity measurement is 561. Of these studies, 496 both collected and reported fidelity data (the remaining 65 studies reported that fidelity data were measured, but did not quantify results). Thus, the total number of comparisons related to measurement of fidelity (n = 561) is different from the total number of comparisons related to reporting of fidelity data (n = 496).
IOA
Trained doctoral students (including the first author) collected initial data from randomly selected issues of 14 journals. This coding was done to determine whether (a) each study in the issue used a single-case research design and, if yes (b) whether fidelity data were collected or reported in the study. Agreement between two coders was assessed separately for each of these variables. Some studies were previously identified as using a single-case research design (see the appendix) by Hammond and Gast (2010). For these studies, agreement for identification was not assessed; agreement for the presence of fidelity measurement was. Agreement was calculated using the number of issues with disagreements divided by the number of total issues (no issue ever contained more than a single disagreement). Using this formula, the identification agreement percentage was 97.8%. The number of disagreements for each doctoral student ranged from 0 to 2, and disagreements by journal ranged from 0 to 3. For 265 (23%) of 1,178 identified articles (containing 1,215 studies), two doctoral students coded whether fidelity data were collected. A disagreement in coding occurred for 9 studies (3%), and one of the authors resolved each disagreement by coding each study a third time. The first author coded all variables related to fidelity measurement and reporting for all identified studies. Four trained graduate students served as IOA coders for 120 of these studies (21.4% of studies, range across journals = 17%-33%). Agreement was assessed separately for each code (variable). The average agreement was 95%, across coders (range = 94%-96%) and codes (92%-100%; see Table 2).
Interobserver Agreement Data by Collector and Code
Fidelity Measurement
The mean percentage of studies including fidelity measurement, of all studies identified using single-case designs, was 44.9% (545/1215). The number of studies measuring fidelity was calculated for each year (1980-2010) to determine whether there were differences in fidelity measurement practices over time. Because there were differences in the number of published studies using a single-case design across years, these counts were converted to percentages and are shown in Figure 1. During the first 8 years, few studies reported fidelity (n = 26, 11% of total single-case published studies). From 1989 to 1993, there was an accelerating trend, which was followed by variable data with no accelerating or decelerating trend from 1994 to 2004. In the most recent 5 years (2005-2010), another accelerating trend occurred. The mean data collection for the earliest measured 5 years (1980-1984) was 12%, the mean for the middle 5 years (1993-1997) was 46%, and the mean of the most recent 5 years (2006-2010) was 68%. Overall, there is a variable but increasing trend in the percentage of single-case research articles reporting fidelity data measurement.

Percentage of Studies With Fidelity Data Collection by Year
Two different reporting trends are apparent across journals, and are depicted in Figure 2 as a percentage of articles reporting fidelity in 10-year increments. Some journals (AJIDD, EC, EXC, FOCUS, JEI, JSE, and TESCE) were excluded from this analysis because each had fewer than 10 published studies in 2 or more of the 10-year periods. Of the included journals, all had overall increases in fidelity measurement between the first and second time periods (1980s-1990s). Four (JBE, ETDD, BD, and RPPSD) had additional increases between the second and third time periods, while three (ETC, JADD, and JABA) had slight decreases or similar percentages between the second and third time periods (1990s-2000s). As shown in Table 3, considerable variability existed across journals for both the number of single-subject studies published (range = 15-350) and the percentage of those studies including fidelity measurement (range = 16.7%-90%) across journals. Although studies related to early intervention may have been published in many of the included journals, most of these studies were likely published in JEI or TECSE. It is important to note that both these journals included a higher percentage of articles measuring and reporting fidelity data (these two journals were two of the top three fidelity-reporting journals; JEI = 86%, TECSE = 76%).

Percentage of Fidelity Measurement Over Time for Journals With 10 or More Studies Published With Single-case Designs for at Least Two of Three Time Periods Shown
Fidelity Reporting by Journal
As an estimate of relative impact, H-factor values were reported for each journal, with higher numbers indicating greater impact. Although a presumption might exist that impact correlates with rigor (e.g., high-impact journals might require measurement and reporting of fidelity), three of the four highest-rated journals (JABA, AJIDD, and JADD) were three of the four lowest-reporting journals related to fidelity. Similarly, three of the four lowest-rated journals (by H-factor; ETC, EXC, FOCUS, JBE) had relatively high levels of fidelity measurement (above the mean of all journals). It may be that the four lowest-reporting journals (JABA, AJIDD, JADD, and RPPSD) include many articles from researchers whose primary affiliations are outside of special education (e.g., in psychology or medicine). In special education, primary funding sources (e.g., Institute of Education Sciences) have placed an emphasis on fidelity measurement; it is possible that the same expectations are not present in other fields.
Characteristics of Fidelity Measurement and Reporting
Although previous reviews of fidelity measurement have primarily reported only whether fidelity data were collected and reported, recommendations exist about how and when fidelity data should be measured and reported. Specifically, these include identifying the measured behaviors, using direct counts, measuring for all participants, measuring in both baseline and intervention conditions, and measuring for both control and independent variable behaviors. In an attempt to determine the extent to which these recommendations are followed, analysis of measurement characteristics was done (n = 561). Even when authors collect fidelity data with appropriate breadth (e.g., across participants and conditions) and with appropriate measures, they may not report sufficient information to readers. Thus, analysis of reporting practices was also done for studies that reported quantitative fidelity data (n = 496). To analyze changes in specific practices over time, 5-year increments were used (with the exception of 1980-1985, a 6-year period). Differences across time are shown in Table 4.
Variations in Measurement and Reporting Across Time
Naming specific behaviors and reporting separately for each
More than half of the studies included in the review (n = 323/561; 58%) named each behavior for which fidelity data were collected. For more than one third of studies (n = 195/561; 35%), behaviors were not named, and it was not possible to determine for what behaviors data were collected. Of 323 studies naming specific behaviors measured, only 165 (33%; 165/491) reported the fidelity data for each variable separately (although 49 named specific variables for measurement and reported 100% fidelity, making report by variable unnecessary). Thus, 43% of studies (n = 214/496) either named specific behaviors and reported 100% fidelity, or named specific behaviors and reported results separately for each behavior. Often this report included a statement that many steps were completed with 100% fidelity, with exceptions noted.
Method of measurement
For behaviors occurring multiple times per session, it is possible to use several methods to measure adherence to procedures; the three common methods are direct counts, checklists, and self-report. The most common of these measurement methods was direct counts (n = 222; 40%), which includes specific counts of behaviors, or measurement of behaviors on a trial-by-trial or interval-by-interval basis. Checklists were also common (n = 150; 27%), and self-reports were less common (n =15; 3%). Five studies reported using multiple methods of measurement (checklist and direct observation). For a large portion of the studies (n = 179; 32%), the method of measurement was not explicitly named. The use of direct counts when measuring fidelity decreased from an initial high of above 50% (1980-1990), but has remained stable for the past 20 years at around 35% to 40% of studies. The use of checklists increased across time, with the last two 5-year time periods (2001-2010) having a percentage of studies using checklists above the overall mean (Table 4).
Measurement and reporting for each participant
Explicit measurement for each participant was done in 311 studies, including 54 studies with only one participant. Almost half of the studies with more than one participant (49%) either explicitly reported failure to collect data for one or more participants or failed to report whether data were collected for each. Of the studies that collected data for all participants, 238 reported fidelity results, with 214 of those studies reporting fidelity results separately for each participant. Thus, 43% of studies reporting fidelity data reported by participant or for the only participant. There were no consistent differences across time in explicit collection of data for each participant (range = 50%-68% across 5-year periods; see Table 4).
Measurement and reporting across conditions
Fidelity across conditions was measured for 307 studies (55%). In some studies, fidelity was measured during only one condition (n = 96; 17%). It was not possible to determine which conditions data were collected for in 158 studies (28%). Of the 307 studies reporting explicit collection across conditions, 3 did not report results, 43 reported 100% fidelity for each condition, and 80 reported overall fidelity. Eleven studies reported fidelity for pre-intervention training in a cascading logic model (in only one condition). The remaining studies (n = 170; 34%) reported results separately for each condition, a practice consistent with previous recommendations. There were no consistent differences across time in explicit collection of data for each condition (range = 48%-58% across 5-year periods; see Table 4).
Type of behavior measured
“Type of behavior” referred to whether the measured behaviors were independent variables, control variables, or both. The most common code for the type of behavior measured was Can’t tell (n = 195; 35%). Typically, authors of these studies reported fidelity data were collected, but reported no information regarding specific behaviors measured and provided no explicit information about whether these were independent, control variables, or both. Almost as many studies (n = 194; 35%) reported measuring independent variables only, or named these variables. Fewer studies measured independent and control variables (n = 156; 28%) or control variables only (n = 16; 3%). There was a considerable accelerating trend in the number of studies not reporting information about the type of behaviors measured, with an increase during each 5-year block from 9% during 1980-1985 to 43% from 2005 to 2010 (see Table 4).
Use of adherence and differentiation comparisons
Percentage of studies using common variations in fidelity data collection (whether independent or control variables were measured and whether those occurred in different conditions; see Table 1), along with additional variations (possible because of inadequate reporting), are shown in Table 5 for noncomparison and comparison designs. Noncomparison designs evaluated an independent variable against a baseline condition, and comparison designs evaluated the relative efficacy of two or more independent variables. About half of all studies (49%) failed to report the type of behaviors for which fidelity data were collected, conditions during which measurement was conducted, or both. In these studies, it is not possible to tell whether adherence or differentiation was assessed, or for what variables these analyses were possible. Data were measured in both conditions (to allow for assessment of adherence and differentiation) for both independent and control variables in 23% of the noncomparison studies and 20% of the comparison studies. Data were reported for adherence and differentiation of the independent variable only in an additional 16% of the noncomparison studies and 19% of the comparison studies. Many studies (n = 96) collected data during only one condition, limiting analysis to adherence during that condition.
Measurement Variations in Comparison and Noncomparison Studies
Note. NR = not reported.
Discussion
The results of this review suggest that there are some changes in fidelity measurement over time that correspond with previously recommended practices. Data for journals specific to early intervention were particularly positive, with many studies in those journals (JEI, TECSE) collecting and reporting fidelity data. There are also considerable differences between recommendations and common practices for some aspects of measurement and reporting. Before considering implications of the results of this review, some limitations should be noted. All relevant disciplines may not have been included in the review (e.g., no journal specific to school psychology was reviewed). Also, journals were selected based on readership and perceived importance, and the quality of studies in these journals may be higher in these journals than in those not chosen. Thus, estimates of fidelity measurement may be inflated. Also, estimates of fidelity for specific journals may not be representative, as relatively small numbers of studies were published in a few included journals. Thus, conclusions made are specific to journals included in the review, and conclusions regarding measurement and reporting practices for specific journals should be made with caution. In addition, it should be noted that collecting data broadly and precisely is not necessarily indicative that data were collected well—the possibility exists that data were not collected consistently over time or that authors did not measure for all relevant control and independent variables. The degree to which authors measured consistently and for important variables was beyond the scope of this review. Finally, the fidelity measurement in these studies focused on whether data were collected on procedural variables. Suggestions have been made that other important considerations exist, such as the quality of intervention and psychometric properties of assessment (Gresham, 2009). These aspects of fidelity were not assessed as part of this review.
The primary question for previous reviews was “To what extent are fidelity data reported”? Contrary to some reviews (Gresham et al., 1993; L. Peterson et al., 1982), and in agreement with others (Griffith et al., 2009; Sanetti & Kratochwill, 2009), this review shows that fidelity measurement, while variable by year, increased over the 30-year measurement period, although this finding was not consistent for all journals (see Figure 2). Variability by journal may explain differential results in previous reviews. In this review, 45% of articles using single-case design reported measuring fidelity. No estimate of fidelity measurement was found in previous reviews that approximated the recent fidelity measurement by studies reviewed here (2006-2010; 68%), which may be the result of recent increases in fidelity measurement in the field. Journals specific to early intervention reported and measured fidelity at higher-than-average rates; we did not evaluate potential reasons for this difference. Both JEI and TECSE have published numerous articles related to fidelity for at least the past 20 years (e.g., Halle, 1998; LeLaurin & Wolery, 1992; C. A. Peterson & McConnell, 1993; 1996; Strain et al., 1992; Wolery, 2011); this emphasis may have led to increased interest in and use of fidelity measurements in these journals.
It is commonly recommended that studies include fidelity measurement, and this review shows that a higher percentage of published studies included this assessment over time. Recently, other common recommendations have been given for variations related to fidelity measurement, although no other review has evaluated the extent to which these recommendations are followed. They include (a) identifying measured behaviors and reporting for each, (b) using direct counts, (c) measuring and reporting fidelity for each participant, (d) measuring and reporting fidelity across conditions, and (e) measuring control and independent variables. Based on these recommendations, four additional conclusions can be drawn regarding specific fidelity assessment and reporting practices.
The first conclusion focuses on the recommendation that authors name and report by specific behavioral variables measured, and two findings are apparent: (a) More than half (58%) of the studies that measured fidelity named specific variables (e.g., specific teacher behaviors) for which data were collected, but (b) only about 43% reported data by behavior. This trend of reporting data by behavior is decreasing over time, despite recent recommendations that this practice should be used more frequently (Gresham et al., 1993; Moncher & Prinz, 1991; Perepletchikova, 2011; Schlosser, 2002). These data are essential for knowing which components of an intervention are likely to be implemented well, and which may require additional training or support. This knowledge can help researchers to design addition training to improve fidelity, or to design interventions teachers can implement well with limited training (e.g., to simplify their interventions).
The second conclusion is related to the recommendation to increase precision of measurement by using direct counts, rather than using self-report or checklist measurement. This review shows that the percentage of studies using checklists consistently increased across time, while the percentage of studies using direct observation and recording consistently decreased. Many of these checklists were completed for behaviors that were not explicitly named; this limits not only the conclusions authors can make, but the confidence readers can have that important variables were measured precisely across conditions. Lack of precise measurement, in turn, limits confidence that procedures were implemented consistently.
The third conclusion that can be drawn is that measurement and reporting for each participant were variable, but not increasing over time, with a mean of 55% of studies measuring for each participant, including studies with a single participant. Even fewer studies (41%) reported results for each participant separately, including studies with 100% fidelity and explicit collection for each participant. Given that much data in single-case research are reported by participant (e.g., dependent variable data), reporting fidelity data separately for each participant would not necessarily involve substantial publication space, as these data could be reported alongside other data (e.g., Wood et al., 2007) or in separate tables (Logan et al., 1998). Thus, despite recent calls to measure and report for each participant, adequate theoretical justification for doing so, and few practical reasons for not doing so, reporting practices are relatively unchanged across time, with many authors failing to report results for each participant.
The fourth conclusion is related to the final two recommendations (measurement and recording for both independent and control variables, across conditions). Only 22% of studies measured both types of variables, across comparison conditions (e.g., baseline and intervention, or Independent Variables 1 and 2). Most studies (51%) failed to report variable type (21%), failed to report conditions during which behaviors were measured (16%), or reported neither (14%). Thus, readers are often unable to determine what comparisons were made; and when comparisons made are transparent, they often include adherence to the intervention condition only. This limited measurement restricts conclusions that can be made, and the failure to explicitly report behaviors collected for and conditions during which fidelity was assessed limits confidence the readers can have that procedures were implemented as described, and that independent variables were differentially implemented across conditions.
Taken together, these four findings suggest that, although fidelity is being measured in a higher percentage of single-case studies (Figure 1), investigators are not measuring more broadly (e.g., across participants and conditions; and for independent and control variables), and may be measuring less precisely (e.g., with checklists rather than direct observational measurement and recording). Thus, authors of studies using single-case design do not seem to be heeding the recent published recommendations, with the relative frequency of some recommended practices (e.g., using direct counts) actually decreasing over time. Determining causes of narrow and imprecise measurement is beyond the scope of this review, but increased precision in fidelity assessment (e.g., using direct counts across conditions) may be avoided by some authors because it could result in the need for more resources (e.g., observer time), although the extent to which this is true has not been evaluated experimentally. It may also be that the relative popularity of fidelity assessment—for example, increased calls for fidelity measurement by funding agencies—has resulted in increased measurement, but has not caused a concurrent increase in awareness by investigators that broad and precise measurement and explicit reporting are necessary practices for making many conclusions related to fidelity and outcomes.
Suggestions for Future Research
Although there has been some consensus about recommendations regarding fidelity measurement and reporting, and although these recommendations are logically and theoretically sound, little research on fidelity measurement and reporting has been done. Specifically, two types of studies are recommended: (a) studies related to fidelity of specific interventions and (b) studies related to measurement and reporting of fidelity. Studies related to specific interventions need to include those that systematically manipulate the levels of fidelity to determine the degree to which variations in fidelity result in variations in outcome, and could also include component analyses of complex interventions. Some of these studies have been conducted (e.g., constant time delay fidelity analyses; Holcombe, Wolery, & Snyder, 1994; Wilbers, 1989), but no systematic fidelity manipulations have been conducted for the vast majority of procedures considered to be evidence based, despite the practical implications these studies may have. Thus, research is needed to determine which interventions will work despite fidelity failures that may be likely in complex early intervention settings with indigenous implementers. For example, research has shown that when reinforcement is provided with high fidelity (e.g., for every desired response), desired responses quickly increase. However, additional research has shown that, following high-fidelity reinforcement, low-fidelity reinforcement is effective in maintaining responding (e.g., use of VR-2 reinforcement schedules; Cooper, Heron, & Heward, 2007). Future research might focus on which components of an intervention package can be removed without decrements in child performance (e.g., in an A-BC-B-BC design), whether systematic errors in procedural fidelity correspond to changes in child behavior (e.g., in an adapted alternating treatments design, with one set of behaviors assigned to a “high-fidelity” treatment, and another set of behaviors assigned to the same treatment, implemented with low fidelity).
In addition to intervention-specific fidelity studies, those that compare variations of fidelity measurement and reporting are needed. Specifically, those that measure the accuracy, reliability, formative utility, and cost of using direct counts to checklist or self-report assessments are needed. For example, it may be that measuring more precisely (e.g., with direct counts) but less often (e.g., 20% rather than 40% of sessions across conditions) may result in data that are equally cost-efficient and more accurate when compared with more frequent checklist assessments. In addition to experimental manipulations, an increase in explicit reporting of data for each condition and participant would allow for reviews that determine the extent to which fidelity varies as a function of these variables, which would inform the continued use of these practices. Thus, additional research and more explicit reporting are both necessary to determine to what extent current recommendations are supported empirically.
Conclusion
General and specific recommendations have been made regarding fidelity: It should be (a) measured, (b) measured broadly (across variables, conditions, participants, and levels of implementation), and (c) measured precisely (with counts derived from direct observation). In addition, recommendations have been made that authors should be explicit in their reporting of fidelity (e.g., naming variables, conditions, and participants for which data were collected). In general, this review suggests that authors of studies using single-case research are heeding the first recommendation, with an increase in fidelity measurement over time. However, the number of studies that measured fidelity broadly and precisely has not consistently increased over time, and, in some cases (e.g., use of checklists), has actually decreased over time. In the absence of empirical evidence that broad and precise measurement is not necessary (e.g., studies showing checklist measurement is equally as reliable, accurate, and useful as direct observation and recording), investigators should measure the fidelity of both independent and control variables, in each condition and for each participant. Likewise, they should measure using direct observation and recording, and report completely procedures and characteristics of fidelity data collection in their publications. Although measurement and reporting were generally more prevalent in early intervention journals, need and room for improvement still exist, particularly for increasing breadth and precision of measurement of fidelity in research.
Footnotes
Appendix
| Included journals | Initial year | Previous titles (1980-2010) | Previously identified a |
|---|---|---|---|
| The American Journal on Intellectual Disabilities | 1980 |
American Journal of Mental Retardation
American Journal of Mental Deficiency |
No |
| Behavioral Disorders | 1980 | None | Yes |
| Education and Training in Autism and Developmental Disabilities | 1980 |
Education and Training of the Mentally Retarded
Education and Training in Mental Retardation Education and Training in Mental Retardation and Developmental Disabilities Education and Training in Developmental Disabilities |
Yes |
| Education and Treatment of Children | 1980 | None | No |
| Exceptionality | 1990 | None | No |
| Exceptional Children | 1980 | None | No |
| Focus on Autism and Developmental Disabilities | 1987 | Focus on Autistic Behaviors | Yes |
| Journal of Applied Behavior Analysis | 1980 | None | Yes |
| Journal of Autism and Developmental Disorders | 1980 | None | Yes |
| Research and Practice for Persons With Severe Disabilities | 1980 | The Journal of the Association for Persons with Severe Handicaps | Yes |
| Journal of Behavioral Education | 1991 | None | No |
| Journal of Early Intervention | 1981 | Journal of the Division for Early Childhood | No |
| Journal of Special Education | 1980 | None | Yes |
| Topics in Early Childhood Special Education | 1981 | None | Yes |
For one issue of each journal, articles using single-case designs were identified by Hammond and Gast (2010). Additional issues and all issues 1980-1983 and 2008-2010 were identified as described in this manuscript.
Authors’ Note:
Jennifer R. Ledford, Department of Special Education, Vanderbilt University; Mark Wolery, Department of Special Education, Vanderbilt University. The authors would like to thank David Gast and Dianna Hammond Baekey for their assistance in locating articles and would like to thank Laura Steacy, Mackenzie Savaiano, Sarah Ivy, Hattie Gore, Debra McKeown, Lydia Bentley, Chagit Edery, Dana Kan, and Lindsay Wilson for their assistance in identifying and coding articles for this manuscript. Preparation of this article was supported in part by the U.S. Department of Education, Office of Special Education and Rehabilitative Services, Leadership Program, Grant H325D070075. However, the opinions expressed do not necessarily reflect the policy of the U.S. Department of Education and no official endorsement should be inferred.
