Abstract
Replication studies in special education are necessary to strengthen the foundation upon which instruction and intervention for students with disabilities are built. J. Jenkins et al. (2017) found intermittent reading fluency progress monitoring schedules did not delay decision-making and were similar in decision-making accuracy to the traditional weekly progress monitoring schedule. Results of the current pilot study, although underpowered, conceptually replicated the original claims and extended their work by investigating their questions in the area of mathematics computation. Implications for research and practice are shared.
Two fundamental understandings of science are the premise that a single study, regardless of its findings, cannot stand alone and be considered “true” (Ioannidis, 2012) or that such findings are completely generalizable to all contexts. Because of these understandings, the foundation upon which the field of special education has been built, it is necessary for researchers to conduct further studies that either confirm or refute such findings. Replication studies provide a systematic approach through which previous findings can be tested. Researchers in the field of special education have recently (and repeatedly) called for an increase in the number of replication studies conducted (e.g., Banerjee et al., 2018; Cook, 2014), with attention to naming the replication effort explicitly (Coyne et al., 2016). While some may be hesitant to label replication studies as such because of the historically diminished likelihood of publication or the emphasis on novel discoveries (Cook et al., 2016), the need to create a clear line of investigations that are easily traceable and scrutinized has the possibility of strengthening the knowledge about what works under what conditions and subsequently inform instruction and intervention practices. This article will explain three big ideas related to replication: types of replication studies, author overlap, and preregistration. Next, we explain our decision to extend the work of J. Jenkins et al. (2017) beyond curriculum-based measurement (CBM) administration schedules in reading and describe our conceptual replication of CBM administration schedules in mathematics. Finally, this article will report our conceptual replication efforts, including ongoing comparisons between the current study (in mathematics) and the original Jenkins et al. study (in reading), and conclude with implications for practice and research.
Types of Replication Studies
Although replication studies exist on a continuum (Rosenthal, 1991), there are two primary types: direct and conceptual. Direct replications intend to exactly reproduce a study’s procedures; conceptual replications employ different methods but are aimed at testing the same hypotheses as a previous study (Schmidt, 2009). Direct replications in special education are relatively nonexistent (Therrien et al., 2016), likely because of the applied nature of working with students and teachers in varying school contexts. Conceptual replications, while still relatively scarce (Therrien et al., 2016) allow researchers to attend to the same hypotheses as the original study, while intentionally altering key factors within the study, such as study participants or the academic domain in which the study is carried out. These types of alterations contribute to the generalizability of findings by describing the conditions under which the original findings hold and the conditions under which they do not. Although many studies in special education may appear to be conceptual replications or may inadvertently replicate existing findings, few are intentionally conducted within a “replication framework” (Coyne et al., 2016, p. 244). For example, studies that constitute a research program were not necessarily designed and conducted to replicate previous findings but to, more broadly, contribute to the knowledge within that body of work. Naming an investigation as a replication frames the authors’ intentions and provides a lens through which to examine the method, analyses, and findings.
Author Overlap
Of the special education studies that are identified as replications, author overlap is common (Lemons et al., 2016; Makel et al., 2016). That is, one or more authors are part of both the original study and the replication. In replications that have author overlap, it is more likely for findings from the original study to be replicated (i.e., confirmed; Makel & Plucker, 2014). Researchers have pointed out the potential limitations of interpreting replicated findings from teams that include author overlap because of the effects of research bias that would be present in both studies, as well as the reluctance author teams might have to publish findings that contradict their previous results (Cook et al., 2016). Given the importance of replication studies in special education, an additional criterion could be to employ a unique set of authors for subsequent studies to mitigate issues of bias and reporting.
Preregistration
Another approach to mitigating bias and reporting is to preregister a research plan, typically through an online platform dedicated to making plans available (e.g., https://osf.io/). Preregistration entails making one’s methodological and analytical plans transparent prior to conducting the study or analyzing the data. Authors articulate their anticipated methods and analyses, and upload this document through an online platform, which makes the plan accessible. The benefits of preregistration include making unpublished findings discoverable for inclusion in meta-analyses (Cook et al., 2018), providing an unambiguous roadmap for direct or conceptual replications, and discouraging post-hoc analytical choices that cater to statistical significance (Cook, 2014). Although there are practical challenges related to adhering to a preregistered plan (e.g., Nosek et al., 2018), the plan is not intended to stifle changes to the method or analyses (that commonly arise when conducting research in applied settings), but rather, to document the researchers’ decision-making. Documenting the researchers’ decision-making within a replication framework carries additional importance as differences between the original study and the replication can be understood in relation to the reported findings.
Weekly Versus Intermittent Progress Monitoring: Replicating J. Jenkins et al. (2017)
Replication efforts should be aimed at studies that address significant issues in the field of special education. CBM is a procedure used to monitor student progress and was developed to improve the educational outcomes of students with disabilities (Deno, 1985, 2003). CBM are short-duration measures administered to capture student progress in key academic areas. These brief measures were developed to serve as indicators of student performance and progress and are predictive of long-term and high-stakes academic outcomes. In mathematics, CBM are typically skills-based measures that include one domain (e.g., magnitude comparison and computation) and mixed skills (e.g., addition, subtraction, multiplication, and division; Kelley et al., 2008). One tenet of CBM is the one-probe-per-week schedule intended for students with the most persistent needs, and it is the most commonly used (e.g., Mellard et al., 2009). Although special education teachers’ use of CBM has increased and remains a staple of teachers’ progress monitoring (PM) practice (Swain & Hagaman, 2020), time to administer probes is a consistently reported barrier (Deno, 2003; Swain & Hagaman, 2020).
In response to this barrier, J. Jenkins et al. (2017) questioned whether the traditional schedule was feasible for teachers to maintain and necessary for timely instructional decisions or whether an intermittent schedule, where the schedule is less consistent or occurs with less frequency than weekly, could be just as effective. Traditionally, an argument in favor of the weekly PM schedule is the ability to make accurate and timely instructional decisions. These decisions are the result of examining students’ short-term growth and determining whether growth is expected (i.e., aligned with or surpassing the growth rate toward the goal) or unexpected (i.e., below the expected rate of growth and not on pace to reach the goal). If growth occurs as expected, the intervention is retained, and the student stays the course. If growth does not occur as expected, an adaptation is made.
In their original study, J. Jenkins et al. (2017) found the majority of intermittent PM schedules were at least as accurate as the traditional weekly schedule and that intermittent PM schedules did not delay decision accuracy. Gesel and Lemons (2020) recently conducted a conceptual replication of the Jenkins et al. study also in the area of reading fluency with minor differences in method and with an additional correlational analysis to determine the relation between teachers’ self-reported instructional changes and students’ growth slopes. Gesel and Lemons concluded that their study replicated Jenkins et al.’s original findings. Other studies have investigated the most timely and accurate PM schedule in reading (e.g., Bulut & Cormier, 2018; Christ et al., 2013), and although other researchers have investigated the impact of intermittent PM schedules in the area of mathematics using computer-adaptive tests (Nelson et al., 2017; Van Norman et al., 2016), at this time we did not find any studies that utilized CBM probes to address these questions. Understanding whether there are content-area differences related to different PM schedules informs the generalizability of previous findings that suggest intermittent schedules can be accurate and result in timely decisions compared with the traditional one-every-week schedule. Furthermore, in comparison with reading, there is simply less research in the area of mathematics (Lembke et al., 2012), therefore findings from this study have implications for future replication efforts and practical implications for what constitutes best PM practice in an understudied domain of special education intervention.
This study was designed within a replication framework, without author overlap, and includes a preregistered plan (DOI: 10.17605/OSF.IO/P3SCN) that was created prior to data collection. The purpose of this study was to conceptually replicate the work of J. Jenkins et al. (2017) and determine the degree to which the studies arrived at similar conclusions about the accuracy and usefulness of different PM schedules within different academic domains. In the spirit of replication, we adopted Jenkins et al.’s research questions:
Method
Differences Between Preregistered Plan, J. Jenkins et al. (2017), and Current Study
Table S1 in the online supplemental materials details the similarities and differences between our preregistered plan and the actual study with a rationale for why differences occurred. The main differences were sample size and the analysis plan. Our original plan was to recruit 70 students and analyze a final sample of 56 students, congruent with the sample size used by J. Jenkins et al. (2017). The onset of a global pandemic resulted in the closure of a second cohort of participating schools and interrupted data collection of a spring cohort (N = 17). Because students in the spring cohort had half as many datapoints as the fall cohort, the spring cohort was not included in analyses given the limited and speculative nature of comparisons that could be made in relation to the Jenkins et al. study. We recognize that even under nonpandemic conditions, a sample of 34 would have still resulted in a meaningfully smaller sample than the original study. However, given the general scarcity of mathematics intervention research (in comparison to reading intervention research; Lembke et al., 2012), this pilot study adds to the mathematics intervention knowledge base and suggests future investigations.
Another meaningful difference was our analysis plan. Originally, we planned to calculate digits correct per minute. However, J. Jenkins et al. (2017) calculated a words-read-correctly-per-week metric, which we matched in our analysis and created a digits-correct-per-week metric. When choosing an expected growth rate Jenkins et al. used existing reading fluency literature to set their expected growth rate at 1.0 word per week. We initially thought a similar expectation would be appropriate, but then discovered that Fuchs et al. (1993) established grade-specific growth rates for the probes used in this study. We used the grade-specific growth rates to reflect not only the difference in probe content (e.g., reading fluency v. mathematics computational fluency) but also to more accurately determine whether students made expected growth in relation to the 25th percentile norms (which reflects, in part, the unique growth rates we might expect from students working off grade level). Furthermore, we used the 25th percentile to reflect the broadest view of students considered struggling in mathematics and because of its common usage in other CBM research (e.g., Clarke et al., 2011).
Table S2 details the differences between this study and the J. Jenkins et al. (2017) study. Points of comparison were taken from Coyne et al.’s (2016) suggested dimensions. Per other recommendations (e.g., Cook et al., 2016), our purpose in sharing this table was to clearly and fully describe the degree to which our study replicated Jenkins et al. and in what ways the studies differed. The main differences between the Jenkins et al. study and this study were the result of the different academic domains. Because our study employed mathematics CBM, there were differences in measures, CBM administration, and students’ instructional levels, compared with the reading fluency measures used by Jenkins et al. Although oral reading fluency (as assessed in the Jenkins et al. study) and computational fluency (as assessed in this study) are meaningfully different, both measures are understood to be conventional measures related to their respective content areas. The research questions addressed in this study are about decision-making accuracy and timeliness and, therefore, differences that may be related to the nature of the measures are of secondary interest.
The other main differences not specifically attributed to the difference in CBM content were the difference in determining students’ instructional level and CBM probe selection. Jenkins et al. used teacher-specified instructional levels and administered probes at that level for each student. We used the teacher-specified instructional level as a springboard for further baseline probe administration. Additionally, J. Jenkins et al. (2017) used probes created by two different vendors to create a set of unique reading fluency probes; we decided to readminister a random set of the same probes (see the CBM Administration subheading for details about determining students’ instructional level and how we created our set of probes).
Sample
In the fall of 2019, teachers were recruited from Districts A and B. Both school districts were located in small rural cities in the U.S. Midwest and had special education populations between 4% and 7%. After obtaining Institutional Review Board and school district approval, we contacted the special education directors at both districts who assisted us in identifying special education teachers. A total of six teachers across the two districts agreed to participate. All teachers were White and female. None of the teachers were using a data-based individualization process or regularly using CBM prior to participating in this study. Teachers identified students who met the following criteria: (a) were in Grades 2–5 and (b) were receiving special education services for either high- or low-incidence disabilities (see Table 1 for student demographics). Once students were identified, teachers sent consent forms home. The total consented sample was small (N = 17), in general, and in comparison with the J. Jenkins et al. (2017) sample (see Table 1 for comparison between student samples). Two students were excluded from the final sample due to behaviors incompatible with study participation (e.g., writing random answers on probes). We considered the research questions in this study to be about issues of assessment and not issues of intervention or implementation. If we had a different purpose additional steps would have been taken to include students who had difficulty completing weekly probes; however, it was beyond the scope of this study to address practical issues of implementation or student behavior. Of the 17 consented students, one student was missing 1 week of data. The difference between the true-growth slopes of students with and without complete data was not statistically significant, t(15) = −0.13, p = .91.
Comparison of Student Demographics for the Current Study and the J. Jenkins et al. Study.
Note. Total N = 17 for current study; N = 56 for Jenkins et al. study. Instructional level not reported for Jenkins et al. study given the difference in CBM content. LD = learning disability; S/LI = speech and/or language impairment; OHI = other health impairment; DD = developmental delay; ID = intellectual disability; AU = autism.
Measure
The MBSP: Monitoring Basic Skills Progress, Basic Math (Fuchs et al., 1999) measures were chosen given empirical evidence supporting their reliability (r = .80–.90), criterion validity (r = .62–.82), and ability to detect student growth over time (Jiban & Deno, 2008). Although there are newer measures available, these particular measures were used because they were available for use across the two school districts and allowed for consistency. Furthermore, other measures, such as the AIMSweb probes used in the J. Jenkins et al. (2017) study, would have required technology beyond the available resources of either participating district. The MBSP measures consist of 30 parallel probes, each of which includes 25 skills-based measure items. Probes range from Grades 1 to 6 and measure computational skills across addition, subtraction, multiplication, and division. Items in the Grades 4 to 6 probes include fractions; items in the Grades 5 to 6 probes include decimals.
Procedures
Administration training
Examiners were five special education graduate students (including the first author) trained and experienced in administering and scoring CBM probes. One, 1-hr training session was conducted by the first author. The training addressed the purpose of the study, introduced examiners to the MBSP probes, reviewed MBSP administration directions, and discussed the weekly administration procedure. MBSP administration directions included scripts for initial and subsequent administrations. For this replication, researchers created supplemental materials, which were referenced in the MBSP administration directions, but not provided (e.g., a sample division problem with quotient written with a remainder, instead of a fraction).
CBM administration
As a starting point, we asked teachers to identify students’ instructional level in mathematics. J. Jenkins et al. (2017) used the teacher-identified instructional level throughout their study. We used the teacher-identified instructional level to administer the first three MBSP probes in the corresponding grade-level set. The first three probes in each grade-level set were only used to determine students’ instructional level and baseline scores and were not readministered for the remainder of the study. An instructional level was determined when a student’s median digits-correct score fell between the 25th and 50th percentiles for fall normative digits-correct scores for that grade level according to MBSP norms (Fuchs et al., 1999). For the majority of students (n = 10; 59%), the teacher-identified instructional level was accurate; other students completed additional sets of three MBSP probes at different grade levels until their instructional level was identified. Given the small sample size, it was not possible to assess within-grade variance.
A random number generator was used to assign a random sequence of the remaining 27 probes to each student at their instructional level. After the remaining probes were administered one time, students completed an additional randomized sequence of probes to account for the 42 total datapoints. Although the choice to readminister a random sample of already-completed probes increased concern about practice effects, repeated probes were possibly separated by a 10-week interval (as recommended by J. R. Jenkins et al., 2005) potentially mitigating this concern. Furthermore, J. Jenkins et al. (2017) recommended this procedure, highlighting the benefit of using forms that were known to be equivalent.
Aligned with the J. Jenkins et al. (2017) study, students completed three MBSP measures every week during Weeks 1 to 11 and six MBSP measures during Week 12. Each week, one graduate student would administer the MBSP measures to a small group (i.e., 1–5 students), in a quiet location, using a standard administration script. When possible, probe administration occurred on the same day and same time each week, though due to unexpected school events (e.g., assemblies) and student absences, there were expected administration variances. Students working at a Grades 1 to 2 instructional level had 2 min per probe; students working at Grades 3–4 instructional level had 3 min per probe; and students working at Grade 5 instructional level had 5 min per probe. Administration fidelity was audio recorded and checked once per graduate student across the 12 weeks of the study and resulted in 100% fidelity for all examiners.
Scoring and reliability
MBSP scoring guidelines were used. Researchers added language in the scoring guide to clarify how to score correct digits when the answer was recorded as a fraction. Aligned with J. Jenkins et al. (2017), all CBM were originally scored by members of the research team, with the majority of original scoring conducted by one special education graduate student. Percent agreement was calculated by dividing the number of agreements by the total number of digits, then multiplying the quotient by 100. The first author rescored all CBM, which yielded 96% initial agreement on number of digits correct; discrepancies were examined and resolved. Next, all CBM scores were double entered into a Microsoft Excel spreadsheet, which yielded 99% accuracy; discrepancies were examined and resolved.
Analyses
Mirroring the analyses conducted by J. Jenkins et al. (2017), we used ordinary least squares regression to calculate seven slopes per individual student: one true-growth slope and six short-term-growth slopes that corresponded to the six intermittent PM schedules being tested. A student’s true-growth slope was calculated by using all 42 PM measures, resulting in the best estimate of the student’s actual growth over time. The six slopes that corresponded to the six PM schedules were (see Table 2 for a visual distribution of probes across schedules): one probe every 1 week (calculated using baseline scores and the first probe administered during Weeks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, and 12 for a total of 15 probes), two probes every 2 weeks (calculated using baseline scores and the first two probes administered during Weeks 2, 4, 6, 8, 10, and 12 for a total of 15 probes), three probes every 3 weeks (calculated using baseline scores and the three probes administered during Weeks 3, 6, 9, and 12 for a total of 15 probes), three probes every 4 weeks (calculated using baseline scores and the three probes administered during Weeks 4, 8, and 12 for a total of 12 probes), three probes every 5 weeks (calculated using baseline scores and the three probes administered during Weeks 5 and 10 for a total of 9 probes), and three probes every 6 weeks (calculated using baseline scores and the three probes administered during Weeks 6 and 12 for a total of 9 probes). Per Jenkins et al.’s rationale, students’ three baseline scores were included in each schedule to “achieve a reliable estimate of baseline performance and ensure a common starting point” (p. 46).
PM Schedules for Decision Points Over Time.
Note. Numbers not in parentheses are accuracy of the current study as a percentage; numbers within parentheses are accuracy of the J. Jenkins et al. study as a percentage; italic numbers are score overlaps as a percentage. Binomial test; no correction for multiple tests. For both studies, score overlap was the same and was calculated as the number of PM scores following baseline (n/39). PM = progress monitoring; NR = not reported.
p < .05. **p < .01.
Following J. Jenkins et al.’s (2017) procedures, slopes were calculated using individual probe scores and a time value that was calculated as the number of days since the first baseline measure, divided by 7. Then, because students completed three probes during each weekly session, 5 min (0.003 days) were added to the second and third probes administered each week to distinguish each probe as occurring at a distinct point in time, which is consistent with the procedures used by Jenkins et al.
Assessing student growth
In addition to students’ true-growth slopes, students’ short-term growth was determined using MBSP rates of growth (Fuchs et al., 1993). Short-term expected growth rates differed by grade level and were .30 digits per week for Grades 1 to 3 and .70 to .75 digits per week for Grades 4 to 5. To calculate students’ short-term growth, we compared their slope that week to the rate of growth associated with their instructional level and determined the growth to be expected (i.e., equal to or greater than the expected rate of growth) or unexpected (i.e., less than the expected rate of growth).
Decision accuracy and timeliness
Next, we examined the relation between each PM schedule and decision-making accuracy. That is, we asked whether deciding to either stay the instructional course or make an instructional change would be the same or different depending on whether the true-growth slope was consulted versus the short-term-growth slopes produced by each intermittent PM schedule. Similar to J. Jenkins et al. (2017), decision accuracy was determined as the percentage of instances in which the true-growth slope and short-term-growth slopes would have resulted in the same instructional decision. Thus, statistical contrasts were conducted to determine which PM schedules were accurate significantly above chance (i.e., 50%) compared with the traditional one-probe-every-week schedule. Finally, the criteria of interest set by Jenkins et al. were when a PM schedule reached either 70% or 75% accuracy, which they called an “intermediate level of decision accuracy” (p. 51), possibly facilitating teachers’ timely decision-making. We analyzed during what week, if ever, each PM schedule reached either accuracy threshold.
Results
Across the 13 weeks, the mean true growth for the sample was .17 digits per week (SD = .39), with a slightly skewed distribution (1.08; Field et al., 2012). The majority of the sample (n = 12; 70.59%) did not achieve their instructional grade-level goal rate of either .30 for Grades 1 to 3 or .70 to .75 digits per week for Grades 4 to 5. True growth was not significantly correlated with either grade level (r = .37) or instructional mathematics level (r = −.23).
Decision Accuracy
Decision accuracy is represented in Table 2. Similar to J. Jenkins et al. (2017), decision accuracy increased over time, though not in a linear fashion (e.g., Week 7). The rows of Table 2 display the intermittent PM schedules and the columns display the probes that would be assessed if an instructional decision were made at that week. We reported accuracy percentages for all weeks of the study, whereas Jenkins et al. started reporting from the first week statistical contrasts yielded a significant binomial test (i.e., Week 4). Statistical contrasts were conducted when at least one intermittent schedule and the traditional one-every-week schedule could be compared. In this study, in 4 of the 9 weeks for which comparisons were reported all intermittent schedules were at least as accurate as the traditional one-every-week schedule, including 11 of the 17 (64.71%) individual contrasts. Table 2 also reports the percentage overlap between the 39 probes used to determine the true-growth slopes and the number of probes used to calculate each short-term-growth slope.
Timeliness
In alignment with J. Jenkins et al., (2017) we evaluated schedules for if and when they met the 70% or 75% accuracy thresholds. Supplemental Materials S3 illustrates the first week each schedule reached either threshold. In this study, the three-every-5-weeks schedule met the 70% and 75% accuracy criteria the earliest (Week 5). The next most timely schedule was the three-every-6-weeks schedule (Week 6). The remaining schedules, including the traditional one-every-week schedule, did not meet the 70% or 75% accuracy criteria until Week 8 or 9.
Discussion
The purpose of this study was to conceptually replicate the study carried out by J. Jenkins et al. (2017), but in the area of mathematics. Research conducted within a replication framework contributes to the science of special education and works to increase the generalizability of that science. We replicated Jenkins et al.’s inquiry about the accuracy and timeliness of intermittent PM schedules compared with the traditional one-every-week schedule. Although to a lesser extent than J. Jenkins et al. (2017), we found the majority of intermittent schedules to be of equal or greater accuracy and timeliness compared with the traditional one-every-week schedule. These findings replicate the original claims made by Jenkins et al. and extend the original findings into another content domain.
Decision-Making Accuracy
Intermittent PM schedules varied in accuracy across time compared with the traditional one-every-week schedule. If making an instructional decision during Week 2, the intermittent schedule would have been more accurate, but not at the 70% criteria; however, during Weeks 3 and 4, the traditional 1-every-week schedule was more accurate, but not at the 70% threshold. For Weeks 5, 6, 8, 9, 10, and 12, the intermittent schedules did perform at least as accurately (except for deciding at Week 6 for the two-every-2-weeks schedule) and met the 75% accuracy threshold. One explanation for the increase in decision accuracy over time is the score overlap. The 42 probes used to calculate a student’s true-growth slope were also used to calculate a student’s short-term-growth slopes using subsets of the same probes. Therefore, the overlap reported in Table 2 illustrates that as later weeks were assessed, score overlap increased.
J. Jenkins et al. (2017) noted the “relatively strong performance” (p. 49) of the three-every-3-weeks schedule, which we interpreted in light of it being one of two schedules for which all reported contrasts were statistically significant, and accuracy was at or above the 75% threshold for the duration of that schedule (i.e., during Weeks 6, 9, and 12). In this study, such a claim was not supported. The only schedules for which those two criteria held up were the three-every-5-weeks and three-every-6-weeks schedules. The Jenkins et al. study also observed the three-every-6-weeks schedule as performing both accurately (i.e., above the 75% threshold) and no timelier than the three-every-3-weeks schedule (i.e., at Week 6). Perhaps one reason the three-every-3-weeks schedule was highlighted was because the three-every-3-weeks schedule would allow teachers to make an instructional decision at Weeks 6, 9, and 12, compared to the three-every-6-weeks schedule, which would result in decisions at Weeks 6 and 12. Increased opportunities for decision-making aligns with one of the core features of using CBM to conduct PM: measures that are sensitive to growth over short periods of time (Deno, 1985).
Timeliness
In this study, the traditional one-every-week schedule reached the 70% and 75% accuracy thresholds in Week 8. Compared with the two-every-2-weeks, three-every-3-weeks, and three-every-4-weeks schedules, referencing the one-every-week schedule would not have delayed or meaningfully expedited decision-making. The three-every-5-weeks and three-every-6-weeks schedules, however, would have resulted in an instructional decision as early as Week 5 or 6, respectively. Supplemental Materials S3 compares the current study and the J. Jenkins et al. (2017) in terms of when different schedules meet different accuracy thresholds across the two studies. The PM schedules assessed in this study reached the 70% accuracy threshold later or at the same time as the Jenkins et al. study. However, when comparing the two studies against the 75% accuracy threshold, the majority of PM schedules reached that threshold earlier or at the same time in the current student compared with the Jenkins et al. study (with the exception of the three-every-3-weeks schedule).
J. Jenkins et al. (2017) and others (e.g., Gesel & Lemons, 2020) have asked what accuracy percentage is the most optimal for decision-making so that accuracy and timeliness are equitably attended to. Jenkins et al. speculated that adopting a higher accuracy criterion (e.g., 80%) could result in delayed decision-making. However, in the current study, adopting the three-every-5-week schedule could have yielded a timely decision (as early as Week 5) that surpassed that higher accuracy criterion; schedules in the Jenkins et al. study did not reach 80% accuracy until Week 12. Jenkins et al. characterized their findings as “a beginning database for guideline development” (p. 50) in terms of what constitutes acceptable accuracy criterion and in relation to when that criterion is met. We view our findings as similarly adding to this discussion with new evidence about the sensitive balance between decision-making accuracy and timeliness. This contributes to existing literature such as that by Christ and colleagues (Ardoin et al., 2013; Christ et al., 2013) that provides suggestions for number of weeks of instruction and number of data points necessary for high-quality decision-making and whether duration of instruction or number of data points should be the primary focus. It should be noted that the Christ work cited here was in the area of oral reading, so there may be differences when examining progress monitoring data in mathematics.
Did This Study Replicate the Findings of the Original Study?
Although there were differences pertaining to both accuracy and timeliness, overall, the findings from this study did in fact replicate those reported by J. Jenkins et al. (2017) both in broad and nuanced ways. In terms of accuracy, Jenkins et al. found a greater majority of individual contrasts (86.67%) were at least as accurate as the traditional one-every-week schedule; this study found a smaller majority of individual contrasts (64.7%) were at least as accurate as the traditional one-every-week schedule. In terms of timeliness (i.e., reaching the 70% or 75% accuracy threshold), in both studies, only one intermittent schedule would have resulted in a delayed decision compared to the traditional one-every-week schedule (the three-every-3-weeks schedule in this study and the three-every-4-weeks schedule in the Jenkins et al. study).
Limitations
Given the unique approach used in this study to evaluate the usefulness of different progress monitoring schedules, the limitations of this study are similar to those identified by J. Jenkins et al. (2017) and the Gesel and Lemons (2020) conceptual replication; the score overlap reported in the original study and the other conceptual replication also apply here. Because the true-growth and short-term-growth slopes were calculated using either all or a subset of the same scores, this investigation could not disaggregate the unique contributions of either score overlap or the PM schedule itself. Possible solutions to this issue include employing a unique set of measures for true growth or extending the duration of the study (e.g., 15 weeks).
In addition to score overlap, differences in the samples included in the two studies should be addressed. In addition to expected differences in student demographic characteristics, the small sample size in the current study limits the interpretation of these data especially in comparison to Jenkins et al. and their generalizability. The onset of a global pandemic resulted in the closure of a second cohort of schools, yielding a meaningfully smaller dataset (N = 17) than the Jenkins et al. study (N = 56). Therefore, we caution readers to bear this difference in mind as results are interpreted. Future iterations of this study should secure a sample more comparable in size with the Jenkins et al. so that comparisons can be made with greater confidence. Additionally, although true growth was not significantly correlated with grade level (r = .37), with a larger sample size, such correlations may have reached significance.
Implications for Research and Practice
As the special education research community continues to prioritize replication studies, the field can systematically expand the depth and breadth of its understanding. The findings reported in this study carry implications for other researchers given that this effort was intentionally situated in a different content area. Future studies should continue to explore the use of mathematics CBM within intermittent and traditional PM schedules, as well as expanding this body of work to include other content areas (e.g., writing and science) and other types of mathematics probes (e.g., single-skill probes). Researchers have noted differences in students’ growth rates across reading, mathematics, and writing CBMs (e.g., Codding et al., 2015), suggesting that a better understanding of how different content area CBMs function within different PM schedules is warranted. However, with the limitations in mind, the findings from this study suggest the possibility that intermittent PM schedules might be a practically useful practice to continue understanding.
The process of conducting replication research necessitates a transparent description of our growing pains (Cook et al., 2018). For us, maintaining the preregistered plan—updating it as methodological changes occurred, revisiting it at pre-determined phases of the study—was not an established routine within our research practice. Therefore, updates to our plan did not occur as frequently as is likely optimal, despite our commitment to increased transparency. Researchers can integrate these new considerations into their existing protocols to take active strides to change the norms around conducting investigations in this field.
Findings from J. Jenkins et al. (2017), Gesel and Lemons (2020), the current study, and others (e.g., Christ et al., 2013; January et al., 2019) have implications for the recommendations made to practitioners about what constitutes best practice for PM. To echo a discussion Gesel and Lemons started, the degree to which group-level decision-making is appropriate can and should be scrutinized. That is, the utility of group-level goal setting or growth expectancy that works for all students in a classroom is likely limited. Ongoing investigations could make comparisons between a group-level growth scheme and an intra-individual growth scheme (Hosp et al., 2016) to articulate the affordances and limitations of either approach.
Finally, there is a persistent need to support teachers and other school-based personnel to collect, interpret, and make decisions based on CBM data (see Fuchs et al., 2021 for a description of innovations to increase teacher engagement with CBM). This study focused on time, which has been a consistent teacher-reported barrier to regular data collection (Deno, 2003; Swain & Hagaman, 2020). Teachers may be experiencing difficulty conducing regular PM in general, though a recent study indicated that over the last 20 years special education teachers increased their use of CBM in reading for PM purposes but decreased their use of CBM in mathematics (Swain & Hagaman, 2020). Perhaps, under time constraints, special education teachers prioritize CBM in one area (e.g., reading) even if CBM is needed in other areas (e.g., mathematics and writing). The results from this study present a possible solution to this real concern. Adopting an intermittent PM schedule might simultaneously decrease demands on teachers’ time, while increasing the accuracy and timeliness of instructional decisions across content areas.
Supplemental Material
sj-docx-1-aei-10.1177_15345084221133730 – Supplemental material for Weekly Versus Intermittent Progress Monitoring in Mathematics: A Conceptual Replication and Pilot Study
Supplemental material, sj-docx-1-aei-10.1177_15345084221133730 for Weekly Versus Intermittent Progress Monitoring in Mathematics: A Conceptual Replication and Pilot Study by Erica N. Mason and Erica S. Lembke in Assessment for Effective Intervention
Footnotes
Acknowledgements
The authors thank the following people for their assistance with data collection: Stacy Hirt, Stephanie Hopkins, Jiyung Hwang, and Elizabeth Thomas. The authors thank Samantha Gesel for feedback on an earlier draft. Finally, the authors thank Bryan Cook, Bill Therrien, and the Consortium for the Advancement of Special Education Research for their guidance on conducting and reporting replication studies.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material is available on the Assessment for Effective Intervention webpage with the online version of the article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
