Abstract
The transition from sounding out unfamiliar words to effortlessly reading connected text does not occur all at once or at the same rate for students. The purpose of this study was to explore the accuracy of three decision rules (data point, median, and trend line) applied to progress monitoring outcomes of alphabetic principle (nonsense word fluency [NWF]) and oral reading rate (curriculum-based measurement of reading [CBM-R]). Outcomes from a sample of U.S. students receiving Tier-2 supports in oral reading and decoding were analyzed to generate model parameters. Scores were simulated for NWF and CBM-R, and decision rules were applied to schedules where one observation was collected per week. The trend-line rule was viable with NWF after 7 weeks and 9 to 10 weeks with CBM-R. Differences in base rates of non-proficiency between measures call into question the utility of NWF to capture student improvement in alphabetic principle as they encounter increasingly complex word types.
Most assessment research in the academic response to intervention literature has focused on the identification of students at risk of later difficulties (Ball & Christ, 2012). Effective universal screening is foundational to ensure that students truly in need of supplemental support receive intervention (Compton et al., 2010). It is also critical to accurately determine whether students receiving supplemental intervention are improving at a sufficient rate or a change needs to be made (Stecker et al., 2008). Similar to errors in universal screening, incorrect evaluations of student progress have negative consequences. Continuing to deliver ineffective interventions may cause students to fall further behind, and continuing supplemental instruction for students no longer experiencing academic difficulties diverts resources from students truly in need (Fuchs & Fuchs, 2017). As we cannot predict which interventions will work perfectly, ongoing data collection and suitable decision-making frameworks are necessary to accurately capture student response to instruction to ensure interventions are appropriately matched to student need (Deno, 2016). The purpose of this study was to compare the accuracy of recommendations from various decision rules applied to progress monitoring data from two types of curriculum-based measures (CBM; Deno, 1985): nonsense word fluency (NWF) and oral reading fluency.
Early Reading Skill Development
The measure that educators use to monitor student progress should closely align with the skill being taught (Shapiro & Guard, 2014). However, as students learn to read, the progression from mastering early skills, such as sounding out individual letter sounds, to more complex skills, such as reading connected text aloud, does not occur at the same time or at the same rate for all students (Compton et al., 2022). The point is accentuated when considering students experiencing reading difficulties (Cummings et al., 2011), who are most likely to be receiving supplemental support.
Fluent reading of connected text depends on accurate and efficient word recognition (Jenkins et al., 2003). For most early elementary students, sublexical (i.e., below the word level) processing increases as letter–sound correspondence, phonemic segmentation, and blending skills are acquired (Ritchey & Speece, 2006). As increasingly complex orthographic-phonological (i.e., spelling–pronunciation) connections are added to and retrieved from long-term memory, word identification becomes more automatic (Ehri, 2002). As students make progress toward oral text reading, pseudoword reading can be used to index the ability to apply letter–sound correspondence knowledge when decoding unknown words (Torgesen et al., 1999). Early unitization in pseudoword reading may predict general reading outcomes over and above growth in individual sound identification (Clemens et al., 2018). As alluded to before, the rate in which students progress toward acquiring more complex skills is not uniform. It is critical that there is alignment between the student’s developmental level in reading acquisition, the skill intervened upon, and the assessment used to index progress toward acquiring skills (Clemens et al., 2020).
Curriculum-Based Measures
CBMs are a family of brief standardized assessments that measure student performance in basic academic skill areas (e.g., reading, writing, and mathematics; Deno, 1985). The most widely used type of CBM is oral reading (CBM of reading [CBM-R]) in which students read connected text for 1 min (Tindal, 2013). The number of words read correct per minute (WRCM) is the primary metric of interest and has displayed robust evidence as an indicator of broad reading competence (Shin & McMaster, 2019). Fluent oral reading requires the successful execution of a number of lower order interrelated skills. When students experience difficulties in oral reading, it may be necessary to intervene upon any number of lower order skills. To monitor response to those interventions, downward extensions of CBMs have been developed to assess more basic reading skills such as phonemic awareness and alphabetic principle (Good et al., 2001). As students master these lower order skills, often by the end of first grade, it is generally expected that most students are monitored with CBM-R through elementary school. Prior to that point, a variety of measures are available to assess student progress.
CBM assessment systems (e.g., aimswebPLUS, NCS Pearson, 2017; Dynamic Indicators of Basic Early Literacy Skills [DIBELS] 8, University of Oregon, 2021; EasyCBM, Alonzo et al., 2006) typically follow a mastery learning model or short-term measurement approach to monitoring early literacy skills (Fuchs & Deno, 1991). Vendors usually recommend administering similar skill-specific measures to monitor growth at certain stages of reading development. Yet recommendations for when to use various measures to monitor student progress differ considerably by vendor. For example, the newer DIBELS 8th edition incorporates measures of NWF from the beginning of kindergarten in addition to word identification fluency and connected text reading by the beginning of first grade. In contrast, the early literacy AIMSweb assessment schedule focuses on the monitoring of more granular skills (e.g., letter–sound fluency, phoneme segmentation fluency, and NWF) until mid–first grade and word identification is not used at any point during the transition to oral reading.
Transitions to different interventions over the course of early literacy development may pose problems in ensuring that assessments are closely aligned to the skills being intervened upon. Alignment between intervention and assessment is relevant, given the substantive differences between measures, one example being NWF and CBM-R. The decision as to whether NWF or CBM-R should be used to measure response to intervention is presumably dependent on whether a student is able to apply appropriate strategies to decode simple unfamiliar words. NWF is more likely to isolate students’ ability to decode, whereas CBM-R is a more comprehensive assessment of reading in context (Fien et al., 2008). After students have acquired foundational skills in decoding, it is often beneficial to allow those students to practice reading connected text while also receiving targeted phonics instruction for more complex word types (Moats, 1998). In turn, there are scenarios in which students may receive a multicomponent intervention, simultaneously targeting reading fluency and phonics (Parker et al., 2022). In situations where the intervention targets multiple skills, there may be ambiguity in the appropriate assessment to monitor progress. In this example, one may collect both NWF and CBM-R data. However, if different rates of improvement are observed between measures, the degree to which supplemental supports would be deemed effective and the resulting decisions might depend on the measure used to monitor progress.
Decision Rules
Regardless of the measure selected to monitor student progress, it is vital that educators regularly collect, plot, and evaluate outcomes to inform treatment decisions. In general, the process begins by creating a goal line, or expected rate of improvement between a student’s initial baseline score and an end of year target or normative growth referent (Shapiro & Guard, 2014). After intervention begins, data are then collected weekly, biweekly, or along some other predetermined schedule (Stecker et al., 2008). Consistent performance below the goal line indicates that a student is not making sufficient progress and a change (e.g., intervention intensification) needs to be made. Scores that consistently exceed the goal line suggest the educator may increase the ambitiousness of the goal, exit the student from the intervention, or, for students receiving intervention on multiple skills, change the intervention focus. Inconsistent performance that falls above and below the goal line suggests that the educator should continue the intervention and collect more data before deciding (Ardoin et al., 2013).
Previous research suggests that educators are more likely to make an instructional change when students are not showing sufficient progress if they are provided with explicit prompts to do so (Stecker et al., 2005). Rather than asking an educator to visually analyze time series data and determine whether to make a change, a decision rule provides an automatic prompt if certain conditions are met. For instance, the data-point rule dictates that if the most recent three, four, or five observations fall below the goal line, a change is considered. If all observations fall above the goal line, the educator considers increasing the goal and, if any other pattern of results is seen, more data are collected before applying the rule (Ardoin et al., 2013). The trend-line rule involves estimating a line that best fit all available observations. If the slope of the trend line exceeds the slope of the goal line, the educator considers increasing the goal. If the slope of the trend line is less than the slope of the goal line, a change is considered. If slopes approximate one another, more data are collected before reestimating trend and applying the rule again (Ardoin et al.).
A large volume of research around data collection conditions and analytic strategies that improve the accuracy of structured prompts, or decision rules, has emerged since 2013, when Ardoin and colleagues (2013) concluded that no empirical studies had been conducted to evaluate the accuracy of existing recommendations for CBM-R decision rules. Although the circumstances in which decision rules are likely to yield more accurate recommendations when used in conjunction with CBM-R data have begun to emerge (e.g., Hintze et al., 2018; Van Norman et al., 2017), the research on decision rule use with other types of CBM is minimal. Yet decision rules developed primarily with CBM-R data serve as the basis for decision rule recommendations for other types of CBMs, such as NWF (Hosp et al., 2016).
Purpose
The purpose of this study was to compare the accuracy of recommendations from common decision rules when applied to concurrently collected CBM-R and NWF data. Previous research suggests that concurrently collected NWF and CBM-R captured growth in unique academic skills (Van Norman et al., 2018). What is unclear, however, is whether these differences translate to differences in the accuracy of evaluations of student response to instruction. In this study, we directly compared recommendations from common decision rules, applied to progress monitoring data collected for similar lengths of time.
Of primary interest was the number of weeks data needed to be collected from each measure for each rule to yield sufficiently accurate recommendations to support decision-making.
Method
A series of simulations were conducted in which true and observed growth were generated for CBM-R and NWF outcomes. Conditions were informed by an analysis of a large extant data set of outcomes from first-grade students who received phonics and oral reading fluency interventions simultaneously. Students were concurrently monitored with CBM-R and NWF. We first describe the sample and the procedure in which interventions were delivered and data were collected. We then provide information related to inferential analyses conducted to identify the conditions for simulation, the data generation process that used those parameters, and finally how the synthetic data were used to evaluate the accuracy of decision rules.
Participants
We analyzed an extant data set of progress monitoring outcomes from Reading Corps, a U.S. multistate, federally funded Americorps program that trained local community members to deliver Tier-2 reading intervention to students identified as at risk of reading difficulties. Results from an independent evaluation of Reading Corps, and description of general procedures, are offered in Markovitz et al. (2022). In a randomized controlled trial, Reading Corps interventions had medium effects on first-grade students’ NWF (effect size [ES] = .81) and CBM-R (ES = .61) scores (Markovitz et al.).
This sample included 579 first-grade students who received phonics and fluency interventions during the spring semester of the 2018–2019 school year. A total of 232 schools across 11 states were represented in the data set. The sample was primarily male (52%) and identified as native speakers of English (85%). In terms of race, most students identified as White (40%), followed by Black (39%), Latinx (8%), Asian (4%), multiracial (2%), and American Indian (1%). Free and reduced-price lunch status was not available at the individual student level per program guidelines. Roughly, 10% showed proficiency by the end of the school on CBM-R by obtaining at least two out of three successive WRCM scores that exceeded the spring target of 82.
Measures
Nonsense Word Fluency
The NWF measure from FastBridge Learning (Christ et al., 2018) was used to screen and monitor progress of students’ decoding skills. The NWF assessment contained consonant vowel consonant (CVC) decodable pseudowords. Students were allowed to read the entire word or sound out individual parts to earn credit. Credit was given for each sound sequence in a word and incorrect sequences were denoted by diagonal lines on the scoring sheet. The number of correct letter sequences (CLS) in 1 min was the outcome of interest. Median alternate form and internal consistency reliability estimates ranged from .69 to .96 and from .74 to .96, respectively, for first-grade students (Christ et al.). Predictive validity with the Group Reading Assessment and Diagnostic Evaluation test was reported as .61 (Christ et al.).
Curriculum-Based Measurement of Reading
To measure oral reading rate for screening and progress monitoring purposes, students read from grade-level CBM-R probes developed by FastBridge Learning (Christ et al., 2018). The number of WRCM was recorded by tutors for each progress monitoring observation. In first grade, FastBridge CBM-R median alternate-form, inter-rater, and internal consistency coefficients range from .74 to .94, .97 to .98, and .91 to .92, respectively. Concurrent and predictive validity using AIMSweb R-CBM as a criterion was .95 (Christ et al.).
Procedure—Assessment and Intervention
Students were enrolled in the program based upon benchmark performance on CBM-R or NWF during the 2018–2019 school year. Students who earned a median WRCM below 40 from three probes during winter benchmarking were included in this sample. Students who scored below 45 CLS on NWF were also included. Both benchmarks were based upon vendor benchmarks for proficiency developed for the program and correspond roughly to the 40th percentile for national Winter benchmark assessment norms (Christ et al., 2018). Student progress was measured once a week with one NWF and one CBM-R probe by Reading Corps tutors following vendor-recommended administration and scoring procedures. Missing data were not coded by tutors as part of program guidelines required them to collect a progress monitoring assessment at the next available session.
Reading Corps delivers interventions in a 1:1 format for approximately 20 to 30 min, 5 days per week. To be included in this study, students had to have received a combination of oral reading fluency and decoding intervention across the entire semester. Students who participated primarily in phonemic awareness interventions were not included. In addition, students who only received phonics intervention or only received oral reading fluency interventions were not included. The decoding intervention (“Blending Words”) targeted students’ ability to blend letter sounds into whole words. The intervention incorporated explicit instruction, corrective feedback, and repeated opportunities to practice blending sounds into words. Oral reading fluency interventions included newscaster reading (Searfross, 1975) and duet reading (Koskinen & Blum, 1986). Intervention and assessment fidelity data were collected for each tutor at least once a month using standardized checklists. Internal coaches corrected any deviations immediately after the tutoring sessions. Considering intervention fidelity, the average percentage of steps correctly followed by tutors across all intervention observations was equal to 97.17% (SD = 7.75%). The average percentage of steps correctly followed by tutors during data collection observations was equal to 98.91% (SD = 3.33%). Furthermore, 100% fidelity was observed across 89% of intervention observations and 100% fidelity was observed across 88% of data collection observations.
Procedure—Data Generation
Data were simulated in two steps. First, a multivariate linear mixed-effects regression model (Thum, 1997) was estimated using the nlme package (Pinheiro et al., 2022) in R (R Core Team, 2021). Fixed effects, random effects, covariance of random effects, and residual variance terms were estimated for CBM-R and NWF outcomes simultaneously. In this study, fixed effects capture the average initial level of performance (intercept) and weekly rate of improvement (slope) for CBM-R and NWF scores. Random effects for intercepts and slopes capture the level of variability of individual students from fixed effects for each measure. Covariance terms capture the strength of relationships between intercepts and slopes within and between measures. Residual variances quantify the level of error for each measure. Fixed and random effects were used to simulate CBM-R and NWF intercept and slope values for hypothetical students using the mvtnorm package (Genz et al., 2021) in R (R Core Team).
To explore the influence of the number of weeks data were collected on decision rule accuracy; 1,000 pairs of intercept and slope values for CBM-R and NWF values were simulated to reflect data collection schedules that spanned 4, 5, 6, . . . 16 weeks (n = 13 weeks), yielding a total of 13,000 sets of CBM-R and NWF intercept and slope values. A schedule where one CBM-R and one NWF observation were collected per week was assumed. The simulated intercept and slope values represented a hypothetical student’s true level of initial performance and true rate of weekly improvement on both measures. For a given duration condition, true scores at each week were estimated by adding the intercept to the product of the slope value and data collection week. The first week value was centered at 0 for all cases. Observed scores at each week, for each case, and for each measure, were also estimated by adding the appropriate residual variance term, based upon the multivariate linear mixed-effects regression models, to each true score. For WRCM, a random error term, with a mean equal to 0 and SD equal to 6.92, was added to each true score. For CLS, a random error term, with a mean equal to 0 and SD equal to 7.41, was added to each true score. Both values were derived from the previous multilevel analysis.
Procedure—Decision Rule Application
Sets of true and observed scores were available for concurrently collected NWF and CBM-R scores. A total of 1,000 hypothetical cases were available for each duration of progress monitoring (4, 5, 6, . . . 16 weeks). Goal lines for outcomes from each measure for each case were constructed by subtracting the spring target score for proficiency based upon vendor technical documents (70 WRCM and 60 CLS; Christ et al., 2018) from the initial observed score, or baseline score, and dividing that value by 16 (i.e., the number of weeks during the spring semester). The resulting goal lines quantified the average rate of weekly improvement in WRCM and CLS, respectively, required to reach the end of year target for that case.
True status of response to instruction for NWF and CBM-R for each case was determined by comparing the simulated true slope value with the slope of the respective goal line. If true slope was equal to or greater than the goal line, a 0 was documented to indicate that the student truly had shown sufficient response to instruction, whereas a 1 indicated that the student truly did not show adequate response to instruction. Next, three decision rules were applied to observed NWF and CBM-R outcomes for each case. First, the data-point rule was used by comparing the three most recent WRCM or CLS observed scores with the respective goal line. If all three observations were below the goal line, a 1 was documented for that case (inadequate progress, make a change); if any other pattern of observations was observed, a 0 was documented (continue the intervention). For the median rule, the median WRCM and CLS score from the three most recent observed scores was compared with the expected score from the goal line at the most recent week of data collection (Parker et al., 2018). If the median score was below the expected score, a 1 was documented. If the median score was equal to or greater than the expected value, a 0 was documented. Finally, a trend-line rule was applied by estimating an ordinary least-squares trend line through all available observed scores for a given measure. The slope of that line was compared with the slope of the goal line for each measure for each case. If the slope of the trend line was equal to or greater than the slope of the goal line, a 0 was documented. If the slope of the trend line was less than the slope of the goal line, a 1 was documented.
Analysis
For each hypothetical case, the true status and the recommendation from a given decision rule were available. Therefore, the number of true positive (TP), false positive (FP), true negative (TN), and false negative (FN) cases for each assessment at each week of progress monitoring was calculated. In this study, a TP case is one in which true growth was less than the slope of the goal line and the decision rule suggested that a change was necessary. An FP is an instance where the true growth parameter was greater than the slope of the goal line, but the decision rule recommendation suggested a change was necessary. A TN case reflects a scenario in which the true growth exceeded the goal line and the decision rule recommendation suggested that a change was not necessary. Finally, an FN outcome meant that the true growth parameter was less than the slope of the goal line, but the decision rule recommendation suggested that a change was not necessary. The number of TP, FP, TN, and FNs at each week for each measure were used to estimate the base rate of nonresponse in a given sample (TP + FN/TP + TN + FP + FN) as well as the sensitivity and specificity of decision rule recommendations from each measure, as a function of progress monitoring duration. Here, sensitivity captures the probability that a rule would suggest a change is necessary, among cases that were truly not responding to instruction (TP/TP + FN). Whereas, specificity (TN/TN + FP) captures the probability that a rule would recommend that a change is not necessary among cases that truly were demonstrating adequate response to instruction.
Results
Descriptive Statistics
Descriptive statistics for weekly rates of improvement based on estimating separate ordinary least-squares regression lines to scores to outcomes from each student are summarized in Table 1.
Descriptive Results for Progress Monitoring Outcomes (N = 579).
Note. NWF = nonsense word fluency (correct letter sequences per minute outcome); CBM-R = curriculum-based measurement of reading (words read correct per minute outcome); SEE = standard error of the estimate; SEb = Standard error of the slope.
On average, seven weekly NWF observations (SD = 1.49) and 11 weekly CBM-R observations (SD = 4.76) were collected for each student. The average initial NWF score (i.e., intercept) was approximately 45.13 CLS per minute (SD = 17.96), with an average rate of improvement (i.e., slope) of 1.59 CLS per minute per week (SD = 2.55). The typical variability of CLS around the line of best fit (i.e., standard error of the estimate [SEE]) was 6.08 CLS (SD = 3.23). The corresponding value describing the stability of the trend line (i.e., standard error of the slope [SEb]) was 1.05 CLS per minute per week (SD = 0.68). For CBM-R, average intercept and slope values were 24.50 WRCM (SD = 15.22) and 1.66 WRCM per week (SD = 1.45). The average SEE for CBM-R was 5.50 WRCM (SD = 2.56) and average SEb was equal to 0.58 WRCM (SD = 0.51).
Pearson’s product–moment correlations calculated between all intercept and slope estimates were statistically significant at the p < .05 level. Positive, moderate relationships were identified between NWF and CBM-R intercept (r = .324, p < .001) and NWF and CBM-R slope (r = .313, p < .001). The relationship between NWF intercept and NWF slope was negative and moderate (r = –.367, p < .001). The correlations between NWF slope and CBM-R intercept (r = .115, p = .006), NWF intercept and CBM-R slope (r = .194, p < .001), and CBM-R intercept and CBM-R slope (r = .100, p = .016) were positive and weak.
Simulation Verification
Results of multilevel modeling analyses are presented in Table 2. To assess the adequacy of the data generation process, a series of multivariate linear mixed-effects regression models were estimated, using observed simulated CBM-R and NWF scores. Separate models were estimated for each progress monitoring duration. Fixed effects for each duration condition were within 1 SE of the original model used to identify conditions for the simulation. Random effect SDs were within +/– .10 of original parameters across models. Correlations between random effects were also within +/– .05 of original values.
Outcomes of Multivariate Multilevel Model Analysis to Identify Parameters for Simulations.
Note. NWF = nonsense word fluency (correct letter sequences per minute outcome); CBM-R = curriculum-based measurement of reading (words read correct per minute outcome); AIC = Akaike information criterion; BIC = Bayesian information criterion.
Classification Accuracy Outcomes for Decision Rule Recommendations Across Measures.
Note. Each row represents 1,000 unique cases. BR = base rate; Sn = sensitivity; Sp = specificity.
Data-Point Rule
Considering the data-point rule, sensitivity tended to be lower than specificity across all durations for NWF and CBM-R (Table 3). The discrepancy was much larger for NWF, where the average difference between sensitivity and specificity was equal to –.48 (SD = .13), but the discrepancy tended to shrink as duration increased and sensitivity increased. For instance, sensitivity and specificity were equal to .26 and .91, respectively, at 5 weeks, whereas sensitivity and specificity were equal to .65 and .98, respectively, at 15 weeks for NWF. A similar, albeit less severe discrepancy in sensitivity and specificity was observed for CBM-R. In contrast to NWF, sensitivity was almost equal to specificity from 12 weeks onward (.80 vs. .82, respectively). Overall, the data-point rule was much more favorable when used in conjunction with CBM-R relative to NWF with sensitivity and specificity exceeding .60 after 8 weeks of data collection for CBM-R compared with .42 and .95 for NWF across the same span.
Median Rule
The median rule consistently outperformed the data-point rule for sensitivity for CBM-R and NWF across durations (Table 3). Specificity tended to be lower across durations for the median rule relative to the data-point rule for NWF and CBM-R. However, in contrast to the data-point rule, the discrepancy between sensitivity and specificity was not as severe across data collection durations for NWF and CBM-R. The average difference in sensitivity and specificity for NWF for the median rule was equal to –.13 (SD = .05) and .26 (SD = .04) for CBM-R, where sensitivity tended to be greater than specificity across durations.
Trend-Line Rule
The trend-line rule yielded the best balance between sensitivity and specificity across durations for NWF (M = –.08, SD = .05) and CBM-R (M = .23, SD = .04), compared with the median and data-point rule (Table 3). For CBM-R, sensitivity tended to exceed specificity and both outcomes tended to increase as duration increased. For example, sensitivity and specificity were equal to .79 and .51, respectively, when 5 weeks of data were available, compared with .96 and .76 when 14 weeks of data were available. Although sensitivity was poor when few observations were available (e.g., .53 at 5 weeks) it increased dramatically shortly thereafter, exceeding .70 at 7 weeks and .80 by 11 weeks. Although sensitivity tended to be higher for CBM-R across durations for the trend-line rule (M = –.12, SD = .04), specificity favored NWF (M = .19, SD = .04).
Discussion
The purpose of this study was to explore the accuracy of progress monitoring decision rules to evaluate student response to instruction from two CBMs, NWF and CBM-R. Previous research suggested that NWF and CBM-R capture unique growth trajectories when students are receiving supports to promote oral reading fluency and decoding (Van Norman et al., 2018). We explored the accuracy of recommendations from three specific rules: data point, median, and trend line applied to concurrently collected NWF and CBM-R progress monitoring data among students receiving structured Tier-2 text reading and decoding interventions. Data were simulated to follow a schedule where one CBM-R and one NWF observation were concurrently collected per week. Previous CBM-R research has suggested a minimum threshold of .70 for sensitivity and specificity to inform low-stakes decisions, such as changing an intervention or changing day-to-day instruction, whereas sensitivity and specificity equal to .90 may be necessary to inform high-stakes decisions such as special education eligibility (Christ et al., 2018). To better contextualize results and address the stated purpose of the study, we compared the requisite number of weeks data had to be collected and the type of rule used to support each type of decision across measures.
Considering NWF outcomes, the trend-line rule required the fewest number of observations (seven) to achieve sensitivity and specificity values that simultaneously achieved .70. The median rule required 10 weeks of data collection to reach that same threshold, whereas the data-point rule did not yield acceptable outcomes across all progress monitoring durations as sensitivity never exceeded .70, even after 16 weeks of data collection (.66). Decision rules used in conjunction with CBM-R outcomes yielded a slightly different pattern of results. First, the data-point rule was viable to use after 10 weeks of data collection, whereas the trend-line rule was supported between 9 and 10 weeks of data collection. The median rule was not supported for low-stakes decisions until nearly 14 to 15 weeks of data collection. High-stakes decisions were only feasible for NWF when using the trend-line rule after at least 15 observations.
Beyond finding conditions in which sensitivity and specificity were both optimized, the tendency for specificity to be greater than sensitivity for NWF outcomes, and the opposite pattern being observed for CBM-R, has implications for types of errors one is prone to make when evaluating progress monitoring outcomes. When evaluating student growth through NWF, decision rule recommendations tended to yield more accurate recommendations regarding whether an intervention should be maintained, compared with whether a change was necessary (i.e., higher specificity). In circumstances where sensitivity was low, such as with the data-point rule, this suggests that recommendations were prone to recommend that an intervention be maintained. This was less of a concern for the median and trend-line rule after approximately 8 or more weeks of data were available.
Sensitivity tended to exceed specificity for CBM-R across most durations for the median and trend-line rule. This suggests that, when considering growth in oral reading, recommendations were more likely to suggest that a change was warranted, compared with the recommendation that the intervention should be maintained. However, similar to outcomes from NWF, the disparity between sensitivity and specificity was minimal across most durations for the trend-line rule compared. In turn, the trend-line rule may be optimal relative to the median rule and data-point rule to evaluate the progress of students receiving an intervention targeting oral reading fluency and decoding. Building upon that point, it may behoove educators to wait at a minimum of 7 to 9 weeks before making drastic changes if a student is receiving oral reading fluency and phonics intervention.
It should be noted that the base rates of students not showing adequate growth on CBM-R was nearly 90% across progress monitoring durations. In turn, the outcomes of the data-point rule may appear to perform better here with CBM-R than in previous investigations because true growth was highly discrepant (i.e., less than) the goal line, which is one of the few scenarios in which the data-point rule yields acceptable results (Van Norman & Christ, 2016). When true growth tends to be closer to the goal line, even marginally more than the base rates observed in this study, the data-point rule tends to yield unacceptable outcomes even after 4 months, or 16 weeks of data collection (e.g., Hintze et al., 2018).
Although not an original focus of this study, the discrepancy in base rates between NWF and CBM-R has broader implications for progress monitoring practices. Only half the hypothetical students demonstrated proficient growth on NWF. This finding suggests that there may be an important gap in skills measured by the present NWF measure that overestimates mastery of decoding skills. Previous research shows that modeling growth on decoding increasingly complex word types, beyond CVC, better approximates improvement measured through WRCM on CBM-R (Van Norman et al., 2018). It seems then that measures of NWF that include more complex word types, such as DIBELS 8 (University of Oregon, 2021), may provide a finer gradient to measure transition to reading connected text. Alternatively, measures of word identification fluency may be a more useful outcome to model as students begin to receive more advanced decoding and eventual oral reading fluency interventions (Fuchs & Fuchs, 2017). Previous research supports the notion that measures of word identification fluency provide a more useful index of reading connected text among Grade 1 students, compared with measures of NWF (Clemens et al., 2014).
Implications for Practice
The results of this study suggest that one should carefully consider what measure to use to monitor student progress. This point is not novel (e.g., Shapiro, 2011) but is worth reiterating. The choice of assessment and progress monitoring procedure should be driven by the immediate goal and context of the intervention, or interventions, being delivered. If, for instance, a teacher is working with a student who is just beginning to learn to decode unfamiliar words and the preponderance of intervention is dedicated to blending and segmenting words, only using NWF outcomes to monitor progress is likely appropriate. If, however, the student is receiving decoding intervention because they have been struggling to read connected text, and the primary purpose of the intervention is to catch a student up to speed as soon as possible, CBM-R data also may be worthwhile to collect, particularly if connected text reading supports are simultaneously being provided during the intervention. Yet the extremely high base rates of non-proficiency suggest that it may be best to not act upon WRCM scores to drive intervention decisions until after students have exceeded a future target score on CLS consistently. Alternatively, transitioning to a measure of word identification fluency to monitor response to decoding intervention after students have shown adequate performance with simple CVC words may be advisable. The scope of word-types assessed through NWF for most vendors (perhaps except the newest version of DIBELS) may be too narrow to use to measure student response to phonics instruction across an entire semester among students beginning to read connected text in their classroom. School leaders should look at the entire battery of CBMs available for educators to use, as well as the nature of skills they assess, before making a large purchase, to ensure that educators are equipped with the necessary tools to accurately chart student progression toward acquiring more advanced literacy skills.
Limitations and Future Directions
The results of this study need to be interpreted considering potential limitations. First, NWF and CBM-R assessments from one vendor were evaluated in this study. If the study were replicated with other assessments from different vendors, each possessing different psychometric properties, different results may have been observed. Second, one type of goal was evaluated in this study, namely, the required rate of weekly improvement to reach a future target. Other types of goals can be used to evaluate student response to instruction. For instance, normative goals can be used to gauge the students’ rate of progress, compared with other students (Shapiro & Guard, 2014). Future studies should explore the influence of different goal-setting procedures’ accuracy of decision rule recommendations. In addition, only one type of decision rule was used with a single outcome in conjunction with the goal line. It may be that a hybrid approach could be used in which one rule is used to inform “rule-out” decisions (e.g., data-point rule) as a first step and the second step could be to determine whether a change is needed (e.g., the trend-line rule). If a mixed pattern of results is observed, the intervention may continue. Similar two-step approaches have been used, where more or less ambitious goal lines serve as referents to improve the accuracy of CBM-R decision rule recommendations (e.g., Van Norman, 2021). Third, a convenience sample of first-grade students receiving structured literacy interventions was used to identify study conditions. Although the program served a wide geographic region, the results here can best be conceptualized as the accuracy of recommendations from decision rules in conjunction with Tier-2 interventions. Different results may be observed when using the said decision rules in conjunction with students receiving more intensive interventions, such as Tier-3 supports or special education programming. Relatedly, it is important to recognize that categorical recommendations from the decision rules in this study are not the only course of action educators may take. There are many ways to intensify existing interventions rather than completing changing an intervention program (Fuchs et al., 2017). In addition, if students persistently do not show improvement, follow-up diagnostic assessment may be warranted to better target the academic supports, consistent with data-based individualization (Lemons et al., 2014). The results of this study may be viewed as the first pass of evaluating response to instruction among students at risk of later difficulties as opposed to students receiving ongoing specially designed instruction.
Although simulations exert a high degree of internal validity, it often comes at the expense of external validity. In this study, several assumptions had to be made to simulate outcomes. Linear growth was assumed on CBM-R and NWF outcomes across the semester. In addition, error was assumed to be homogeneous for each measure. Future studies should investigate the implications of alternate conditions for simulations on the results observed here. Nevertheless, this initial investigation can be used as a launching point for those subsequent studies.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The research reported here was supported in part by the Institute of Education Sciences, U.S. Department of Education, through Grant R305A210027 to the University of Wisconsin-Madison. The opinions expressed are those of the authors and do not represent views of the Institute or the U.S. Department of Education.
