Abstract
Curriculum-based measurement of oral reading (CBM-R) is used to index the level and rate of student growth across the academic year. The method is frequently used to set student goals and monitor student progress. This study examined the diagnostic accuracy and quality of growth estimates derived from pre–post measurement using CBM-R data. A linear mixed effects regression model was used to simulate progress-monitoring data for multiple levels of progress-monitoring duration (6, 8, 10, . . ., 20 weeks) and data set quality, which was operationalized as residual/error in the model (σε= 5, 10, 15, and 20). Results indicate that the duration of instruction, quality of data, and method used to estimate growth influenced the reliability and precision of estimated growth rates, in addition to the diagnostic accuracy. Pre–post methods to derive CBM-R growth estimates are likely to require 14 or more weeks of instruction between pre–post occasions. Implications and future directions are discussed.
Keywords
Evaluation of Progress-Monitoring Outcomes and Trend Lines
CBM was developed as an approach for assessing the effectiveness of interventions related to special education, enabling teachers to formatively evaluate their instruction. CBM specifically aims to index the level and trend of student achievement within the basic skill areas of reading, mathematics, written expression, and spelling (Deno, 1985, 1986, 2003; Deno & Mirkin, 1977). Thirty years of research has since contributed to developments that have established CBM as a set of procedures that are uniquely designed to improve student achievement when coupled with instructional modifications (Stecker, Fuchs, & Fuchs, 2005) within problem solving (Shinn, 2008) or response to instruction (RtI) frameworks (D. Fuchs, Fuchs, McMaster, & Al Otaiba, 2003). Within a RtI framework, these procedures function to screen, benchmark, and monitor achievement and instructional effects across the primary grades (Deno, 2003).
Specific to CBM-R, robust theoretical (L. S. Fuchs, Fuchs, Hosp, & Jenkins, 2001; Shinn, Good, Knutson, Tilly, & Collins, 1992), empirical (L. S. Fuchs et al., 2001; Stecker et al., 2005; Wayman, Wallace, Wiley, Ticha, & Espin, 2007), and psychometric evidence (Wayman et al., 2007) exists. Within the problem-solving model, educators must identify the most relevant evidenced-based practices (viz., instructional manipulation), and observe and evaluate the effect through progress monitoring (Deno, 2002). In this sense, progress monitoring is the application of single-case design principles (e.g., Kazdin, 1982) to evaluate the effects of educational programming on student progress. Deno (2002) concluded that “such an approach might be termed ‘formative’ in contrast to diagnostic prescriptive” (p. 360).
Research provides modest support for progress-monitoring practices. In general, the research literature indicates that progress monitoring improves student outcomes if time series data are collected, graphically depicted, and evaluated with predefined rules for decision making (Deno, Fuchs, Marston, & Shin, 2001; L. S. Fuchs & Fuchs, 1986; L. S. Fuchs, Fuchs, & Hamlett, 1989; L. S. Fuchs, Fuchs, Hamlett, & Whinnery, 1991; Stecker et al., 2005); however, sparse evidence is available on the technical adequacy of prescribed decision rules associated with progress monitoring (Ardoin, Christ, Morena, Cormier, & Klingbeil, in press). Moreover, data derived from recently published research indicate that the results from progress-monitoring data might be less stable than previously anticipated (Christ & Ardoin, 2009; Francis et al., 2008; Jenkins, Zumeta, Dupree, & Johnson, 2005; Poncy, Skinner, & Axtell, 2005). That is, the variability of student performances across time points is often relatively large in magnitude. As a result, the time series data set is difficult to interpret, yields instable growth estimates, and negatively influences decision accuracy. Given such ambiguity and instability, interim and pre–post methods of progress monitoring should be considered and evaluated as potential alternatives (Jenkins, Graff, & Miglioretti, 2009; Jenkins & Terjeson, 2011). It is the purpose of this study to evaluate pre–post assessment schedules across a variety of progress-monitoring conditions.
Progress-Monitoring Schedules
It might seem odd to consider a pre–post CBM-R assessment schedule as an alternative to weekly progress monitoring (e.g., collecting CBM-R data at the beginning and end in contrast to weekly data collection). Ongoing continuous progress monitoring was a fundamental tenet associated with CBM-R from its inception (Deno, 1985; Deno, Marston, & Tindal, 1985; L. S. Fuchs, Deno, & Mirkin, 1984). Indeed, most CBM-R progress-monitoring research to date focused on continuous schedules of data collection whereby one or more CBMs-R were collected weekly. Researchers have rarely attended to the burden of continuous data collection or the potential alternatives, such as pre–post assessment schedules to evaluate instructional effects.
Early research suggests that progress monitoring (i.e., organization, directions, administration, scoring, graphing, and analysis) takes approximately 2.5 min per child to collect one CBM-R per week, which is approximately 10% of instructional time (Wesson et al., 1988). Time can add up quickly if more than one CBM-R are collected weekly. Researchers recently reflected on the burden of CBM-R data collection for progress monitoring. For example, in their review of the literature, Stecker et al. (2005) identified multiple studies where researchers failed to achieve high integrity for weekly progress monitoring among participating teachers. In response, researchers explored options to make CBM-R progress monitoring more efficient (Jenkins et al., 2009; Jenkins & Terjeson, 2011). The objective of such research is to “lighten the assessment burden for teachers and still provide accurate growth information” (Jenkins et al., 2009, p. 159). Jenkins and colleagues recommended the use of a monthly data collection schedule using multiple passages on each occasion.
Very little is known about alternate schedules of CBM-R progress monitoring. Most research institutes the standard weekly or twice weekly assessment schedule. The results of recent research suggest that the duration of instruction might be more influential than the schedule and number of CBMs-R collected (Christ, Zopluoglu, Monaghen, & Van Norman, in press). Christ et al. evaluated a variety of data collection schedules, which ranged from very dense (one CBM-R per day) to very sparse (three CBMs-R per month). Their results indicate that intervention effects need time to substantiate (Christ, Zopluoglu, Long, & Monaghen, 2012). More data over a brief period of time did little to improve the precision of progress-monitoring outcomes. Such results provide impetus to evaluate less dense and burdensome schedules of progress monitoring.
Christ et al. (2012) used regression to evaluate the relative impact of the progress-monitoring duration, quality of the data set (i.e., residual/error), and the density of the data collection schedule (i.e., CBMs-R per week). Results indicate that the 67% of the variance in the precision of growth estimates was associated with the progress-monitoring durations, which spanned 2 to 20 weeks. Only 22% was associated with the quality of the data set (i.e., good versus very good quality data set) and 4% was associated with the schedule. Further analysis indicated that the influence of duration was most substantial over the first 5 weeks of data collection and was less influential thereafter. Analyses of durations in the range of 6 to 20 weeks indicate that 44% of the precision of growth estimates was associated with duration. An additional 41% was associated with residual, and 10% was associated with schedule. Such results indicate that the duration of instruction is substantially more important than the frequency of data collection when estimating growth. It may be that pre–post assessment methods might be less burdensome and result in equitable growth estimates using CBM-R.
Influential Progress-Monitoring Variables
Student performances on CBMs-R are influenced by a variety of factors. The students’ performance and teachers’ use of data influence the quality of progress-monitoring outcomes. This study will examine various durations of instruction, quality of the data set (residual), and two methods to analyze pre–post CBM-R assessment data.
Duration of instruction
CBM-R progress-monitoring data are typically collected weekly; however, Christ et al. (2012) observed that the duration of instruction and progress monitoring influenced the precision of assessment outcomes. The pre–post assessment schedule used in this study will extend the work of that prior study by fixing the number of CBM-R observations to six (i.e., three pre and three post) and systematically manipulating the duration of instruction from 2 to 20 weeks.
Quality of the data set
CBM-R is sensitive to a variety of characteristics (Colon & Kranzler, 2006; Derr-Minneci & Shapiro, 1992; Derr & Shapiro, 1989) but most notably characteristics of the instrumentation or passage set. Passage sets of inconsistent difficulty are likely to yield imprecise estimates of weekly growth (Ardoin & Christ, 2009; Christ, 2006; Hintze & Christ, 2004). Recent research indicates that performances might fluctuate by as much as 40 words read correct per minute (WRCM) when administered alternate grade-level passages, and the expected deviations from a student’s mean (or median)-level performances across repeated administrations might approximate ±10 to 15 WRCM (Christ & Ardoin, 2009; Francis et al., 2008; Jenkins et al., 2005; Poncy et al., 2005). As described by Jenkins et al. (2009), “even concurrent measures of [CBM-R] vary from passage to passage, challenging the assumption that passages are equivalent” (p. 245). That inconsistency of passage difficulty—and/or other limitations to standardization of assessment conditions—undermines the efficiency of progress monitoring and the utility of outcomes to guide educational decisions.
Mean and median methods
Curriculum-based approaches to assessment have long relied on the use of a median value from across three administrations (Shinn, 1989; White & Haring, 1980). Presumably, the median method functions to limit the detrimental impact of extreme scores. In a limited set of three observations on each of only two occasions, both the median and mean procedures were used to evaluate the impact of pre–post growth estimates.
Purpose
This study examined a pre–post CBM-R assessment schedule as an alternative to ongoing progress monitoring. Pre–post assessment might provide equitable outcomes with fewer resources allocated to assessment. That is, instead of weekly data collection for a period of 14 weeks, it might be feasible to collect data in Week 1 and then again in Week 14.
The first part of this study examined the quality of data gathered by evaluating the relationship between observed growth and true growth across a variety of pre–post assessment conditions. The conditions included the duration of time elapsed between pre–post assessment occasions, residual (error) associated with CBM-R, and methods to estimate observed growth. Both standard correlations and common estimates of precision were examined to evaluate the quality of growth estimates. Correlations yielded estimates of the rank-order consistency between observed/estimated growth and true growth. Reliability calculations yielded another estimate of rank-order consistency. Finally, the magnitude of correspondence between true and observed growth estimates was examined with mean error (ME; aka, mean residual) and root mean squared error (RMSE). The second part of this study evaluated the diagnostic accuracy of decisions made across a variety of pre–post assessment conditions. The conditions were the same as those of the first part of the study. Areas under the curve (AUCs) were computed to determine the diagnostic accuracy of the data gathered under the different assessment conditions. In addition, decision thresholds for observed growth rates were calculated for 50th, 20th, and 15th percentile of performance. Decision thresholds refer to the observed growth rates that best classify student performance at or below the given percentile. The outcomes of this study will provide empirically based recommendations for the method of calculating growth, optimal time between pre–post assessments, and requisite quality of progress-monitoring data sets to yield useful results. Decision thresholds for the different percentiles of performance within different assessment conditions will be offered as well.
Method
Statistical Model for Data Generation
Interpretation and use of progress-monitoring data requires longitudinal modeling. A modern approach to modeling longitudinal data is provided by linear mixed effects regression (LMER; see, for example, Verbeke & Molenberghs, 2000). The LMER model was used to simulate responses of hypothetical individuals over time. Thus, LMER was considered to be the true model that generated the observed growth curve data for each participant. The LMER model for the simulation was as follows:
In this model, Y ijk is the kth (k = 1, 2, 3) observed level of performance for individual i (i = 1, …, N) at time point j (j = 1, …, n i ). Equation (1) is slightly different from the usual LMER model and includes an additional subscript k at the error term. This modification is chosen to be able to model the successive observations of individual i at time point j. “(β0 + b0i) + (β1 + b1i) × Time ij ” represents the true score of individual i at time point j, and the observed scores of individual i differ by adding a random error component for the kth observation at time point j. If we drop the k subscript from the model, (1) only allows to model one observation per time point, and the observed score of individual i differs from the true score by adding a single random error component for time point j.
In our simulation, there were no missing data, so n i = n. The parameters β0 and β1 are group-level regression coefficients known as fixed effects. β0 is the fixed effect for the intercept, and β1 is the fixed effect for the slope. The terms b0i and b1i are random effects representing the individual deviations from the fixed effects of the group. b0i is the deviation of ith individual’s intercept from the group intercept, and b1i is the deviation of ith individual’s slope from the group slope. It is the mix of fixed and random effects that is the genesis for the term mixed effects regression. Finally, ϵ ijk is the residual—or random error—associated with the kth observation for the ith individual at the jth time point.
There were common assumptions used in conjunction with the LMER model to generate data (Fitzmaurice, Laird, & Ware, 2004; Verbeke & Molenberghs, 2000). The random effects have a joint-normal distribution with means of 0 and variance–covariance matrix G. The variances of the random intercepts and slopes are denoted as Var(b0i) and Var(b1i), respectively, and the covariance between them as Cov(b0i, b1i). The random errors are assumed to have a multivariate normal distribution with means of 0 and variance–covariance matrix of R with the constant error variances (
Equation (1) was used to generate CBM-R data for schedules with three CBMs-R per time point for an individual. The random errors associated with the successive k observations at time point j for person i were assumed to be correlated, and the random errors associated with the observations at different time points are assumed to be independent. See Table 1 for the corresponding error variance-covariance matrix used in the current study. The assumption that errors are uncorrelated across time points is well founded in psychometric theory. The assumption that errors are correlated within time points is founded within research and theory. That is, there are characteristics of the examiner, examinee, and setting that influence CBM-R outcomes but are difficult to specify. In general, it might be expected that examiner and examinee characteristics such as motivation, interest, mood, alertness, excitement, proximal prior exposure to similar tasks (e.g., reading instruction), and the like may influence outcomes. A student might have a good day or bad day. The examiner’s presentation of directions, stimuli, scoring, and timing procedures might vary systematically across days, which might influence CBM-R performance. Finally, the characteristics of the setting are likely to fluctuate. CBM-R is typically administered within a classroom, hallway, library, or some similar setting. Progress-monitoring assessments are rarely administered routinely within a standardized testing room, so administrations and performances are subject to the variations of those less controlled environments. The characteristics of any one of those settings are likely to fluctuate across occasions so that it is quiet with limited distractions 1 day and less so on another day.
Error Variance–Covariance Matrix for (1) With k Subscript.
Note.
To simulate CBM-R progress-monitoring outcomes, values for the parameters in Equation (1) were specified. The parameters used in the simulations were based on the values derived from a large empirical data set (described below). Generating data based on known parameters allows for sampling from specific populations and the assessment of the effect of sampling vagaries. Repeated measures of hypothetical individuals were generated based on assumptions as described above with known values for the following parameters: β0, β1, Var(b0i), Var(b1i), Cov(b0i, b1i),
Parameter Values for Data Generation
The parameter values for the simulation were derived from empirical data and the expert judgment of the researchers. The population parameters were derived by fitting LMER models to a large progress-monitoring data set of second- (n = 1,517) and third-grade students (n = 1,561) with a demographic distribution of approximately 46% female, 2% special education, 53% White, 17% Black, 8% Hispanic/Latino, 6% Asian/Pacific Islander, and 2% American Indian/Alaska Native across samples. These data were collected in the Midwest through a federally funded project designed to provide supplemental (Tier 2) reading interventions to elementary students at risk of reading difficulties. All data collectors were trained to criterion with AIMSweb training materials and assessed for administration fidelity using the Accuracy of Implementation Rating Scales, (Shinn & Shinn, 2002). Interrater reliability data were not available, but reported estimates typically approximate or exceed .95.
Based on the parameter estimates derived from the empirical data (Table 2), it was determined that the average number of WRC per minute at the beginning of the progress-monitoring period across the two-grade levels approximated 40 WRC per minute with a standard deviation (SD) of 12.2 WRC per minute (var = 150). Therefore, model parameters were set with a fixed effect group intercept, β0 = 40, and the variance for random intercepts, Var(b0i) = 150. These were constants in the model across the experimental conditions for simulation. Likewise, the average weekly slope estimate approximated 1.5 WRC per minute with a SD of 0.63 (var = 0.40). Given those estimates, the fixed effect for group slope, β1 = 1.5, and the variance of the random slopes, Var(b1i) = 0.40, were used in the model across conditions for simulations. The correlation between the random effects (ρ01) was fixed to .20. The correlation between errors associated with the k successive observations at time point j for individual i (ρ) was approximated to .30. These estimates approximate those from a large extant data set where 500 students were administered 65 passages within 10 days with as many as 10 passages administered on each occasion (Christ, Ardoin, Eckert, & White, 2010a, 2010b).
Parameter Estimates for Progress-Monitoring Data Using Linear Mixed Effect Regression.
Note. n = number of progress-monitoring cases per data set and simulation condition. Parameters estimates derived from large data sets of progress-monitoring data.
300 batches of 30 iterations each for a total of 9,000 simulated cases per condition.
Independent Variables
There were three independent variables, the residual variance,
The number of weeks between measurement was set to one of eight levels (n = 6, 8, 10, 12, 14, 16, 18, 20). Two pre–post assessment methods were used to estimate individual growth from simulated CBM-R data sets. These were pre–post mean (PP-mean 1 ) and pre–post median (PP-mdn 2 ). Both methods appear in the research and professional literature, and they are also the most commonly used in the practice.
Dependent Variables
Slope accuracy
Four dependent variables were selected to evaluate the correspondence between observed growth estimates and true growth along with reliability, bias, and precision. These were (a) validity or the correlation between estimated and true slopes, (b) the reliability or the squared correlation between estimated and true slopes (Crocker & Algina, 1986; Nunnally, 1970), (c) the ME, and (d) the RMSE between true and observed slopes.
The true slope for the ith hypothetical subject is defined as
Finally, the ME 3 and the RMSE 4 were calculated for estimated and true slopes. ME was an estimate of bias and RMSE was an estimate of precision. In both cases, values that deviate substantially from zero are indicative of observed estimates that have considerable unreliability in indexing true growth. RMSE is a frequently used measure in simulation studies to index differences between predicted and true values.
If the estimated slope consistently overestimates the true slope, then ME will tend to be negative. If the estimated slope consistently underestimates the true slope, then ME will tend to be positive.
Diagnostic accuracy
Several descriptive measures were used to evaluate the diagnostic accuracy of decisions based on estimated slopes. All of these measures were used to assess how well the ordinary least squares (OLS) estimated slopes predicted categorization based on true slopes. Those measures included sensitivity (or true positive predictive power [TPP]) and specificity (true negative predictive power [NPP]), positive predictive power (PPP) and negative predictive power (NPP), AUC, and phi coefficient (Silberglitt & Hintze, 2005; Swets, 1988; Swets, Dawes, & Monahan, 2000). The criteria applied to interpret AUCs are variable, values are considered excellent, good, fair, or poor within ranges of 0.90 to 1.0, 0.80 to 0.89, and 0.70 to 0.79, respectively, and are generally consistent with National Center for Response to Intervention standards (NCRtI, 2009). In addition, an optimum decision threshold was selected to balance the sensitivity and specificity. As recommended in the literature, the point of the receiver operating characteristic (ROC) curve where the slope is equal to 1.0 is described as S-optimal (Soptimal; Swets et al., 2000). That point coincides with a neutral decision threshold to balance the various criteria that define decision accuracy. Distinct from decision thresholds, three cut-off points for true slopes were identified based on research and national standards (NCRtI, 2009): 50th (1.5 WRCM), 20th (0.96 WRCM), and 15th percentiles (0.84 WRCM). A more detailed discussion of cut points can be found within Christ et al. (in press).
Simulation Design
There were 64 distinct conditions to simulate and evaluate for residual (4 levels) × number of weeks between measurement (8 levels) × evaluation method (2 levels). For each experimental condition, repeated measures for 300 batches were generated based on the LMER model and the given parameters previously described. Each batch contained 30 examinees to approximate a typical size of a school-based class and a common sample size to yield stable mean estimates. In all, there were 9,000 hypothetical students in 300 batches for each experimental condition. After generating repeated measures and saving the true individual slopes used to generate data for the hypothetical examinees, the estimated individual slopes were calculated with each of the two methods. Then, dependent variables as defined above were calculated for each condition.
Analysis
Descriptive statistics are presented within each experimental condition for validity, reliability, precision, and diagnostic accuracy. The independent variables of duration (D) and residual (R) were entered into regression models as predictors of the precision of growth estimates (RMSE). The dependent variable in the analysis was the log transformation of RMSE to accommodate the skewed in the original data set and achieve normality. These transformations were valid representations of the RMSE in the regression model because the variables used to compute RMSE were normally distributed and the log-transformed values of RMSE were approximately normal (Harwell et al., 1996). The same analysis was run separately for the estimation methods used in the study, PP-mean and PP-mdn.
The full model included the two main effects (duration and residual) and the interaction effect (duration × residual). After fitting the full model, the reduced model was fitted by excluding the nonsignificant effects. Finally, the independent contribution of each significant predictor explaining the variance in log RMSEs was examined. Models were examined with all data across conditions for duration, residual, and schedule. As a post hoc procedure, the analysis was run for a restricted set of results that excluded the briefest durations of 2 to 5 weeks. The post hoc analysis was conducted after it observed that duration accounted for an unexpected and disproportionate amount of the variance for RMSE. Simultaneously, it was observed that RMSE was substantially reduced in the first weeks of data collection, because none of the results across studies supported progress-monitoring durations less than 6 weeks; a restricted model seemed relevant and more valid from which to derive conclusions and recommendations.
Results
As a check on the validity of the simulation methods, the estimates of individual growth were compared with the simulation parameters, which corresponded with the true rates of growth. The mean values across conditions were all within ±0.05 of 1.5 WRCM except for the conditions with error of 15 and 20 WRCM and 2-week duration, which are consistent with very poor quality progress-monitoring estimates generated for brief and sparse data sets. This provides evidence that growth estimates were not biased in any but the most extremely low-quality conditions. Figure 1 illustrates the relationship between true and observed growth estimates after 6- and 14-week durations.

Scatter plots of Pre-Post Mean Observed Estimates (Y-axis) and True Slopes (X-axis) for 6- and 14-week durations and residuals of 5, 10, 15 and 20 WRCM.
The conditions with fewer measurements and larger residuals yielded substantially more variability. The observed SDs are comparable with the dispersion parameter value of 0.63 (var = 0.40) WRCM per week. In general, the SD for observed estimates was larger within conditions that were brief in duration or large in residual. For example, an 8-week duration for PP-mdn and PP-mean, within the poor data set condition, yielded SDs that were 300% of the true SD (0.63) or 2.43 and 2.27 WRCM per week, respectively. Similarly, a 20-week duration within the good data set condition yielded SDs of 0.85 and 0.83 (133% of the true SD), respectively.
Validity, Reliability, and Precision
Robust estimates of validity (≥.70) were observed after 10-week durations within only the very good and good data quality conditions (Table 3). Reliability coefficients approximate criterion for low-stakes decisions (≥.70) after 14 weeks within only the very good data quality condition (Table 3). The values for reliability and validity were consistently less robust within PP-mdn conditions as compared with PP-mean conditions. The same was observed for RMSE whereby the values were approximately 7% less precise within the PP-mdn condition (Table 4). For example, the RMSEs for PP-mdn and PP-mean were .62 and .57 after a 10-week duration within the very good set condition.
Validity and Reliability for Estimates of Growth.
Note. PP-mean = pre–post mean method; PP-mdn = pre–post median method. Standard error of estimated values <0.01.
Validity is the correlation of true and observed scores.
Reliability is the squared correlation of true and observed scores.
Root Mean Square Error Between True and Observed Estimates of Growth.
Note. PP-mean = pre–post mean method; PP-mdn = pre–post median method. Standard error of estimated values <.03.
Duration and residual regressed on precision
The log transformation of RMSE values for PP-mean and PP-mdn methods approximated to a normal distribution. The results of the regression analysis for mean and median pre–post methods were almost identical except for the magnitude of the regression coefficient for the intercept term (Table 5). The model overall explained 96.3% of the variance in log RMSEs. Both duration and residual explained a significant amount of variance in log transformation of RMSEs. Within the reduced main effect models, the duration explained about 39% and residual explained about 57% of total variance in log RMSEs.
Summary of Regression Analysis: Duration and Residual Regressed on Precision.
Note. A log transformation of RMSE was used as a dependent variable for the analysis.
p < .001. **p < .01. ***p < .05.
Diagnostic Accuracy
AUCs approximated the NCRtI criterion for good decision accuracy (≥.80) after an 8-week duration within the very good condition and a 14- to 16-week duration within the good condition (Table 6; supplemental table available on request). Those results generalize for decisions associated with the 50th, 20th, and 15th percentiles, which corresponded with true score criterion values of 1.5, 0.96, and 0.84 WRCM per week, respectively (see “Method”). In contrast, optimal decision thresholds were identified for all conditions and are presented in supplemental material (available on request from the first author). The results across PP-mdn and PP-mean were substantially similar, so only PP-mean tables were included (supplemental tables are available from the first author). The optimized decision thresholds for the 50th percentile did not require adjustments, indicating 1.5 WRCM improvement per week as optimal; however, decision thresholds at the 20th and 15th percentiles were increased from 0.96 and 0.84 to approximately 1.1 and 1.0 WRCM improvement per week, respectively. The results generalized across duration, residual, and conditions.
Pre–Post Mean: Estimates of AUC Across Progress-Monitoring Conditions.
Note. AUC = area under the curve. Criterion for AUC of 0.70 to 0.79, 0.80 to 0.89, and 0.90 to 1.0 was poor, good, and excellent; results are generally equivalent to those for pre–post median conditions. Shading corresponds to number of weeks where AUC ≥ 0.80.
The AUC exceeded .80 after at least 8 weeks of progress monitoring between pre–post measurements in the very good quality data set condition. In that condition, the TPP/sensitivity results indicate that 73% to 75% of those cases with deficit true slopes (<cut point) were accurately classified by the observed slope (<optimized decision threshold). The NPP/specificity results indicate that similar percentages were accurately classified. The use of the Soptimal method provided for a balanced decision threshold, which optimized TPP/sensitivity and NPP/specificity. Other corresponding statistics for diagnostic accuracy were all within an acceptable range except for PPP. Specifically, NPP values indicate that 72% to 95% of excess observed slopes (>optimized decision threshold) were correctly classified. PPP values indicate that only 34% to 73% of deficit observed slopes (<optimized decision threshold) were correctly classified.
Discussion
This study evaluated CBM-R pre–post assessment outcome conditions, which include the duration of time elapsed between pre–post assessments (i.e., duration of instruction), the quality of data set (i.e., residual associated with CBM-R), and growth estimate methods (i.e., PP-mean or PP-mdn). Analysis of duration and residual indicates that they accounted for 39% and 57% of the variance in precision, respectively (Table 5). The residual accounted for the majority of the variance in precision; however, there was a substantial effect associated with the duration between pre–post assessment occasions. That observation illustrates that instructional effects require time to substantiate and establish sufficient magnitudes before they are precisely estimated with CBM-R.
The results of this study support some general conclusions: (a) It is necessary to optimize administration conditions and instrumentation to establish a very good or good quality data set because lesser quality data sets are insufficient to guide instructional decisions; (b) growth estimates from CBM-R pre–post data collection requires a minimum duration of 8 to 10 weeks, but longer durations are often necessary; (c) the PP-mean method yielded slightly improved growth estimates when compared with PP-mdn, especially with shorter durations; and (d) decision thresholds should be set more stringently, such that slopes of 1.0 to 1.1 WRCM per week might indicate risk of a deficit RtI if 1.5 or more WRCM per week was expected. These recommendations are addressed in more detail below.
Quality of the Data Set
The potential impact of low-quality CBM-R progress-monitoring data sets on growth estimates and instructional decisions are too rarely addressed in the research literature. The results of this study provide support that it is necessary to minimize the magnitude of residuals in progress-monitoring data sets to guide instructional decisions. It seems residuals less than 10 WRCM might facilitate valid and reliable growth estimates. It is likely—although not certain—that the variability in passage difficulties across alternate forms contributes substantially to the magnitude of the residual. Pre–post assessment methods might provide a solution because the same passages might be used across occasions, which effectively controls for difficulty. Consistent with the recommendations that derive from previous research (Ardoin & Christ, 2008; Jenkins et al., 2009), it seems reasonable to use the same passage set for pre–post assessments. Estimates of delayed test–retest reliability (3 months) with the same passages are excellent (≥.95; Ardoin & Christ, 2008) such that pre–post assessment with the identical passages could approximate the very good data set condition (standard error of the estimate [SEE] ≤ 5 WRCM), whereas there are no other data collection scenarios that seem to approximate the very good condition.
The foregoing discussion will focus on very good and good quality data set conditions with residual of 5 and 10 WRCM, respectively. Data sets with residuals of greater magnitude are not useful for educational decision making.
Duration Between Pre–Post
Jenkins and colleagues (2009) and Jenkins and Terjeson (2011) evaluated and subsequently recommended that the use of intermittent progress-monitoring practices alleviate the burden on educators who seek to evaluate growth. Christ et al. (2012) then observed that the precision of growth estimates was substantially influenced by the duration of progress monitoring and that relationship with duration was more influential than either the residual or schedule of CBM-R data collection. The results of this study provide evidence that duration of instruction does influence the quality of growth estimates independently of the number of CBMs-R. The validity, reliability (see Table 3), precision (Table 4), and diagnostic accuracy (Table 6) of growth estimates improve as the duration of instruction was extended. This is consistent with the results from regression analysis, which indicate that the duration of instruction between pre- and post-occasions substantially influences precision (Table 4).
Validity approximated .70 after 8- and 18-week durations within very good and good quality data set conditions. Reliability approximated .70 after 14 weeks within the very good conditions and never approximated .70 in any other data set quality condition. It seems that 1- to 2-month durations were insufficient. The impact of duration on the quality of progress-monitoring outcomes is also illustrated with estimates of precision, which is the standard error of the slope (Table 4). Precision within the very good and good conditions improved from approximately 1 and 2 WRCM per week to approximately 0.5 and 1.0 as the duration was extended from 6 to12 weeks; therefore, a 68% confidence interval within each data set condition would approximate ±0.5 or 1.0 of the observed growth rate after a 12-week duration.
The results of this study converge with similar previous work that suggests duration is a substantial and influential variable (Christ et al., in press). These results support the conclusion that more data over a brief period of time is inferior to less data over a longer period of time. This study extended on the previous work in that the number of CBMs-R was fixed to three at each pre- and posttest occasion. This isolated the effect for duration unlike previous research that evaluated ongoing progress-monitoring schedule so that longer instructional durations always coincided with more progress-monitoring occasions and, therefore, a larger data set.
Mean and Median Methods
The PP-mean method was superior to the PP-mdn method; however, as the quality of the data set improved and the duration increased, differences between the two methods were marginal, and trivial when the data set was of at least good quality data set (e.g., σ ≤ 10). Growth estimates based on lower quality data sets should not be used to guide educational decisions. These findings are similar to that of Christ and colleagues (2012) and Christ and colleagues (in press), which suggest that the estimation of individual rates of growth can be disturbingly inaccurate when there is a combination of a poor quality data set and a relatively few number of observations.
Decision Threshold
In the current study, diagnostic accuracy for PP-mean method also approximated good decision accuracy after an 8-week lapse between assessment occasions in the very good condition. Decision thresholds were optimized with an Soptimal method when increased from the values of 0.96 (20th percentile) and 0.84 (15th percentile) as recommended by research and national standards, to 1.1 and 1.0, respectively. The decision threshold at the 50th percentile did not require adjustment. At the 50th percentile, this indicates that 1.5 WRCM growth per week is a good indicator of student performance, whereas 1.1 and 1.0 WRCM growth per week are good indicators of students performing at the 20th and 15th percentiles, respectively.
Given the decision threshold adjustments, corresponding statistics for diagnostic accuracy were within an acceptable range except for PPP under those same conditions for the PP-mean method. Specifically, NPP values for a duration of 10 weeks, across all decision thresholds and levels of data set quality, indicate that for cases where the observed slopes predicted deficit true slopes, 57% to 95% of those predictions were accurate. The phi coefficient was in the moderate range; however, as stated, the magnitudes of PPP generally indicated that many of the positive predictions based on observed slopes were inaccurate. For example, for the PP-mean method after a 10-week lapse between pre–post assessment occasions in the good quality data set condition, the PPP values indicate that for cases where the observed slopes predicted a deficit true slope, 28% to 65% of those predictions were accurate. The range was broad with values of 28%, 34%, and 65% for decisions at the 15th, 20th, and 50th percentiles, respectively. The selection of a more conservative decision threshold, such as 0.84 WRCM per week at the 15th percentile, would result in a reduction of sensitivity from 69% to 63% and an improvement in PPP from 28% to 30%. In contrast, increasing the duration between pre–post measurements improves PPP from 28% to as much as 42% with 20 weeks of data and increases sensitivity from 69% to 81%.
Convergence With Previous Findings
The results provide further evidence that even alternative practices to continuous progress monitoring, such as pre–post assessment, should be carefully reviewed and evaluated to ensure that outcomes are of sufficient quality to guide educational decisions. Despite the need for growth estimates, threats and limitations associated with growth estimates are well established in the literature (Cronbach & Furby, 1970), and those threats are especially severe when growth is estimated for the individual student rather than a student group. The effects of an instructional program are more stable when the group response is analyzed. It is much more difficult and tenuous to evaluate growth—and RtI—on the individual student level.
Whereas the PP-mean method provides a more accurate growth estimate against the PP-mdn method, according to analogous indexes obtained in Christ et al. (2012), OLS provided superior growth estimates using 1 CBM-R per week. More specifically, a very good data set with 12 CBMs-R from weekly progress monitoring that was analyzed with OLS yielded growth estimates within ±0.32 WRCM per week of true growth. In that case, the true rate would fall within the range of 1.18 to 1.82 approximately 68% of the time. As the duration increases to 20 elapsed weeks, growth estimates using OLS become increasing more accurate over those obtained using the PP-mean method (±0.19 and ± 0.27, respectively).
In a subsequent evaluation of progress-monitoring practices, Christ et al. (in press) examined a variety of data collection schedules using OLS to estimate growth. The schedules examined included one CBM-R on three and five occasions per week, three CBMs-R on one and two occasions per week, and three observations once a month. Regardless of schedule, durations of less than 4 weeks were not supported. When one CBM-R is collected on five occasions per week or three CBMs-R twice weekly, OLS provides adequate growth estimates after 12 weeks in the very good condition. This is a 2-week improvement over pre–post assessment evaluated in this study and one CBM-R on one occasion per week as evaluated in Christ et al. (2011). For diagnostic accuracy, the AUC approximated criterion for good decision accuracy after 6 weeks in the very good condition when one observation is collected on five occasions per week and when three observations are collected twice weekly. A total of 9 weeks are required for the monthly condition.
Limitations and Future Directions
A large sample of progress-monitoring data across a variety of conditions is difficult to collect, especially when the quality of the data is used as an independent variable condition. Simulation is useful to generate those data that are otherwise prohibitively difficult to collect in practice; however, the validity of the findings and implications depends on the veracity of the model and the parameters. The application of the LMER model has precedence in recent CBM-R research to examine annual growth (Christ et al., 2012) and instrumentation effects (Francis et al., 2008). The judgment of researches was used to specify the particular values used in the simulation. Although the researchers were fairly confident in their judgment, ongoing and future research is necessary to evaluate the degree of convergence across other simulations and fieldwork. In addition, at the time of this study, there were few published studies that applied LMER for large populations. This study relied on parameter estimates from a relatively large data set that emerged from practice (N = 3,078, see Table 2).
Expert judgment was also necessary to define the levels for the number of data points and residual. The very good (σϵ = 5) and very poor (σϵ = 20) quality data sets were optimistic and pessimistic, respectively. Low level of residual may indeed be attained if ongoing research and development define and establish more optimal conditions for assessment, whereas high level of residual is unlikely to occur except in the most extreme—and undesirable—progress-monitoring conditions. If compared with prior research (Ardoin & Christ, 2009), high-quality passage sets, such as Fomative Assessment Instrumentation and Procedures for Reading (FAIP-R) and AIMSweb, might generate good data sets, whereas low quality passage sets, such as Dynamic Indicators of Basic Early Literacy Skills (DIBELS), might be expected to generate poor data sets.
Future research is necessary to evaluate other methods, such as LMER, to estimate individual slopes and conditions associated with progress monitoring (i.e., schedule). A LMER model was not used in this study to generate observed growth estimates because a subset of the dependent variables functioned differently when used to evaluate LMER—so the results were not comparable across additional types of methods.
Concluding Comments
Multiple studies illustrate that the quality of the data set and the duration of progress monitoring influence the quality of growth estimates. It seems that long durations of instruction are necessary to detect instructional effects. It is left to evaluate if and how reliable, valid, and precise growth estimates might be derived more efficiently. The PP-mean method requires less teacher time, less student time, and fewer alternate passage forms. It is also more likely to establish very good quality data sets with residual of 5 WRCM because the same passages can be used at pre- and posttest occasions—so instrumentation issues are minimized. It might also be easier to optimize two administration occasions rather than weekly administrations. Such optimization would ensure that distractions are eliminated, directions are consistent, and the student’s disposition toward the assessment is consistent across occasions. Deviations from those high-quality conditions will result in deteriorated progress-monitoring outcomes.
CBM-R cannot be viewed as an informal or loosely standardized procedure. It is only within optimized or very good assessment conditions that CBM-R might yield valid, reliable, and precise growth estimates. In the case of pre–post assessment, that is likely to take 14 to 16 weeks (3 to 4 months) of instruction between assessment occasions.
Footnotes
Declaration of Conflicting Interests
Dr. Christ receives consulting income from publishing companies, which may commercially benefit from the results of this research. This relationship has been reviewed and managed by the University of Minnesota in accordance with its Conflict of Interest policies. Dr. Christ is also the primary developer and director of the Formative Assessment System for Teachers (FAST;
).
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
