Abstract
As a school-wide framework, Multi-Tiered Systems of Support (MTSS) relies on the prevention and early identification of students at risk of academic failure. Approaches to early identification of students in need of support include the administration of universal screening assessments and the analysis of existing student data such as attendance, grades, office discipline referrals, and prior performance on statewide assessments. However, there is little research that directly compares the accuracy and reliability of these approaches, particularly in middle grades. This investigation provides a direct comparison of curriculum-based measures in reading and the examination of archival data at the middle school level for the identification of students at risk for academic failure. Data were collected for students in Grades 7 (n = 197) and 8 (n = 237). Data were analyzed through hierarchical logistic regression using statewide reading achievement tests as the dependent variable. Results inform how data from universal screening assessments and existing sources can be used to accurately and efficiently identify students in need of academic support.
To provide effective intervention and supplemental services to students who struggle, schools must first identify those in need of support. Early and accurate identification of students at risk of developing persistent skill deficits provides schools the best opportunity to intervene early and put students back on track for success. One approach to identifying such students is through the use of universal screeners. Universal screeners are brief assessments of basic skills used to determine which students need additional supports and services (Hosp, Hosp, & Dole, 2011). Early identification through screening procedures aids in quickly identifying individuals who, without intervention, would “develop serious and chronic academic problems” (D. Fuchs, Fuchs, & Compton, 2013, p. 265).
Many schools administer universal screeners in the form of curriculum-based measures (CBMs). CBMs are robust formative assessments of academic skills, standardized for difficulty within and across grade levels (Hosp et al., 2011). Measures such as oral reading fluency (ORF) provide an estimate of a student’s overall reading proficiency by measuring the number of words read per minute on a grade-level appropriate text. CBM scores are then compared with predetermined benchmark cut-scores. This comparison aids in determining whether a student requires intervention (D. Fuchs et al., 2013).
Screening in Secondary Schools
At the secondary level, universal screening may be used to identify students in need of support and catch academic or behavior problems that may otherwise go unnoticed. Middle and high school students can have considerable skill gaps that fall beyond the scope of prevention and early-stage interventions used in elementary settings. Gaps in skills, such as decoding and fluency, may present problems that affect learning in other subject areas. Such deficits can impede a student’s ability to learn subject area content and function independently with grade-level text.
Fortunately, a number of studies have made inroads to improving effectiveness of screening mechanisms in middle school. Recent research has explored both the technical adequacy of screening measures and novel approaches to screening such as multiple gating procedures and composite scores that use CBM and extant data. In the last decade alone, more than 30 studies have been published that address some aspect of technical adequacy or predictive validity for screening in middle grades. Data from research across 15 different states support the use of CBM as a predictor of outcomes on statewide high-stakes reading tests (Kilgus, Methe, Maggin, & Tomasula, 2014; Yeo, 2009). Findings generally support the use of ORF, Maze, and other CBMs as useful tools in the screening process with correlations to state reading tests ranging from r = .42 to .96 (Baker et al., 2015; Barth et al., 2012; Espin, Wallace, Lembke, Campbell, & Long, 2010; Nelson, Van Norman, & Lackner, 2016; Ticha, Espin, & Wayman, 2009; Yovanoff, Duesbery, Alonzo, & Tindal, 2005).
For students in seventh and eighth grade, Baker et al. (2015) reported that ORF explained 42% to 53% of variance of outcomes on state tests and classification accuracy of .83 to .85 area under the curve (AUC). AUC is the probability of predictors correctly categorizing outcomes such as proficient versus not proficient. Accuracy and reliability of CBM screening may also be strengthened when CBM data are used in concert (e.g., combining ORF data with measures of vocabulary and comprehension). Recent studies have shown improved accuracy of screening results in correctly predicting outcomes on statewide reading achievement tests over use of CBMs (Baker et al., 2015; Nelson et al., 2016). Even amid concerns regarding the appropriateness of some CBMs for secondary students (Baker et al., 2015) and the deceleration of fluency growth as students age (Espin et al., 2010), research into the technical adequacy and predictive validity supports the use of ORF and Maze assessments for students in middle school.
Early Warning Signs (EWSs)
Despite evidence from a number of studies and reviews that support the use of CBMs for students in middle schools (Baker et al., 2015; Barth et al., 2012; Codding, Petscher, & Truckenmiller, 2015; Denton et al., 2011; Yeo, 2009), practitioners do not necessarily view CBM as a satisfactory screening mechanism. Logistical considerations of assessment, time, and the ongoing debate regarding the appropriateness of using basic reading skills such as fluency as a proxy for reading ability in secondary grades (Baker et al., 2015; Espin et al., 2010) have fueled calls for alternative screening measures that may be better fit to the needs of students and teachers in middle and high schools.
One promising alternative is the use of EWSs. EWS uses existing data as an indicator of risk of failure and drop-out. Data from attendance records, grades, state assessment results, and office discipline referrals (ODRs) are consistently good predictors of eventual drop-out (Kennelly & Monrad, 2007) and may be helpful in identifying students who require academic or behavioral supports.
For decades, researchers and practitioners alike have explored the connections between school failure and predictors of failure such as early literacy skills, cognitive ability, and behavior characteristics (Baydar, Brooks-Gunn, & Furstenberg, 1993). Presumably, early identification of students at risk of school drop-out will enable educators to provide timely, effective intervention that will keep students in school and put them back on the path to graduation (Balfanz, 2009; Balfanz, Herzog, & Mac Iver, 2007). As Neild, Balfanz, and Herzog (2007) aptly summarize, “Many students who drop out send strong distress signals for years” (p. 28). Even 5 to 10 days absent per semester are associated with increased rates of drop-out (Heppen & Therriault, 2008).
Neild and Belfanz (2006) found that the number of absences accrued in the first 30 days of high school was one of the strongest predictors of drop-out when compared with other risk factors such as gender, race, age, and scores on standardized tests. In a long-term study with Philadelphia Public Schools, researchers found that among students who eventually dropped out of school, 85% exhibited patterns of absences that developed in sixth grade and increased throughout the remainder of their careers in middle school (Neild et al., 2007).
Similarly, patterns of inappropriate behavior are also associated with learning difficulties and drop-out. Comorbidity of academic and behavior problems is well documented, spanning literature in drop-out prevention, special education, behavior management, health, neuroscience, and others (Barry, Lyman, & Klinger, 2002; McIntosh, Goodman, & Bohanon, 2010; Reinke, Herman, Petras, & Ialongo, 2008; Tobin & Sugai, 1999). In an exploration of drop-out predictors using grades and suspensions, researchers found that a single suspension increased the likelihood that a student would drop out by 77.5% (Suh & Suh, 2007).
Scores on statewide achievement tests are another data source with known relation to future academic success. In an analysis of 10 different measures of reading, Denton and colleagues (2011) reported that of 10 measures, the best predictor of the 2007 Texas reading comprehension accountability test (TAKS) scores was 2006 TAKS scores (AUC = 0.82). In addition, 44% of variance on the TAKS was explained by TAKS scores from the previous year. Reed, Wexler, and Vaughn (2012) even argue that state assessments alone may be an efficient screening tool for predicting which students are currently at risk of not reaching grade-level benchmark and eventual drop-out, particularly when CBM testing is not feasible.
Comparing Screening Methods
The majority of research on CBM as a screening tool has focused on students in elementary grades (Baker et al., 2015; Codding et al., 2015; Denton et al., 2011; Yeo, 2009). Similarly, most research on EWSs has focused on students in Grades 9 through 12, particularly in ninth grade, when the highest rates of drop-out occur (Franklin & Trouard, 2016). This leaves educators in middle schools with comparatively less guidance concerning which methods are most suitable for students in Grades 6 through 8. Even with what is known about the reliability and predictive validity of screening measures, there are often “serious inefficiencies” in implementing screening procedures resulting in “unacceptably high rates of false positives (or students who appear at risk but are not)” (D. Fuchs et al., 2013, pp. 265–266).
However, recent research has begun to explore novel screening methods in middle school. For example, Nelson et al. (2016) examined ways in which schools could integrate existing data sources and CBM data to efficiently screen middle school students for reading and math problems. Using a combination of teacher rating scales and state test scores for students in sixth through eighth grade, researchers examined the classification accuracy of multiple procedures when using each predictor independently and in concert in the form of a composite score. Results indicated that prior state test scores were consistently better at predicting future reading achievement than all other measures, including ORF and teacher rating scales. The combination of state test data and teacher rating scales as a composite score yielded minimal and inconsistent improvements in sensitivity and specificity across each grade level. Although useful, there are many other potential screening mechanisms that have yet to be empirically investigated with middle school students, including the use of extant data sources such as course grades, attendance, and discipline records.
The current investigation extends the literature by directly comparing the classification accuracy of existing screening assessments (EWS and CBM) for early identification of students at risk of reading problems and how data from each method may be used to create a more efficient and accurate approach to screening for students in middle schools. The research question for this investigation is as follows:
Method
Participants and Setting
All data for this investigation were collected from a suburban middle school in the Midwest. The school served 434 students in Grades 7 (n = 197) and 8 (n = 237). Enrollment included 48.8% male and 51.2% female. Enrollment by race was reported in the following categories: 4.7% Asian, 24.1% African American, 36.8% Caucasian, 19.6% Hispanic, and 14.6% of students identified as two or more races. Less than 1% of students were categorized as other or unreported. Additional demographic information includes 64.3% of students eligible for free or reduced-price lunch (FRL), 2.2% categorized as English-language learners, and 15.7% enrolled in special education.
Data Collection
Data included student-level demographic information on grade level, gender, race, eligibility for FRL, disability status, disability classification, and status as English-language learners. District staff provided data for attendance, ODRs, course failures, and assessments. Assessment data for the Michigan Education Assessment Program (MEAP) reading subtest were provided in Fall 2012 and Fall 2013. At the time of this investigation, MEAP was administered annually to all public and charter schools in the state of Michigan. Data from the MEAP included raw score, scaled score, and proficiency level.
All students enrolled during the assessment period participated in all assessments. All missing scores in the dataset were attributed to changes in enrollment status for individual students (e.g., students who enrolled after the conclusion of testing) or students who did not take the statewide reading assessment in compliance with their Individualized Education Plan (IEP).
The participating school regularly provided formal training to all staff involved in direct administration of CBM and state assessments. All teachers participated in administration of state assessments. Not all teachers participated in administration of all CBM assessments, though every teacher was required to participate in administration of at least one measure. At the time of this investigation, the school did not systematically collect fidelity data or calculate interrater agreement in scoring assessments and data entry; therefore, researchers were unable to determine fidelity of testing procedures or integrity of the data provided.
Reading Screening Measures
Reading–Curriculum-Based Measure (R-CBM)
R-CBM is a measure of ORF. The wide range of available testing levels and reported correlations from .60 to .80 with standardized reading comprehension makes ORF a popular screening assessment. Barth et al. (2012) reported ORF correlations with external reading measures in the range of r = .64 to .68 (e.g., Test of Sentence Reading Efficiency and Maze reading comprehension) and test–retest reliability of .89 or greater for students in Grades 6 through 8.
All R-CBM assessments for the current investigation were done using the AIMSweb browser-based scoring system. R-CBM was given one on one. A group of five teachers, paraprofessionals, and instructional coaches conducted all administrations of R-CBM.
During screening with R-CBM, each student read three grade-level passages timed at 1 min each. Median scores were reported to the researcher in the form of (a) number of words read correctly per minute, (b) number of errors, and (c) percentage of words read correctly. All three R-CBM passages were equated for difficulty and evaluated for grade-level difficulty by the assessment publisher. The same three passages were used for all students at each grade. Complete technical information, including statistics regarding internal and external validity, is available at http://www.aimsweb.com/wp-content/uploads/aimsweb-Technical-Manual.pdf.
easyCBM Multiple-Choice Reading Comprehension (MCRC)
easyCBM MCRC is an untimed assessment that measures reading comprehension on grade-level appropriate text passages. After reading a passage, students answer factual, inferential, and analytical questions. For students in seventh and eighth grade, MCRC contains a total of 20 questions. MCRC is designed to be group administered and may be administered via online assessment modules or hardcopy formats.
Unlike other forms of CBMs such as ORF and Maze, which assess basic skills that correlate with grade-level content skills, MCRC directly measures grade-level reading comprehension. The format of easyCBM MCRC (e.g., a typical multiple-choice reading test) is the most common format of reading comprehension assessment (Andreassen & Bråten, 2010), similar to many other measures of reading comprehension, including many statewide assessments of reading throughout the country (Deboy, 2013). Reports on easyCBM MCRC at seventh grade show technical adequacy in alternate form reliability, Rasch item analysis, and split half reliability (range = .12–.63; Irvin, Alonzo, Lai, Park, & Tindal, 2012). Saez et al. (2010) report overall classification accuracy for easyCBM MCRC for seventh grade ranging from AUC = .77 to .81 for the test forms used in triannual screening.
Group administration of easyCBM MCRC was done electronically using iPads. All students took the easyCBM MCRC during the predetermined benchmark assessment period. Scores for easyCBM MCRC were recorded in the dataset as raw scores. Technical documentation for easyCBM MCRC is available from Behavioral Research & Teaching at http://www.brtprojects.org/publications/technical-reports/.
Maze reading comprehension
Maze is a CBM that assesses silent reading fluency and basic reading comprehension. Maze was group administered in hardcopy format via AIMSweb during language arts classes. The same passage was used for all students at each grade. Students took the standard assessment associated with their current grade level. Scores were provided to researchers for the total words marked correctly, total errors, and percentage accuracy. Further technical information regarding Maze assessments through AIMSweb is available at http://www.aimsweb.com/wp-content/uploads/aimsweb-Technical-Manual.pdf.
For Maze, students read a grade-level appropriate passage in which every seventh word has been replaced by three boldface words in parentheses. Students read the passage and circle the word in parentheses that best completes the sentence. Maze is timed at 3 min. Students receive a score indicating the number of words chosen correctly and the number of words marked incorrectly. Maze is generally considered a moderately to highly reliable assessment of reading (Barth et al., 2012; Espin et al., 2010). Espin et al. (2010) report validity coefficients r = .70 or greater in predicting scaled scores on statewide assessments of reading achievement for students in eighth grade.
EWSs
Attendance
Attendance is an EWS with direct connection to academic proficiency and successful graduation. When students are not in attendance, their access to content is limited to what they can manage independently. Allensworth and Easton (2005) cite absences as a predictor of failure for students in ninth grade, noting a negative relation R2 = −.51 between absences and grade-point average.
Attendance data were provided for all students. Per school policy, all teachers were required to record student absences and late arrivals at the beginning of each of seven class periods. For the purposes of this investigation, all period absence codes (e.g., E = excused absence, U = unexcused absence) were collapsed into a single code (A = absence). All period absences were aggregated to yield the total period absences accumulated in the most recent school year.
ODRs
ODR is the term associated with all major and minor disciplinary incidents that occur within the school environment. ODRs are well-known indicators of students’ behavior that are statistically significant predictors (r2 = .65–.91, p ≤ .001) of whether students are on track academically for graduation (Tobin & Sugai, 1999).
School staff tracked ODRs for all students using the School-Wide Information System (SWIS) online behavior tracking system (pbisapps.org). All instances of lunch detention, after-school detention, Saturday-school, in-school suspensions, out-of-school suspensions, expulsions, and referrals to the school’s time-out room were logged as ODRs.
Failing grades
Without passing grades, it is unlikely that students will accrue sufficient credits for graduation and will not reach acceptable levels of academic proficiency required for postsecondary education or the workplace (Casillas et al., 2012). When compared with all other students, students who exhibit even a single-course failure in ninth grade show increased rates of future-course failure, decreasing test scores, and increases in behavior problems (Cohen & Smerdon, 2009).
All final course grades from the prior school year were provided to the researcher as letter grades (A, B, C, D, and F). The total number of course failures was calculated for each student. Although the participating school considered a grade of D as passing, the inclusion of D as a failure in screening procedures decreases the likelihood of misidentifying risk status for students who may be near failure but are technically passing courses. Inclusion of D grades as failures is also consistent with research and policy recommendations of EWSs (Balfanz, 2009; Balfanz et al., 2007).
Reading outcome measures
The MEAP was the statewide system of assessments for academic progress for the state of Michigan. MEAP annually assessed grade-level content in reading, writing, math, science, and social studies in the fall of each academic year for third through ninth grades. Fall administration of MEAP occurred approximately 2 weeks after the close of the fall universal screening period.
As the statewide reading assessment, MEAP was aligned with the state-adopted standards for reading at each grade level. All students in third through ninth grade were required to take MEAP unless the student had an active IEP that indicated taking the statewide reading tests was not educationally appropriate for the student (Michigan Department of Education [MDE], 2012). One student with an IEP in the dataset was marked by the school as eligible for an alternative state assessment and did not take MEAP reading.
Results for MEAP reading were provided to the researcher in the form of scaled scores and proficiency levels. Proficiency levels were determined by the Michigan Department of Education (MDE), Bureau of Assessment and Accountability (BAA). Scaled scores for each student were organized according to predetermined performance criteria. Performance levels include 1 = advanced, 2 = proficient, 3 = partially proficient, and 4 = not proficient. Only students scoring advanced and proficient are counted as proficient under Michigan State Accountability Criteria.
Technical documentation for MEAP reports statewide classification accuracy in reading of 73.9% for seventh grade and 74.7% for eighth grade. Item response theory (IRT) statistics for reliability range from α = .80 to .82 for seventh and eighth grade (MDE, 2012).
For the 2014–2015 school year, MEAP was discontinued by the MDE. As of Spring 2015, the statewide assessment of reading achievement and accountability was changed to the Michigan Student Test of Educational Progress (M-STEP).
Data Analysis
With the primary outcome of consideration being accuracy in predicting proficiency on high-stakes assessments, the outcome of greatest interest is dichotomous (proficient or not proficient). For hypothesis testing of categorical and continuous predictors on dichotomous outcomes, logistic regression is the most appropriate method of analysis (Peng, Lee, & Ingersoll, 2002). Scores for the dependent variable (MEAP reading) were collapsed into binary outcomes. MEAP proficiency levels (1) advanced and (2) proficient were combined into a single category of (1) proficient. MEAP proficiency levels (3) partially proficient and (4) not proficient were combined into a category of (0) not proficient.
Hierarchical logistic regression was used to assess the accuracy and value added of CBM and EWS variables, individual and in combination, in predicting outcomes on MEAP. Predictor variables were added in a step entry for each block. To control for variance due to demographic variables, ethnicity, gender, and students qualifying for FRL were entered in Block 1 of each analysis. The inclusion of demographic variables in subsequent blocks was contingent upon statistical significance. Significant (β value, p < .05) demographic variables were retained in subsequent blocks. Nonsignificant variables (β value, p > .05) were removed from the model.
Order for block entry of CBM variables was based on the results of prior CBM research for students in middle schools (Baker et al., 2015; Stevenson, 2015). Variables were entered in order of the strength of predictive relation with the outcome measures, with the strongest predictors entered first and the weakest predictors entered last. Block 2 included easyCBM MCRC. Block 3 included R-CBM. Block 4 included Maze reading comprehension. The CBM prediction model was run independently at each grade level.
EWS variables were analyzed using the same procedures as described for CBM variables. Block entry of variables for EWSs data was based on variance explained and likelihood of correctly predicting school drop-out from extant data (Balfanz, 2009; Balfanz et al., 2007; Heppen & Therriault, 2008), and the use of past scores on state assessments to predict future scores on state tests (Casillas et al., 2012; Denton et al., 2011; Reed et al., 2012).
Block 2 included the most recent MEAP reading scale score (MEAP-Y1). Block 3 included the total number of ODRs for the prior school year (Tobin & Sugai, 1999). Block 4 included attendance in the form of the total number of period absences from the prior school year. Block 5 included the total number of course failures. Course failures include grades of both D and F, commensurate with the findings of Balfanz and colleagues (2007). EWS analyses were run independently at each grade level.
Results
EWSs
Test statistics for model fit and variance explained for outcomes on MEAP reading are summarized in Table 1 (seventh grade) and Table 2 (eighth grade). Chi-square and Hosmer–Lemeshow (H–L) tests were performed for goodness of fit.
Model Fit for Seventh-Grade EWS Variables on MEAP Reading.
Note. Classification cutoff = 0.5. n = 158. EWS = early warning signs; H–L = Hosmer–Lemeshow; FRL = students qualifying for free and reduced-price school lunch; MEAP-Y1 = Michigan Education Assessment Program reading subtest; ODR = office discipline referral; attendance = total number of period absences; course failure = total number of courses failed.
Gender and ethnicity were nonsignificant (p > .05) in Block 1, FRL was retained in all subsequent block entries of variables.
Model Fit for Eighth-Grade EWS Variables on MEAP Reading.
Note. Classification cutoff = 0.5. n = 229. EWS = early warning signs; MEAP = Michigan Education Assessment Program; H–L = Hosmer–Lemeshow; FRL = students qualifying for free and reduced-price school lunch; ODR = office discipline referral; attendance = total number of period absences; course failure = total number of courses failed; MEAP-Y1 = MEAP reading scale score.
Gender and ethnicity were nonsignificant (p > .05) in Block 1, FRL was retained in all subsequent block entries of variables.
For seventh grade, omnibus tests showed significance in Block 1 (FRL), χ2(1, N = 158) = 27.56, p < .001, and Block 2 (MEAP-Y1), χ2(1, N = 182) = 78.882, p < .001. Although the overall model itself retained significance in each block, the addition of predictors in Block 3 (ODRs), χ2(1, N = 182) = 1.65, p = .198, Block 4 (attendance), χ2(1, N = 158) = 3.00, p = .083, and Block 5 (course failure), χ2(1, N = 182) = .728, p < .3931 was nonsignificant. Classification results for seventh-grade EWSs to predict MEAP reading showed an accuracy rate of 60.1% for FRL alone in Block 1 (see Table 3). FRL was retained in subsequent block entries. The addition of MEAP-Y1 (prior performance on MEAP reading) in Block 2 increased classification accuracy to 74.7%. Adding ODRs in Block 3 increased accuracy to 75.9%. The addition of attendance in Block 4 produced no increase or decrease in accuracy. In Block 5, course failure decreased classification accuracy down to 74.1%.
Logistic Regression Results for Seventh-Grade EWS on MEAP Reading.
Note. Classification cutoff = 0.5. n = 158. EWS = early warning signs; MEAP = Michigan Education Assessment Program; FRL = free and reduced-price lunch; ODR = total number of office discipline referrals; attendance = total number of period absences; MEAP-Y1 = MEAP reading scale score.
For eighth-grade EWS, omnibus tests showed significance in Blocks 1, χ2(1, N = 222) = 13.16, p < .001, and 2, χ2(2, N = 221) = 123.25, p < .001. The addition of ODRs, attendance, and course failure to the model in subsequent blocks produced an overall model that was significant even though the addition of each individual variable was nonsignificant. Classification results for eighth-grade EWS to predict MEAP reading showed a larger-than-expected accuracy for FRL at 64.4% in Block 1 alone (see Table 4). The addition of MEAP-Y1 in Block 2 increased classification accuracy to 84.7%. The addition of ODRs, attendance, and course failure in subsequent blocks produced no increase or decrease in classification accuracy.
Logistic Regression Results for Eighth-Grade EWS on MEAP Reading.
Note. Classification cutoff = 0.5. n = 229. EWS = early warning signs; MEAP = Michigan Education Assessment Program; FRL = free and reduced-price lunch; ODR = total number of office discipline referrals; attendance = total number of period absences; MEAP-Y1 = MEAP reading scale score.
CBMs
Block entries of CBM variables were entered in the following order: demographic variables, easyCBM MCRC, R-CBM, and Maze. Test statistics for model fit and variance explained are summarized in Table 5 (seventh grade) and Table 6 (eighth grade). Chi-square and Hosmer–Lemeshow (H–L) tests were performed for goodness of fit.
Model Fit for Seventh-Grade CBM Variables on MEAP Reading.
Note. Classification cutoff = 0.5. n = 180. CBM = curriculum-based measures; MEAP = Michigan Education Assessment Program; H–L =Hosmer–Lemeshow; FRL = students qualifying for free and reduced-price school lunch; MCRC = Multiple-Choice Reading Comprehension; R-CBM = reading-CBM.
Gender was nonsignificant (p > .05) in Block 1; ethnicity and FRL were retained in all subsequent block entries of variables.
Model Fit for Eighth-Grade CBM Variables on MEAP Reading.
Note. Classification cutoff = 0.5. n = 226. CBM = curriculum-based measure; MEAP = Michigan Education Assessment Program; H–L = Hosmer–Lemeshow; FRL = students qualifying for free and reduced-price school lunch. MCRC = Multiple-Choice Reading Comprehension; R-CBM = Reading-CBM.
Gender and ethnicity were nonsignificant (p > .05) in Block 1; FRL was retained in all subsequent block entries of variables.
For seventh-grade CBM predicting MEAP reading, omnibus tests showed significance in Block 1 (FRL + ethnicity), χ2(1, N = 182) = 15.083, p = .002, Block 2 (easyCBM MCRC), χ2(2, N = 182) = 40.018, p < .001, Block 3 (R-CBM), χ2(3, N = 182) = 63.374, p < .001, and Block 4 (Maze), χ2(4, N = 182) = 64.528, p < .001, for the overall model. The addition of easyCBM MCRC in Block 2, χ2(2, N = 182) = 25.314, p < .001, and R-CBM in Block 3, χ2(3, N = 182) = 23.355 p < .001, was significant. The inclusion of Maze, χ2 = (4, N = 182) = 1.154 p < .282, in Block 4 was nonsignificant.
Results of classification accuracy for seventh-grade CBMs to predict MEAP reading showed an accuracy rate of 66.1% for FRL and ethnicity in Block 1 (see Table 7). FRL and ethnicity were retained in subsequent block entries. The addition of easyCBM MCRC in Block 2 increased classification accuracy to 67.8%. Adding R-CBM in Block 3 increased classification accuracy to 73.9%. Adding Maze in Block 4 resulted in a decrease in classification accuracy to 72.8%.
Logistic Regression Results for Seventh-Grade CBMs on MEAP Reading.
Note. Classification cutoff = 0.5. n = 180. CBM = curriculum-based measure; MEAP = Michigan Education Assessment Program; FRL = students qualifying for free and reduced-price school lunch. MCRC = Multiple-Choice Reading Comprehension; R-CBM = Reading-CBM; Maze = Maze Reading Comprehension; MEAP-Y1 = MEAP reading scale score
Gender was nonsignificant (p > .05) in Block 1; ethnicity and FRL were retained in all subsequent block entries of variables.
For eighth-grade CBM predicting MEAP, H–L tests in Blocks 1 and 2 were nonsignificant. Block 3, which included easyCBM MCRC and R-CBM, was significant, H–L χ2(8, N = 222) = 16.1, p = .041 (see Table 6). Omnibus tests showed significance in Blocks 1, χ2(1, N = 222) = 12.525, p = .006; 2, χ2(2, N = 222) = 13.16, p < .001; 3, χ2(3, N = 222) = 104.681, p < .001; and 4, χ2(4, N = 222) = 110.119, p < .001. The addition of easyCBM MCRC, R-CBM, and Maze in subsequent blocks was significant at each point of entry. Classification results for eighth-grade CBM showed an accuracy of 65.9% for FRL in Block 1 alone (see Table 8). The addition of easyCBM MCRC in Block 2 increased classification accuracy to 79.6%. In Block 3, R-CBM improved accuracy to 84.1%. In Block 4, the addition of Maze reduced overall classification accuracy to 82.7%.
Logistic Regression Results for Eighth-Grade CBM on MEAP Reading.
Note. Classification cutoff = 0.5. n = 226. CBM = curriculum-based measure; MEAP = Michigan Education Assessment Program; FRL = students qualifying for free and reduced-price school lunch. MCRC = Multiple-Choice Reading Comprehension; R-CBM = Reading-CBM; Maze = Maze Reading Comprehension; MEAP-Y1 = MEAP reading scale score.
Gender and ethnicity were nonsignificant (p > .05) in Block 1; FRL was retained in all subsequent block entries of variables.
Secondary Analysis
Race, gender, and socioeconomic status were included in the block entry of variables to control for variance that may be attributed to demographic characteristics with established correlations to below-average reading achievement. The inclusion of such factors to isolate the variables of interest is a hallmark of rigorous statistical analyses in social science. However, including demographic characteristics in academic screening couldopen a school to criticism of institutional bias and even racism. With these practical implications of risk screeningin mind, analyses were rerun without demographic variables.
Results without inclusion of race, gender, or FRL predictors showed results similar to those described above. Classification results for seventh-grade EWSs to predict MEAP reading showed an accuracy rate in Block 1 (MEAP-Y1) of 78.8%. The addition of ODRs, attendance, and course failure to the model produced an overall model that was significant even though the addition of each variable was nonsignificant. Adding ODRs in Block 2 decreased classification accuracy to 77.9%. The addition of attendance in Block 3 further decreased classification accuracy to 75.7%. In Block 5, the addition of course failure produced a decrease in classification accuracy down to 74.8%.
For eighth-grade EWS predictors, classification results showed an accuracy rate in Block 1 (MEAP-Y1) of 82.0%. The addition of ODRs, attendance, and course failure to the model was nonsignificant in Blocks 2 through 4 with a classification accuracy range of 82.0% to 82.9%.
Using CBMs to predict MEAP reading in eighth grade, easyCBM MCRC in Block 1 produced a classification accuracy of 78.8%. The addition of R-CBM in Block 2 was significant (p < .001) and increased overall classification accuracy to 81.9%. Adding Maze in Block 3 was significant (p < .009) and increased overall classification accuracy to 82.7%.
For seventh grade, CBM predictors produced a classification accuracy of 68.3% for easyCBM MCRC in Block 1. The addition of R-CBM in Block 2 was significant (p < .001) and increased overall classification accuracy to 75.0%. Adding Maze in Block 3 was nonsignificant (p = .19) and increased overall classification accuracy to 76.7%.
Although there were differences between analyses that included demographics and those that did not, the overall findings were consistent. EWS and CBM produced comparable rates of classification accuracy with EWS performing slightly better in both seventh and eighth grades.
Discussion
Summary
This investigation compared two different models for the identification of students in need of reading support: one based on extant data and the other based on additional-reading-specific screening tests. A search of extant literature revealed a need to compare methods regarding use at the middle school level. There is currently no consensus on the most appropriate way to identify middle school students at risk of underperformance in reading.
Classification accuracy between EWS and CBM returned a less than 1% difference (seventh grade = 0.2%, eighth grade = 0.6%). The EWS model accounted for greater variance explained in outcomes than the CBM model. This is not surprising, given that the EWS model included previous scores on state assessments. With previous scores on statewide reading assessments being a predictor of future scores on state reading assessments (Denton et al., 2011), it makes sense that prior state test scores account for the largest proportion of variance in the EWS prediction model. Findings of this investigation show that variance explained by prior state test scores on future state test scores in reading (seventh = 55.1%, eighth = 58.1%) are notably higher than results reported by Denton et al. (2011). In their analysis of scores for eighth-grade students in Texas, researchers found that 44% of the variance explained on future state test scores was accounted for by the previous year’s state test scores. Results of Denton et al. (2011) and the current investigation support the assertion made by Reed and colleagues that data from annual statewide reading assessments should be “an integral part of the universal screening system” for reading in secondary schools (Reed et al., 2012, p. 49).
Results of CBM analysis indicate easyCBM MCRC accounted for the largest proportion of variance. This is perhaps not surprising, given that easyCBM MCRC and MEAP reading are both multiple-choice reading tests that assess factual, inferential, and evaluative questions. The analysis in this investigation did not explore the effects of test format as a moderator though there may be some advantages to using one multiple-choice reading test to predict outcomes on another multiple-choice reading test.
The addition of R-CBM provided an incremental increase that was small but consistently significant across grade levels. As a measure of ORF, R-CBM measures a different aspect of reading than does easyCBM MCRC. EasyCBM MCRC focuses on passage reading, recall, and application; R-CBM is a test of speed and accuracy (L. S. Fuchs, Fuchs, Hosp, & Jenkins, 2001).
Including Maze in the model failed to produce consistent or statistically significant contributions to the classification accuracy. Although the lack of consistent positive results with Maze may initially seem problematic, it is important to consider that Maze is a combined measure of silent reading fluency and reading comprehension. With the two variables entered immediately prior to Maze being measures of reading comprehension (easyCBM MCRC) and fluency (R-CBM), it is possible that variance potentially attributable to Maze was already accounted for by R-CBM and easyCBM MCRC.
For the EWS models, it is clear that prior scores on state assessments in reading account for the vast majority of variance explained on future outcomes of state reading assessments. This pattern occurred at both seventh and eighth grades for outcomes on MEAP reading. These findings are consistent with similar analysis conducted by Denton and colleagues (2011) that found prior scores on state reading assessments in Texas to be the best overall predictor of future scores on state reading assessments compared with nine other commonly used reading assessments.
The addition of ODRs, attendance, and course failure produced only incremental nonsignificant increases in variance explained and classification accuracy. This result is somewhat unexpected, given the reported strength of behavior, attendance, and course failure as predictors of high school drop-out (Balfanz, 2009; Balfanz et al., 2007; Neild et al., 2007). The key to the weak value added by attendance, behavior, and course failure may be in part due to comorbidity of academic and behavior problems. Because variables included in EWS are likely to be functionally related, students who struggle in one area are likely to struggle in another. For example, students with poor attendance (>10% absences) are more likely to fail courses than those with adequate attendance. This means that there is likely some overlap in the groups of students predicted at risk if the EWS variables were each explored independently. Results of the investigation by Neild et al. (2007) allude to this in reporting that the chance of a student dropping out of high school was 75% for students who exhibit one or more warning signs. However, the more likely cause of the weak value added by attendance, behavior, and course failure may be due to the fact that the dependent variable, in this case, is state test outcomes and not high school drop-out. State achievement tests and high school drop-out are related but nevertheless demonstrably different variables.
From a practical standpoint, it is important to consider the balance between the value added by including ODRs, attendance, and course failure in the model and the costs associated with collecting and managing such data. Data from extant sources are routinely collected and often required for state reporting purposes. Therefore, it may be beneficial to keep ODRs, attendance, and course failure as part of screening procedures. The incremental benefits in accuracy of prediction may outweigh the time, effort, and resources necessary to include these data. In sum, if schools use an EWS model for identification of students in need of reading support, and the data on ODRs, attendance, and course failure are readily available, school personnel should consider such data for screening purposes.
Although results of the current investigation show that EWS and CBM categorize students with similar accuracy, results do not support the elimination of CBM testing. It may be tempting to conclude that CBMs are no longer required for screening purposes. However, there are situations in which CBMs are necessary due to extant data that are inaccurate, incomplete, inaccessible, or do not exist. In these cases, use of CBMs is recommended. Furthermore, providing students the essential services for success requires understanding exactly which services students need. Even with sufficient data infrastructure, it is difficult to pinpoint specific areas of concern without a proximal assessment of reading.
Finally, there are several barriers that may inhibit schools’ ability to use extant data as a method for screening. In many cases, teachers, interventionists, and support staff may not have direct access to records of state assessment scores, absences, ODRs, and course failure. Likewise, schools may not have sufficient data infrastructure and/or expertise necessary to compile EWS data in a format that is suitable for screening purposes. There are technological options that make compiling EWS data less challenging; however, implementation often requires additional time, training, and funding.
Limitations
Data for this investigation came from a single school serving two grade levels (seventh and eighth). The sample size (n = 434) for both grades combined is not atypical in the field of CBM research (Yeo, 2009), but nonetheless could benefit from a larger sample across more schools and a wider geographic area. Because data were limited to a single site, it was not possible to control for school-level, district-level, or regional-level effects in the analysis.
Another limitation is the exclusive use of secondary data. Data were provided directly by the participating school, and it was not possible to assess the fidelity of testing procedures and the accuracy of data entry. Likewise, as students’ assessments were not available to researchers, procedures for interrater agreement of hand-scored assessments were not available. A study on training and fidelity of test procedures found that 8% of all ORF administrations resulted in procedural errors that compromised the data (Reed & Sturges, 2012). Similarly, Cummings, Biancarosa, Schaper, and Reed (2014) found that up to 16% of variance in ORF scores may be attributed to variation between test proctors. Although there is no evidence in the current study of poor implementation fidelity or data quality that would compromise the findings, ensuring fidelity of implementation is an essential component of scientific inquiry and cannot be assumed. Lack of data on the fidelity of testing procedures and data integrity means that results must be interpreted with caution.
It is also important to recognize that this analysis of EWS and CBM focuses on only one of the many functions of CBM and EWS data. CBM in particular can be useful for progress monitoring, evaluation of instructional effectiveness, and as part of the referral process for special education, among others. Although EWS and CBM appear to be comparable methods for identification purposes, there are many other aspects of such data sources that warrant thorough empirical investigation.
Suggestions for Future Research
Results of this study suggest that extant data sources, including prior scores on statewide reading achievement tests, may be a feasible alternative to the use of CBMs in reading for universal screening of students in middle school, when CBM screening is not possible. However, it is presently unclear how the methods tested in this study will play out in practice for school personnel. Arming schools with accurate and efficient methods for identification of students in need of support is only the first step in intervention and prevention of persistent academic difficulties. The ways in which school personnel use CBM and EWS data to make decisions regarding services for students and allocation of resources are equally important. Future research is needed that explores the practical implications and feasibility of using CBM and EWS as a mechanism for universal screening in middle schools.
In addition, results of the current study fuel the need to explore more accurate, reliable, and efficient screening systems. Some options include the use of multiple gating procedures (Nelson et al., 2016) and advanced systems that combine extant data and CBM data into a single integrated system of screening and formative assessment.
Footnotes
Acknowledgements
The author wishes to thank Dr. Cynthia Okolo for her support in completion of this study.
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
