Abstract
We conducted conceptual replications for an online-administered general outcome measure known as critical content monitoring in science. Totally, 547 Grades 4 to 6 students from a rural Louisiana school district participated. Research questions addressed criterion validity, diagnostic accuracy, growth, and social validity concerns. Criterion measures included science and reading/literacy subtests from a state accountability test and a nationally standardized test. Findings indicated moderate to strong concurrent and predictive criterion validity correlations for the overall scores, fair diagnostic accuracy statistics for fall and winter and moderate for spring benchmark scores, variable growth scores across grades, and favorable social validity comments from teachers. Study limitations were listed. Research and practice implications were discussed.
Critical content monitoring (CCM; Mooney, McCarter, Russo, & Blackwood, 2013) is an online-administered general outcome measure used in science and social studies classrooms. It was designed to provide teachers and students with an efficient and effective means of documenting learning in upper elementary and secondary school content courses. An adaptation of vocabulary matching (Espin & Deno, 1994-1995), CCM is one of 11 instruments that are named and/or evaluated in the general outcome measurement literature for science and social studies content.
General outcome measurement is an assessment framework characterized by a long-term learning focus and prescriptive measurement procedures (L.S. Fuchs & Deno, 1991). When developed, general outcome measurement was contrasted with the then and still now more traditional short-term and skill-based nature of subskill mastery measurement. General outcome measurement has a demonstrated track record in terms of differentiating regularly performing students from those at risk and documenting academic growth, particularly in the areas of reading and math (L. S. Fuchs, 2017). General outcome measurement probes are created in two ways (L. S. Fuchs, 2004). One method identifies robust (and stronger in magnitude) indicators of curricular proficiency (e.g., oral reading fluency, ORF; Deno, Mirkin, & Chiang, 1982). An alternate method systematically samples from the skills/knowledge that comprise an annual curriculum (e.g., mathematics concepts and applications; L. S. Fuchs, Fuchs, & Zumeta, 2008).
In the general outcome measurement literature for science and social studies content, robust indicators are those measures that have students engage academically with relevant content through reading, listening, and/or writing channels, similar to what students do when completing science or social studies coursework in school settings. In the measurement of science and social studies content, general outcome measurement robust indicator instruments include ORF, traditional maze, concept maze, content maze, sentence verification technique, and statement verification for science (Busch & Espin, 2003; Hosp & Ford, 2014; Johnson, Semmelroth, Allison, & Fritsch, 2013; Royer, Hastings, & Hook, 1979; Wayman, Wallace, Wiley, Tichá, & Espin, 2007). Meanwhile, curriculum-sampling measures target the academic vocabulary of a given subject. Science and social studies measures include vocabulary matching, CCM, and key vocabulary (Espin & Deno, 1994-1995; Mooney, McCarter, Russo, et al., 2013; Mooney, McCarter, Schraven, & Callicoatte, 2013; Vannest, Parker, & Dyer, 2011).
As indicated, CCM serves as a brief academic indicator of a student’s overall success mastering the expectations of a science or social studies course and uses curriculum-sampling principles, similar to math and spelling general outcome measurement tools. Academic vocabulary is the measurement index because knowledge and use of academic content is essential to classroom practice. Townsend (2015) describes academic vocabulary as academic word knowledge and a useful component of the broader academic language, defined as “the specialized language, both oral and written, of academic settings that facilitates communication and thinking about disciplinary content” (Nagy & Townsend, 2012, p. 92). Academic vocabulary, which explains unique variance in achievement after controlling for general vocabulary knowledge in middle school students (Townsend, Filippini, Collins, & Biancarosa, 2012), is made up of general and discipline-specific terms. General academic words, such as structure, function, method, and experiment, are used across disciplines (Nagy & Townsend, 2012). Discipline-specific academic words, such as mitosis and cytoplasm, are typically presented within a specific subject area (Nagy & Townsend, 2012).
Academic vocabulary and language are communicative currency (Alexander, n.d.) in that they are central to teacher instruction, student activity and learning, and classroom discourse. They are what is shared, used, taught, and learned as students read textbooks and complete assignments and teachers lecture on and/or discuss facts and concepts, and develop, deliver, and evaluate student work. Content classroom success occurs with meaningful use of academic vocabulary and language. Nagy and Townsend (2012) contend that academic vocabulary and language proficiency enable students to more completely access the meaning in academic text and discussion, act like scientists and historians, and better demonstrate achievement in school. From an assessment perspective, differences in students’ skill levels can be quantified in terms of their academic vocabulary proficiency, with differentiation of achieving from struggling students (Baker, Kame’enui, Simmons, & Simonsen, 2007) and those with adequate versus inadequate learning histories (Beck, McKeown, & Kucan, 2002).
In upper elementary and secondary classrooms, competence has always been demonstrated to some degree by passing coursework consisting of assignments and teacher-made tests comprising academic vocabulary. Over the last two decades, school-level success has also been associated with the number of students who meet state-level standards on standardized, state-developed content tests. From a practical assessment perspective, demonstration of academic vocabulary’s robustness has been evident in the stronger correlations with a relevant criterion that general outcome measures of academic vocabulary such as CCM and vocabulary matching have demonstrated over competing measures (e.g., ORF, maze; Espin & Foegen, 1996; Mooney & Lastrapes, 2016; Mooney, McCarter, Schraven, et al., 2013).
As a general outcome measure, CCM is administered using a standardized administration and scoring format. Students access individual probes online and answer multiple-choice questions that include a definitional stem and vocabulary alternatives within a specified time frame. Traditionally, students take up to 5 min to complete as many of 20 questions as they are able and/or motivated to answer. At test end, students have access to their score.
Rationale for Continued Inquiry
To date, CCM has demonstrated its efficiency in two ways. First, its administration and scoring functions are managed online, with scores provided immediately and without teacher, student, or researcher time and effort directed at scoring. Second, teachers access multiple alternate forms of the probe since they are developed ahead of time. Indication of CCM’s empirical effectiveness includes statistically significant and moderate in magnitude correlations with standardized content achievement tests across science and social studies content and testing formats (Mooney, McCarter, Russo et al., 2013; Mooney, McCarter, Russo, & Blackwood, 2014). In addition, when CCM has been compared with other general outcome measures, its correlations with standardized criterion measures have been descriptively stronger across both predictive (i.e., fall and spring benchmarks) and concurrent (i.e., spring) criterion validity circumstances (Lastrapes & Mooney, 2017; Mooney & Lastrapes, 2016; Mooney, Lastrapes, Marcotte, & Matthews, 2016).
Given the formative nature of the general outcome measurement literature for science and social studies content collectively and CCM specifically, we believed that continued scientific inquiry regarding CCM was warranted. Recent scholarly attention has been directed toward replication research in the field of special education. Systematic replication is considered vital to the accumulation of scientific knowledge, bolstering of confidence in findings that exist, and identification of research- and evidence-based practice (Travers, Cook, Therrien, & Coyne, 2016). While the replication emphasis in special education has been on intervention inquiry, replication of relevant assessment scholarship merited attention. That included general outcome measurement for science and social studies content, in part due to its limited history. The systematic development of bodies of evidence through direct and conceptual replications allows for greater consumer confidence in the claims that are made by researchers (Travers et al., 2016).
The present study incorporated conceptual replications of CCM validity research (Mooney & Lastrapes, 2016) with larger and more representative samples than had been included previously. Participant pools in individual content general outcome measurement studies have generally been small in number, with less than 200 students in a given study and singular in grade level, locale, and content area (e.g., Espin et al., 2013). For CCM, previous research studies have included 50 to 100 participants in single grades and school settings targeting both science and social studies content (Mooney & Lastrapes, 2016; Mooney, McCarter, Russo et al., 2013, 2014). The present study focused on science content and involved multiple schools, teachers, and grade levels and considerably larger student numbers, including students with disabilities. Replication was conducted using both a state and a national standardized measure of achievement.
Conceptual replication of CCM inquiry was designed and informed by the broader general outcome measurement literature. First, analysis of the similarity of general outcome measurement scoring patterns to those of relevant criteria has been documented in the vocabulary matching literature. That is, Mooney, McCarter, Schraven, and Haydel (2010) demonstrated that 9 of 10 significant or nonsignificant differences in variable patterns for the statewide accountability test scores were evident for vocabulary matching in sixth-grade social studies content. We hypothesized that similarities would exist in the present criterion-predictor comparison in science content. Second, as noted previously, criterion validity studies in the published literature have targeted single grades to date, focusing primarily on middle/junior high school students. Moderate concurrent correlations have been reported for individual samples in Grades 5 through 7 and 10 (e.g., Espin et al., 2013). Mooney and Lastrapes (2016) reported moderate predictive correlations for CCM fall (r = .56) and winter (.51) benchmark probe scores with those of spring statewide accountability test scores. We expected to see moderate concurrent and predictive correlations across Grades 4 to 6 for the larger sample.
Third, diagnostic accuracy calculations in the content general outcome measurement literature have been limited, with statistics first reported by Johnson et al. (2013) for a content maze probe in middle school science content. The present study provided the first set of data for the curriculum-sampled measures of academic vocabulary. Fourth, growth data have only been reported for vocabulary matching studies (e.g., Espin et al., 2013; Espin, Shin, & Busch, 2005) and focused on weekly administrations over 11- to 25-week time frames. What has yet to be reported, both in the general outcome measure literature for science and social studies content and for CCM specifically, are benchmark growth scores. Benchmark data are necessary elements of response to intervention frameworks in that they facilitate screening and/or progress monitoring decisions. Fifth, we built upon the social validity comments from students (Mooney, McCarter, Russo et al., 2013) by soliciting the perceptions of the teachers responsible for probe administration.
Finally, we tested the potential for academic vocabulary to serve as a vital indicator of learning (Deno, 1985) by calculating concurrent criterion validity correlations for English language arts (ELA) statewide test scores using CCM content probes. Deno (1985) hypothesized that curriculum-derived probes could serve as quick checks of student performance and progress. As the general outcome measurement development process was applied to science and social studies content, Espin and Deno (1994–1995) discovered that a timed measure of academic vocabulary (i.e., vocabulary matching) was a better predictor of content performance than reading indicators ORF and maze in science and English content. If academic vocabulary serves as a valid indicator across science, social studies, and ELA content and with older students, then correlations between academic vocabulary-derived probes and meaningful criterion could be useful catalysts for secondary school response to intervention frameworks. We hypothesized that ELA correlations would be smaller in magnitude than those with science scores since the terms and definitions were science focused.
Overall, the following research questions were addressed:
Method
Participants
Participants included 547 Grades 4 to 6 students from seven schools in a rural south Louisiana district of about 4,700 students, approximately 9% of which were students with disabilities. District fourth (n = 158), fifth (n = 247), and sixth graders (n = 142) who completed all benchmark CCM probes and had state achievement scores comprised the sample. The demographic breakdown was 50% female, 68% non-White, 86% lower socioeconomic status (SES), and 7% students with disabilities. Less than 1% of the entire district were English language learner. The mean age was 11.9 (SD = 1.8). At the request of the district, a subset of the sample (N = 84) in two of the district’s low-achieving schools also completed science and reading comprehension subtests of the online abbreviated Stanford Achievement Tests–Tenth Edition (SAT-10; Pearson Education, n.d.). The number was limited to the funds available to finance the cost of test administration. The makeup of the subset, whose average age was 11.6 years (SD = 1.9), was 56% female, 98% African American, 96% low SES, and 4% students with disabilities.
Measures
Three benchmark CCM probes, two criterion measures, and a teacher social validity measure were administered. Both criterion measures were compared to CCM across multiple administrations. The criterion measures for the larger sample were the science and ELA tests of the Louisiana Educational Assessment Program (LEAP; Louisiana Department of Education [LDE]; fourth grade) and the integrated LEAP (iLEAP; fifth and sixth grades). The smaller sample was administered the iLEAP/LEAP science and ELA tests and the SAT-10 science and reading comprehension tests.
CCM
Probe content was generated from a list of terms at each grade level, organized by curricular unit, and collected from science textbook glossary sections. A small group of practicing teachers recommended to the first author as both content knowledgeable and highly effective teachers reviewed the body of terms for content validity. Researcher-created probe terms were organized in terms of content strand (e.g., science as inquiry, life science) before being randomly selected. Totally, 30 terms and definitions were included in each probe. Probe makeup matched the proportional breakdown for each grade-level strand reported in the state pacing guide. Terms and accompanying definitions were entered into a Qualtrics online survey software system in a multiple-choice format. Previous criterion validity correlations were moderate in magnitude (r = .36 to .67; Marston, 1989) with national and state test in science and social studies content in multiple fifth-grade level convenience samples (Mooney & Lastrapes, 2016; Mooney, McCarter, Russo, et al., 2013, 2014). Alternate-forms reliability correlations for a series of 20 CCM forms were small to moderate in magnitude, with the mean and median reliability correlations in 190 comparisons .56 (SD = .09) and .56, respectively, with a range of .21 to .73 (Mooney, McCarter, Russo, et al., 2013). The probes used in the present study were not evaluated for alternate-forms reliability statistics prior to being administered.
iLEAP/LEAP Grades 4 to 6 criterion-referenced test
The stated purpose of iLEAP/LEAP is measurement toward Louisiana’s academic standards in ELA, math, science, and social studies (LDE, n.d.-b, n.d.-c) for all students in Grades 4 through 9. The science test includes multiple-choice questions and is untimed, with content breakdowns different across grades (LDE, n.d.-b). The ELA test includes reading and writing tasks across fiction and nonfiction texts. Students were expected to read texts and answer questions as well as respond to writing prompts. Achievement level descriptors were unsatisfactory, approaching basic, basic, mastery, and advanced. Technical adequacy data for the iLEAP test were accessed from the LDE website. For Grades 5 and 6, Cronbach’s alpha levels of .87 and .88 were reported for science and .87 and .89 for ELA as reliability evidence of the 2014 test’s internal consistency (LDE, n.d.-a). For the Grade 4 LEAP, Cronbach’s alpha levels were .85 and .90 for science and ELA, respectively (LDE, n.d.-d). State-provided validity data were described in terms of a content validity process that was delineated (LDE, n.d.-a, n.d.-d).
SAT-10 abbreviated online tests
The abbreviated form of the online SAT-10 is a standardized, norm-referenced achievement test battery for students in kindergarten through 12th grade that measures reading, mathematics, spelling, language, listening, science, and social studies performance. Grade-level science and reading comprehension tests were administered to participants. Publishers described the content tests as aligned with national and state content standards and reflective of current practice. The untimed, 30 multiple-choice-question science test assessed knowledge of science as inquiry and life, physical, and earth sciences. The reading comprehension test, also untimed and 30 questions in length, assessed literary, informational, and functional reading purposes and multiple modes of comprehension, including initial understanding, interpretation, critical analysis, and awareness and usage of reading strategies (Pearson Education, n.d.). The test-derived scaled scores were used in the present study. The scaled score is vertically equated across each subject test, reportedly allowing for the tracking of performance across grades (Pearson Education, n.d.). Carney (n.d.) provided a positive review of SAT-10’s capacity to measure public school achievement, reporting alternate-form reliability and content validity evidence for SAT-10. In the present study, criterion validity was evidenced by respective Grades 4 to 6 correlation between the SAT-10 and iLEAP/LEAP science tests of .60, .50, and .61. In ELA, the Grades 4 to 6 correlations were .77, .72, and .67, respectively.
Social validity
Teachers were presented with a researcher-developed survey at the end of the school year and asked to reflect back on the CCM formative assessment process. Only 15 of 31 teachers completed the survey due to the fact that it was disseminated at a training late in the year that was sparsely attended. There were no responses to follow-up email requests to complete the survey. Participants answered two questions. The first was “In your opinion, would the information gained from the CCM procedure be helpful in informing your practice in developing: Course curriculum, establishing progress toward goals, measuring progress toward goals, deciding when to change instruction, communicating student progress to other school personnel, to parents, and to students?” The teachers responded to each subcategory with Likert-type responses (1 = not helpful, 2 = somewhat helpful, 3 = helpful, 4 = very helpful). The second question was “How time consuming were the CCM procedures?”
Procedures
Students took CCM benchmark assessments in October, January, and May. The assessments were delivered via Qualtrics, an online survey software system. Researchers administered the initial October CCM probe with the teachers over the course of a week to facilitate fidelity of implementation. During subsequent administrations, district teachers administered the tests independently during a testing window. The second author conducted informal observations of testing at each of the schools and reported anecdotal evidence that CCM implementation went as planned. However, no formal fidelity checks were conducted. Teachers administered the statewide accountability test in April of the academic year in accordance with guidelines established by the state department of education. Researchers administered the SAT-10 test shortly before state testing.
Data Analysis
Descriptive statistics for all measures were calculated. To examine mean differences among subpopulations on the criterion measures and CCM probes, each measure was examined separately as a dependent variable with demographic variables as independent variables. Using separate multiple regressions, comparisons were examined for the dichotomous demographic categories gender, education classification (students with disabilities vs. students without disabilities), and SES status (high vs. low). Race was treated as a dichotomous variable limited to White and African American students, as only 16 students were in a different racial category. This group of 16 was excluded from the analysis regarding race but were included in the other analyses. For the correlational analysis, bivariate correlation coefficients were calculated to determine the criterion and predictive validity for CCM with iLEAP/LEAP or SAT.
For the diagnostic accuracy analyses, state test data were provided to the researchers as both a continuous score and an “achievement” string variable. Unsatisfactory and approaching basic were considered failing by the state, whereas basic, mastery, and advanced were deemed passing. This variable was transformed into a dichotomous “Pass/Fail” variable. To determine a cut score for the benchmark CCM scores, a Receiver Operating Characteristic (ROC) curve was calculated for the score that came closest to .80 sensitivity and .20 1-specificity level which was then chosen as the cut score. A new dichotomous variable was computed for each probe, with “1” if a CCM score fell below the cut score and “0” if the score was at or above the cut score.
Interpretation of the area under the curve (AUC) was viewed as an effect size of the diagnostic accuracy of the scale with scores between .50 and .69 considered “low,” .70 - .89 “moderate” and .90–1 “high” (Kilgus, Riley-Tillman, Chafouleas, Christ, & Welsh, 2014). Statistics were calculated for true positive (TP), students who were correctly identified as at risk; false positive (FP), students incorrectly identified as at risk; true negatives (TN), students correctly identified as not at risk; and false negatives (FN), students incorrectly identified as not at risk. Sensitivity (SN) was the proportion of students at risk who were correctly identified (TP / [FN + TP]). Specificity (SP) referred to the proportion of students not at risk who were correctly identified as not at risk (TN / [TN + FP]). Positive predictive power (PPP) was the proportion of students correctly identified as at risk (TP / [TP + FP]). Negative predictive power (NPP) was the proportion of students correctly identified as not at risk (TN / [TN + FN]). The prevalence or base rate referred to the proportion of students who were identified as at risk on the given probe ([TP + FN] / N) and the correct classification (CC) described the proportion of the overall sample who were correctly identified ([TP + TN] / N). Finally, Cohen’s Kappa (K) referred to the proportion of agreement between the test and actual condition of the subjects (at risk vs. not at risk) beyond that accounted for by chance, where K values of ≤ .2 was considered “poor,” .21–.40 “fair,” .41–.60 “moderate,” and ≥.61 “good” (Forstmeier & Maercker, 2007; Kilgus et al., 2014).
For growth, a linear mixed model was conducted to determine if CCM accounted for students’ growth over time and if there were differences in growth rates (Espin et al., 2013). The level one variable was time and the level two variable was individual students. First, an unconditional (null) model was calculated to determine if there was enough variation over the participants to warrant this type of analysis. The null model is the equivalent of a one-way analysis of variance (ANOVA) with no predictors, only a random effect (Shek & Ma, 2011). Once this was established, time was entered into the model as a fixed factor to establish the mean growth rate, and then tested as a random factor to determine if differences in the model could be explained by variability between individuals. Finally, the model was tested to see if there were differences among the grades and the interaction of time and grade.
Results
Descriptive Statistics and Demographic Comparisons
Table 1 provides descriptive test data for the three CCM probes and both the state achievement and SAT-10 tests. Overall mean CCM probe scores increased across benchmark time periods, from 10.8 (SD = 5.1; 95% confidence intervals [10.4, 11.2]) in October to 13.2 (SD = 5.1; [12.8, 13.6]) in January to 15.9 (SD = 5.5; [15.4, 16.4]) in May. Each of the three individual grade levels also showed incremental increases in scores out of 30 questions/points possible from fall to winter to spring. The greater proportion of growth came from winter to spring for fourth and fifth grades and for fall to winter for sixth grade.
Critical Content Monitoring and Criterion Distributions of Means, Standard Deviations, and 95% CI by Grade.
Note. CI = confidence intervals; i/LEAP-S = Louisiana state achievement test—Science; i/LEAP-E = Louisiana state achievement test—English language arts; SAT-10-S = Stanford Achievement Test, Tenth edition—Science; SAT-10-E = Stanford Achievement Test, Tenth edition—English language arts.
Table 2 compares the predictor and criterion tests with respect to subgroup scoring pattern differences. There were more comparable scoring patterns for the state test than the national test. Differences among subgroups for the CCM probe scores matched those of the science state achievement test for each benchmark for gender, disability status, and SES but not race. CCM scoring patterns matched those of the state ELA test in disability and SES but not race or gender. CCM scores across all benchmarks only matched those of SAT-10 science for gender. There were no matches reported between CCM scores and SAT reading comprehension.
Comparison of CCM Benchmark Probes and Criterion Tests With Respect to Differences in Subgroup Mean Scores.
Note. CCM = critical content monitoring; B = unstandardized coefficient; SEB = Standard error of the coefficient; β = standardized coefficient; LEAP = Louisiana Educational Assessment Program; SES = socioeconomic status; ELA = English language arts; SAT-10 = Stanford Achievement Tests–Tenth Edition.
p < .05. **p < .01.
Concurrent and Predictive Validity
Table 3 displays correlations between the benchmark CCM probes and iLEAP/LEAP and SAT tests for science and ELA. For science, the concurrent correlation between CCM and iLEAP/LEAP was .50, while the concurrent comparison with SAT-10 was .67. Across grades, correlations ranged from .44 in sixth grade to .66 in fourth grade. Predictive overall CCM science correlations ranged from .38 to .42 for the state test and .40 to .58 for the national test. In ELA, the concurrent correlation between CCM and iLEAP/LEAP was .66, while the concurrent comparison with SAT-10 was .75. Across grades, correlations ranged from .55 in sixth grade to .70 in fifth grade. Predictive overall CCM ELA correlations ranged from .51 to .56 for the state test and .41 to .64 for the national test.
Correlations Across CCM Probes and Criterion Measures—Science in Bold in Lower Matrix, ELA in Upper Matrix.
Note. Science scores in bold in lower left of the matrix, ELA scores in upper right of the matrix. CCM = critical content monitoring; ELA = English language arts; LEAP = Louisiana Educational Assessment Program; SAT = Stanford Achievement Tests; Oct = October CCM, Jan = January CCM, May = May CCM.
not significant.
Correlations marked * significant at p < .05; all other correlations significant at p < .01.
Diagnostic Accuracy
Table 4 provides a summary of diagnostic statistics for the benchmark CCM probes with the state achievement test. In science, the May CCM cut score of ≤14 correctly classified 75% of participants as passing or failing the test based on state criterion levels. The cut scores increased for each successive benchmark period along with the correct classification rate. In ELA a similar pattern was indicated. The October ELA cut score was different than that determined for science, but the January and May scores were the same across content areas. Kappa statistics increased as the year progressed with values that ranged from fair to moderate.
Diagnostic Accuracy Statistics for CCM Probes Across Benchmarking Periods for Science and ELA.
Note. CCM = Critical content monitoring; ELA = English language arts; SN = sensitivity; SP = specificity; PPP = positive predictive power; NPP = negative predicting power; CC = correct classification; BR = base rate/prevalence; K = Cohen’s kappa; AUC = area under the curve; 95% [CI] = 95% confidence intervals for AUC.
Growth
The model building process started with an unconditional or “null” model, in which there were no predictor variables and the intercept was allowed to vary randomly. Calculation of the intraclass correlation determined that there was enough variability in individuals to warrant a linear mixed model. Time was entered into the model as a fixed factor to establish the mean growth rate on CCM. Akaike Information Criteria (AIC) was used to evaluate the model at each step. Decreases in magnitude of AIC indicate a better model and guided the model building process (Bell, Ene, Smiley, & Schoeneberger, 2013). The intercept was centered at the fall benchmark assessment, the initial point of data collection, to establish if there were individual differences in students’ CCM scores at the beginning of the study (see Table 5).
Growth as Measured by CCM.
Note. Intraclass correlation = .45. Entries show parameter estimates and standard errors in parenthesis. Estimation method = maximum likelihood. CCM = critical content monitoring; AIC = Akaike Information Criteria.
AIC value increased, indicating worse model.
p < .05.
The null model (Model 1) revealed that the intercept was statistically significant with a magnitude of 17.8 at initial time of testing. The intraclass correlation (14.3/[17.5+ 14.3]) indicated that about 45% of the total variation in the CCM scores was due to significant differences between the participants. To determine if CCM was an indicator of average growth for students, the predictor time was added as a fixed effect to the model (Model 2). The result was statistically significant and indicated that for every benchmarking period, students showed an average improvement in CCM scores by 2.6 points. To determine if CCM growth varied across students, time was added as a random effect. Because the AIC increased (Model 3), that indicated that Model 2, with fixed effects for time, was the better model. Finally, to examine if there were differences among the three grades, the main effects for grade and the interaction of time and grade were examined. Model 4 shows the results, with the lowest AIC indicating the better fit. The interaction of time and grade was significant, meaning that there was a difference in sixth graders (specifically) depending on when they were tested. With fourth grade as the reference category, it appears that fifth graders scored 2.8 points higher each marking period compared to the fourth graders, and sixth graders scored .83 higher, depending on the marking period (see Table 6).
Parameter Estimates for Growth Model 4 With Grade and Time as Predictors.
Note. Fourth grade reference category.
p < .05.
Social Validity
For the 15 responders, the majority regarded CCM as very helpful in communicating student progress to parents while the majority of respondents regarded CCM as helpful in communicating student progress to students and school personnel, establishing goals, measuring progress toward goals, and deciding when to change instruction. Totally, 68% of responders thought the CCM procedure was not very or not at all time consuming, with 25% indicating that it was time consuming.
Discussion
The present study provided a series of replications of CCM validity inquiry. Research questions targeted L. S. Fuchs’ (2004) Stages 1 and 2 (of 3) general outcome measurement development areas of single score performance (Stage 1) and multiscore growth potential (Stage 2), respectively, replicating science findings and extending inquiry to ELA across grades. We summarize validity findings prior to discussing study limitations and implications for the online general outcome measurement index of content learning known as CCM. Validity is a vital consideration in test development, defined as “the degree to which evidence and theory support the interpretations of test scores entailed by proposed uses of the test” (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014, p. 9). Replication research is valuable in discussing both the quality and quantity of evidence in support of test score interpretation.
CCM science scores were compared against those of a statewide accountability and a nationally standardized test across three upper elementary/lower middle school grades. All previous research had targeted single grade-level samples. Three benchmark scores from October, January, and May were evaluated. On the surface, a measure of face validity was provided in that across the year scores increased, indicating the possibility that exposure to instruction and course content resulted in improved scores on separate probes that are designed to be equivalent in difficulty. Increases occurred in overall mean CCM scores as well as those for each grade level (see Table 1). A similar logical pattern of growth was demonstrated in the scores of students with and without disabilities. The lone exception was the fall-winter comparison in fourth grade for students with disabilities, wherein winter scores were smaller in magnitude than fall scores. Also logically, mean scores for students with disabilities were lower across all periods than those of students without disabilities.
Regarding patterns, there was variability in agreement between spring CCM and the criterion measures in looking at subgroup patterns in science achievement (see Table 2). That is, patterns were similar for the state exam but not for the national test. Across gender, SES (i.e., low vs. high), and exceptionality status (i.e., students with disabilities vs. students without disabilities), CCM performance patterns mirrored those of the state exam. The only difference came in the area of race, where African American–White differences were statistically significant for the state test and not so for CCM. For the national test, however, only the nonsignificant differences in scores across gender matched. Pattern differences between the state and national tests may have been due to the fact that CCM content was created by teachers focused on developing critical state-level content. The two tests may also have had different areas of emphasis given the fact that there were only moderate correlations—.50 to .61—at each grade level. Overall, the pattern analysis is believed to be the first evaluation for CCM. Analyses have been conducted for vocabulary matching, with that paper-pencil measure of academic language demonstrating compelling evidence that general outcome measurement test scores could mirror criterion test performance (Mooney et al., 2010).
Single static correlations with meaningful criterion measures were more variable in magnitude than hypothesized yet generally in the expected moderate range. Previous CCM science concurrent correlations with statewide and national tests ranged from .36 to .55 across two homogeneous, high-achieving fifth-grade convenience samples (Mooney & Lastrapes, 2016; Mooney, McCarter, Russo, et al., 2013). For a more representative sample that included students with disabilities, the overall across grades correlations ranged from .50 (state test) to .67 (national test). The larger correlation compares favorably with those of vocabulary matching, which has the strongest research support in the general outcome measurement for science and social studies content assessment literature. Grade-specific CCM science correlations were variable, ranging from .44 to .76, with both those relationships from the sixth-grade sample. Four of the six concurrent correlations were stronger in magnitude than the range previously reported.
Prior predictive criterion correlations in science content ranged from .51 to .56 for a 20-question CCM probe in fifth grade (Mooney & Lastrapes, 2016). The present findings ranged from .08 to .58 across 12 grade-specific comparisons for a 30-question probe, with the median correlation .52. Moderate correlations were evenly distributed across state and national test comparisons.
Findings extended the CCM technical adequacy literature in three significant ways. First, linear associations were also reported for ELA. Prior inquiry had examined CCM’s tenability as a measure of science and social studies content. Our findings indicated that in spite of the science content of the measures, concurrent and criterion correlations with literacy-oriented criterion measures were generally in the moderate range and often stronger in magnitude than the science results. The overall concurrent correlations ranged from .66 to .75 (see Table 3). For the grade-specific correlations, four of the six ELA concurrent correlations were larger than those for science, including all three state comparisons. However, 7 of 12 grade-specific comparisons favored science. All of these comparisons were descriptive in nature.
The second extension of the literature addressed diagnostic accuracy statistics, an area where scarce attention has been directed in the general outcome measurement for science and social studies content literature. Johnson et al. (2013) reported 69% classification accuracy for a sample of seventh-grade students who were administered winter benchmark content maze probes prior to the spring state science test. Our findings for CCM demonstrated fall and winter correct classification rates of 67% and 70%, respectively, in science for a cut score that, similar to Johnson et al. (2013), balanced sensitivity and specificity statistics (see Table 4). Correct classification rates increased across benchmark periods and content areas, with spring science and ELA figures at 75%.
The final area where CCM technical adequacy research was extended related to growth statistics. Like diagnostic accuracy, growth research in this area is sparse, with all four previous studies targeting vocabulary matching and weekly growth rates. Most recently, Espin et al. (2013) reported a .63 matches a week growth rate for a seventh-grade sample administered science probes over 14 weeks. Weekly rates have been variable across four growth studies targeting science or social studies content. Mooney, McCarter, Schraven, et al. (2013) collected data across a school year in sixth-grade social studies content and found that growth rates differed in the fall and spring semesters. Borsuk (2010) noted differences in growth rates for students with and without disabilities. Our CCM research targeted benchmark scores and compared growth rates across grades. Findings indicated that while a positive overall growth rate of 2.6 correct choices per benchmark period was reported, rates differed across grades and times tested (see Table 5). Grade-level differences seem logical given that commercial general outcome measurement norms (e.g., AIMSweb) across instruments indicate different weekly growth rates both across and within grades. Within grades differences were noted for ORF Grades 2 to 6 benchmark scores by Silberglitt and Hintze (2007).
The present study also marked the first time that CCM growth data were analyzed for students with disabilities. There were statistically significant growth rate differences from first to last benchmark period for students with and without verified disabilities. Recognizing that there was tremendous disparity in the group sizes, students with disabilities, whose initial October benchmark scores were statistically significantly lower than students without disabilities, reported an average growth rate that was 3.3 points lower than their peers without disabilities. Results, then, support CCM’s capacity to measure the performance and progress of students with disabilities, who, in this case, were tested and educated in general education settings. That is, first, as expected, initial performance scores for students with disabilities were lower than those of peers without disabilities. Second, the statistically significant differences in CCM scores for students with and without disabilities mirrored those of state accountability test results indicated across both science and ELA test scores. Third, the growth rate differences for students with and without disabilities paralleled differences reported by Borsuk (2010) across 24 weeks of vocabulary matching testing in middle school science content.
Limitations
Four limitations likely impact the evaluation of study findings. First, the purposeful sample of predominantly African American and lower SES students, while large, multigrade, and quite likely representative of urban public school populations, was not representative of the larger U.S. population demographically, thereby limiting the potential generalizability of findings. Second, the particular probe administration program we utilized, Qualtrics, did not allow for immediate teacher access to useable data as the assessments were put online by the researchers, which may have affected subsequent CCM scores over the course of the study. Third, our analysis did not control for students’ reading abilities, which did not allow for determination of whether or how reading skill affected performance. Finally, alternate-form reliability analyses were not conducted for the probes used for this study. Therefore, it was not known whether the probes were statistically equivalent or not. This is important because normal development of CCM probes does not guarantee statistical equivalence (Mooney, McCarter, Russo, et al., 2013).
Implications for Research and Practice
Previously, we have suggested that indicators of academic vocabulary such as CCM and vocabulary matching be considered viable measures of content assessment frameworks (Mooney & Lastrapes, 2016; Mooney et al., 2016). That research-informed assertion is bolstered by the present conceptual replication findings. First, there are now three studies that report criterion-related validity correlation magnitudes for CCM with meaningful science criterion measures in the moderate to large range (as specified by Marston, 1989) for a large, multigrades, and more representative sample of public school students that included a small number of students with disabilities. The content area science literature also includes Espin et al. (2013) in which generally strong in magnitude findings were reported for vocabulary matching for an older sample of demographically different students. Moreover, the content literature includes evidence of moderate to strong criterion validity correlations with national- and state-level measures in social studies (Espin Busch, Shin, & Kruschwitz, 2001; Mooney, McCarter, Russo, et al., 2013, 2014) across upper elementary and middle school grades. In addition, present data demonstrate moderate to strong correlations to relevant criterion in ELA and reading comprehension across Grades 4 to 6.
Second, and we believe more importantly, a growing literature is highlighting that, when directly compared, academic language-oriented predictor measures such as CCM and vocabulary matching have consistently provided stronger in magnitude correlations with meaningful criterion than those of other predictors. For example, and likely expectedly, Mooney, McCarter, Schraven, et al. (2013) reported findings for social studies content wherein a measure of social studies vocabulary (i.e., vocabulary matching) produced statistically significantly stronger concurrent validity comparisons than did a reading measure (i.e., traditional maze) and a writing measure (written expression-Curriculum Based Measurement) in content criterion comparisons. More recently, Lastrapes and Mooney (2017), Mooney and Lastrapes (2016), and Mooney et al. (2016) have compared CCM to ORF, traditional maze, sentence verification technique, and written retell for science content, both in concurrent and predictive criterion validity circumstances, and found that the academic vocabulary measure provides descriptively stronger correlation coefficients than other predictors in most all cases. So there is technical adequacy justification, albeit predominantly at the static score level (L. S. Fuchs, 2004).
Finally, the present data strengthened the argument for academic vocabulary measure use by demonstrating predictive, growth, and implementation capabilities. While not powerfully accurate, fall and winter CCM benchmark scores were able to distinguish participants who “passed” and “failed” the state test. CCM benchmark scores also displayed evidence that individual and collective student learning can be captured through repeated administration of timed probes. The diagnostic accuracy and growth analyses came following teacher-directed administration of the probes across professionals, grades, schools, and benchmark periods. Teachers and not solely researchers demonstrated the ability to carry out a year-long assessment implementation process. Collectively, these data suggest that academic vocabulary general outcome measures can be used to systematically document student learning in response to intervention frameworks similar to the more recognizable reading and math measures. Systematic research in the design, implementation, and evaluation of response to intervention frameworks for older age students and in science and social studies content areas is warranted given the variability in system implementation (e.g., D. Fuchs & Fuchs, 2017) and the differences necessary for older students (e.g., L. S. Fuchs, Fuchs, & Compton, 2010).
With the general outcome measurement reading and math history of inquiry as guidepost to future investigation in the science and social studies content areas, the notion of systematically designing and completing conceptual replications in support of generalizability suggested by Coyne, Cook, and Therrien (2016) can also guide development. While significantly more technical adequacy research remains necessary, systematic program inquiry in logistical feasibility and instructional effectiveness (Deno & Fuchs, 1987) by diverse researcher–practitioner groups are logical next steps.
In terms of logistical feasibility, we provided evidence that content and/or secondary school teachers can successfully direct benchmark data collection efforts with researcher guidance and support. Determining how best to collect, summarize, and use benchmark data will be beneficial to the hope of more widespread adoption of preventive assessment and instructional frameworks in the content areas and at the secondary school level. Research in the general outcome measurement for science and social studies content area, and particularly the academic vocabulary measure area remains researcher driven. For such assessments, and the potential data-based decision making and intervention processes that could be associated with their use, the involvement and leadership of practitioners is sorely needed. Diversity in the size and makeup of schools and districts would allow for development of a systematic implementation replications. Input from practitioners might first address how to best incorporate academic vocabulary assessment and intervention into science and social studies courses.
In terms of instructional effectiveness, L. S. Fuchs (2004) has made the point that curriculum-sampled general outcome measures better enable teachers to target skills—in this case, content—for remediation. Academic vocabulary is relevant remediation content given its status as academic currency (Alexander, n.d.), its centrality in classroom discussion and action (Nagy & Townsend, 2012), and its ability to distinguish between stronger- and weaker-performing students. Academic vocabulary use also seems timely as proponents of the Common Core State Standards (National Governors Association Center for Best Practices & Council of Chief State School Officers, 2010) in ELA advocate for instructional strategies such as close reading (Brown & Kappes, 2012) to be used with informational text. Few studies (e.g., Boudreaux-Johnson, Mooney, & Lastrapes, 2017) have incorporated academic vocabulary measures as indicators of effectiveness. A body of systematic, conceptual replications would inform stakeholder efforts to improve instruction and/or achievement across the assessments tiers that comprise the suggested or implemented response to intervention frameworks in secondary schools. A starting place may be the application of data-based individualization to students with disabilities in inclusive science and social studies courses.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The research was supported by grant funding from the Louisiana State University College of Human Sciences and Education’s Dean’s Circle and Louisiana Board of Regents, Louisiana Systemic Initiatives Program.
