Abstract
The authors used a large data set (N = 1,011,549) to examine literacy growth over a single school year comparing general education (GenEd) students to three high-risk subgroups: English language learners (ELL), those with a specific learning disability (LD), and those identified as both LD and ELL (LD-ELL) in students in Grades 3–10. The authors were particularly interested in whether variability existed between initial status and the growth trajectories of the three high-risk groups on measures of spelling, fluency, and reading comprehension across the school year and whether this variability was differentiated because of socioeconomic status (SES) as defined by free and reduced lunch (FRL) status. Results indicate that all high-risk groups began the year at substantially lower levels than their GenEd peers, with the largest differences seen between the LD-ELL students and the other subgroups. Further results suggest that students who are in the high-risk subgroups and also qualify for FRL perform significantly worse than their peers in similar risk status groups who do not qualify for FRL, demonstrating the significant impact of SES on academic outcomes for all groups.
One fifth of the school-age population speaks a language other than English (U.S. Department of Education, 2003); estimates predict that by 2030 approximately 40% of students will be from homes in which English is not their first language. Of students who come from non-English-speaking home environments, 77% speak Spanish as their first language (Hemphill & Vanneman, 2011); this population is currently the fastest growing in the United States. Often, students who do not speak English as their first language are identified as English language learners (ELL). This identification or classification varies widely from state to state, but in most cases an ELL student is someone whose English language skills are limited to the point where, in general, he or she cannot benefit from general education instruction without special support in English language development (Ortiz & Kushner, 1997). According to the Center on Educational Policy (2009), the performance gap between ELL and non-ELL students is significant and persistent nationwide, with particular discrepancies in literacy achievement. Even when ELL children are in school from early elementary school, only 7% of fourth graders and 3% of eighth graders score at or above proficient on reading assessments, as compared to 37% and 35% of native English speakers (National Center for Education Statistics [NCES], 2012). Although these national statistics are informative, they often do not compare subgroups beyond ELL and non-ELL to include demographic data such as socioeconomic status (SES) and identification of learning disabilities (LD), nor do they take into account school- and classroom-level effects. The current article attempts to do just this. By utilizing a large data set, we investigate literacy growth in ELL and non-ELL populations taking into account SES and identification for LD and school- and classroom-level effects, in an effort to provide a more thorough and perhaps a more meaningful description of literacy growth in ELL students.
An LD-ELL student is someone who is identified as ELL because of home language practices and/or assessments given by a district and who has also been identified as impaired according to a district definition of LD. The prevalence of LD-ELL students in the nation’s schools has not always been reported accurately, as typically LD-ELL students have been grouped with LD students in district and statewide databases. Organizing district- and state-level data in this way has made it difficult to determine the prevalence of LD-ELL students as well as examine academic outcomes for this population. Researchers have stressed the importance of disaggregating special education descriptive data by grade level and language status (Artiles, Rueda, Salazar, & Higareda, 2002); however, studies that report data in this way are rare. Even more rare are evaluations of literacy growth of the LD-ELL group of students with comparisons to other subgroups of students such as ELL only, LD only, and general education (GenEd) students.
Recent data collected by the Office of English Language Acquisition suggest that there are fewer ELL students receiving special services than would be expected (U.S. Department of Education, 2003; Zehler, Fleischman, Hopstock, Pendzick, & Stephenson, 2003). Overall, the data show that 9.9% of ELL students receive special education services compared to 13.3% enrolled in schools (U.S. Department of Education, Office of Special Education and Rehabilitative Services, 2002). National statistics have long shown a disproportionate representation of minorities in special education programs, both underrepresentation and overrepresentation, depending on the special education category. Many factors contribute to this disproportionate representation, including inappropriate language and educational assessments, lack of cultural and linguistic knowledge, cultural differences (Ortiz & Maldonado-Colon, 1986; Salend, Duhaney, & Montgomery, 2002), language abilities of the students (Gunderson & Siegel, 2001), and a lack of language support systems for students who come to school speaking a language other than English (Artiles et al., 2002; Artiles, Rueda, Salazar, & Higareda, 2005; Artiles & Trent, 1994). ELLs tend to be underrepresented for special education services for LD in early elementary school and overrepresented after fifth grade (Artiles et al., 2002). A more recent study found that language-minority students were underidentified for special education in kindergarten and first grade but overrepresented by third grade (Samson & Lesaux, 2009). One main reason for this is confusion between language proficiency and actual cognitive impairment in content areas such as reading, writing, and math. Because of the heterogeneous nature of the ELL population, it is difficult to gauge “normal” second language acquisition, and unfortunately a lack of English proficiency has sometimes been interpreted as limited intelligence or disability (Klingner, Artiles, & Mendez Barletta, 2006). When students are younger, teachers may hesitate to refer ELL students for special education services because it is difficult to determine if the learning problems students are having are the result of second language acquisition or cognitive impairment. As students get older and as their proficiency in the English language increases, if cognitive impairment is present, manifested in the widening of achievement in content area skills between ELL students and their monolingual and other bilingual peers, they are more likely to be recommended for special services. Limbos and Geva (2001) found that teachers were better able to identify monolingual English-speaking children as at risk for reading problems than they were able to identify ELL at risk for reading difficulties in first grade. However, as students got older, and presumably learned more English, their ability to identify monolingual and ELL students having reading difficulties was the same.
Existing research suggests that placement in special education services is correlated with language proficiency levels, with ELL students with lower levels of English proficiency being placed in special education services at higher rates than their more proficient ELL peers (Artiles et al., 2002). In addition, extant data show that the majority of ELL students who are struggling academically have reading difficulties (Klingner et al., 2006), and approximately 66% of ELL students who receive special services are classified as LD (Zehler et al., 2003). This is not to assert that all ELL and LD children have reading problems, but researchers estimate that approximately 80% of the population identified as LD have reading difficulties (Snow, Burns, & Griffin, 1998).
From a practical standpoint, it is often difficult to differentiate if limited English proficiency is interfering with academic growth or masking an LD, making the identification of these students complex. At the policy level, as a recognition of consistent disproportionate prevalence of minority students in special education, the 2004 Individuals with Disabilities Education Improvement Act (IDEIA) definition of LD clearly states that it does not include a “learning problem that is primarily a result of . . . environmental, cultural or economic disadvantage” or being “limited English proficient.” Given cultural considerations and the fact language-minority students are more likely to come from low-income backgrounds (Capps et al., 2005), there is an inherent difficulty for special education practitioners and administrators when identifying ELL students with LD, especially in the early grades when discriminating between language development and cognitive impairment can be difficult.
Risk Profile of LD-ELL Students
Longitudinal and cross-sectional data suggest that students who are identified as both ELL and LD have a high-risk profile. Learning English as a second language places students at high risk for poor language skills (Tabors, Paez, & Lopez, 2003), placement in special education (Bennet, 1999; Oviedo & Gonzalez, 1999), academic failure (Casanova, 1992), and dropping out of school (NCES, 2002). At the same time, students who are identified as LD have much lower levels of literacy achievement in secondary school, with 65% performing below basic on reading in eighth grade (NCES, 2007), and have higher rates of school dropout, typically 2 to 3 times higher than those of their peers (Morrison & Cosden, 1997; U.S. General Accounting Office, 2003; Young & Browning, 2005). In addition, LD students are less likely to attend postsecondary education, entering higher education at only one tenth of the rate of their peers (Wagner, Newman, Cameto, Garza, & Levine, 2005; Young & Browning, 2005). Understanding the growth trajectory of students who are identified as both LD and ELL in critical literacy areas is essential to begin to build our capacity to serve this high-risk population.
The literature on growth for at-risk students offers two contradictory theories: the cumulative growth model, better known as the Matthew effect, and the compensatory growth model. The Matthew effect as it applies to reading (Stanovich, 1986) postulates that students who enter school at a lower level, over time, are unable to catch up to their higher achieving peers. That is, students who enter school at a higher advantage tend to grow at more advanced speeds than those who are more disadvantaged, allowing for a widening of the achievement gap between the two groups. The Matthew effect is often used as an argument for concentrating the limited resources schools have into early identification and intervention for students who are struggling in the early grades in reading-related skills. The compensatory growth model argues that children who enter school at lower levels will initially grow at higher rates than their peers who entered with higher reading skills, allowing them to begin to catch up with their higher achieving peers over time, or at the very least the discrepancy between the more high-risk and not at-risk students remains equivalent over time (Aunola, Leskinen, Onatsu-Arvilommi, & Nurmi, 2002).
Research investigating LD students in the context of these two different theories has produced mixed results. Some studies have shown that in areas such as word reading and comprehension discrepancies between LD and GenEd students increase over time (McKinney & Feagans, 1984; McNamara, Scissons, & Gutknetch, 2011); however, two other studies found that the reading achievement gap did not increase over time between LD and GenEd students (Baker, Decker, & DeFries, 1984; Scarborough & Parker, 2003). In theory, students who are identified as LD should be receiving intensive individualized instruction; therefore, one would assume there is a trend toward a decrease in the discrepancy between LD and GenEd students. However, perhaps because interventions are not intensive enough or as effective as we might think, schools may not be successful in closing the achievement gap. Morgan and colleagues investigated Matthew effects with Hispanic students and found that this population enters school with much lower levels of reading skills and grow at a slower rate than their Caucasian peers over time (Morgan, Farkas, & Hibel, 2008), indicating some support for the Matthew effect among the ELL population. In addition, the above-described national reports show that large discrepancies exist between ELL and non-ELL students in reading-related tasks. To our knowledge there have not been any studies that have investigated literacy growth trajectories specifically for the LD-ELL population compared to ELL, LD, and GenEd students across grade levels.
The Present Study
Districts are increasingly identifying ELLs for special services (Artiles & Ortiz, 2002). This population of students, given its unprecedented growth across the country and the unique learning challenges they bring to classrooms, is a group that warrants further study to enable a deeper understanding of developmental trends in key literacy areas. To our knowledge, there have been very few studies, if any, that have been able to evaluate literacy growth with data that are categorized in such a way that LD-ELL children can be separated from LD students. Therefore, the present study is an exploratory study to examine literacy development across a school year in high-risk subgroups, including LD-ELL students between Grades 3 and 10. There are important theoretical and practical implications for investigating growth in these high-risk subgroups. Theoretically, examining growth rates allows the field to better understand the development of literacy in these subgroups, which has potential to lead to more targeted, effective intervention. Practically speaking, increasing discrepancies between the high-risk subgroups and the GenEd students would highlight the need for more intensive individualized intervention. Contrasting the LD-ELL students to other high-risk groups may give the field better information on how to best serve the needs of this particular subgroup of students, who are often treated with interventions developed for monolingual students with LD, with little consideration for their ELL status and the unique needs of this population of students.
We aimed to compare initial status and growth rates, across a single school year, of ELL students, LD-ELL students, LD students, and GenEd students in three essential literacy areas (reading comprehension, text reading efficiency, and spelling). This study is unique in that it includes a very large sample size (N = 1,011,549) across Grades 3 to 10, which allows us to adequately examine achievement trends across high-risk subgroups. In addition, we examine literacy growth of these subgroups when they are separated by eligibility for free and reduced lunch (FRL), an SES indicator. Examining students eligible for FRL and those who are not allows this study to investigate two important things. First, we are able to descriptively explore prevalence rates of the high-risk groups separated by their eligibility for FRL. Extant data suggest that children who come from lower SES backgrounds (Artiles et al., 2005; Blair & Scott, 2002) and language minorities (Shifrer, Muller, & Callahan, 2011) and Hispanics who come from lower income backgrounds (Shifrer et al., 2011) are identified at a higher rate as LD. In addition, including FRL status in our analyses of literacy growth allows us to examine if differences between groups are more or less evident as the result of SES. A recent study by Kieffer (2011) demonstrated that students who come from similar SES backgrounds, regardless of ELL status, were similar in their risk profiles for reading failure, with low-income monolingual and low-income language-minority students having significantly elevated risk profiles. In addition, the study showed that longitudinally English reading scores of language-minority learners who come from low-income backgrounds are similar to those of their monolingual peers who come from similar economic backgrounds. This indicates that it may not be language-minority status but instead SES that predicts reading outcomes.
Research Questions
In this exploratory study, we analyzed data from the Florida Assessments for Instruction in Reading (FAIR; Florida Department of Education, 2009a, 2009b) using the Reading Comprehension, Text Reading Efficiency, and Word Analysis subtests in Grades 3 and 10. Our research questions are as follows: (a) What do the descriptive data show us about the prevalence of LD-ELL, ELL, and LD students across grade levels? Does the prevalence change when considering students who are eligible for FRL and those we are not? (b) Does variability exist in initial status and growth on the three subtests across grade levels? and (c) Do differences in fall performance, slope, and spring performance on the three subtests exists across the three high-risk subgroups (ELL, LD, and LD-ELL)? Do such differences become more pronounced when we consider SES as a moderator of the subgroups? How do the growth trajectories of the high-risk subgroups compare to those of typically developing or GenEd students?
Method
Participants and Procedures
Participants were 1,011,549 students across elementary (3rd grade n = 165,582, 4th grade n = 138,626, 5th grade n = 137,494), middle (6th grade n = 112,676, 7th grade n = 114,102, 8th grade n = 110,200), and high schools (9th grade n = 123,124, 10th grade n = 109,745). See Table 1 for sample size broken down by classification status and FRL eligibility. Students were enrolled in 40,404 classes and 1,112 schools across 67 districts in Florida. The demographic composition of students was as follows: 52% male, 52% White, 22% African American, 19% Latino, 4% Multicultural, 2% Asian, and less than 1% Native American. Just more than half of the sample, 57%, was eligible for FRL. Overall, 14.0% of participants were identified as qualifying for exceptional student education (ESE), 6.5% as limited English proficient (classified as ELL), 8.0% as specific learning disabled, and 0.6% as both LD and ELL (LD-ELL). Classification for ESE was determined at the school and district levels, using state-adopted criteria.
Frequency of Total Sample and Subgroups by Grade Level.
Note: ELL = English language learner; GenEd = general education; SLD = specific learning disability.
Data for this study were collected during the 2009–2010 academic school year. Students between Grade 3 and Grade 10 were administered the FAIR three times per year to monitor progress toward Florida reading standards. Assessment windows of 25 testing days for each assessment period are set by the Florida Department of Education in consultation with district superintendents. Fall testing occurred between the end of August 2009 and the beginning of October 2009, the winter window was between the beginning of December 2009 and the end of January 2010, and the spring window was between the end of March 2010 and the middle of May 2010. Those students who were required to take the FAIR assessment fell into three categories: (a) students who had never taken the FCAT in Grades 3–10 who were new to Florida public schools, (b) students in Grades 4–10 who previously failed the FCAT (Level 1 or 2), and (c) students who passed with a very low score (low Level 3). Most districts administered the Reading Comprehension screen to students who had passed the FCAT (Level 4 and 5), mainly to obtain the Lexile® score associated with the Reading Comprehension screen.
The FAIR in Grades 3–10 is a computer assessment, which is typically administered by the reading/language arts teacher in computer labs or classrooms. Students complete all three subtests in one class period. The population of students taking the FAIR Reading Comprehension task reflects the state demographic and achievement distribution. Once students complete FAIR, teachers may download reports on student results and access instructional recommendations from the Progress Monitoring and Reporting Network (PMRN), a website and database of deidentified student information on reading performance maintained by the state of Florida. Data for these analyses were obtained from the PMRN. A FAIR instructional toolkit is available for teachers, with professional development provided on its use.
Measures
The FAIR measures utilized for this study were (a) the Reading Comprehension screen, (b) the Text Reading Efficiency Task (hereafter called Fluency), and (c) the Spelling Task. A general description of FAIR in grades kindergarten through 10 can be found in Foorman, Torgesen, Crawford, and Petscher (2009). Background on the development and psychometrics of FAIR is provided in the FAIR Technical Manual (Florida Department of Education, 2009a) and a report from the Buros Testing Company (2010). All three tasks are computer administered and require a live Internet connection and headphones. Directions are provided for each task, along with practice items and feedback. Currently FAIR’s reading comprehension screen is utilized to predict the Florida Comprehensive Assessment Test (FCAT; Florida Department of Education, 2008). All three measures from the FAIR were administered in the fall (approximately September to October), winter (approximately November to January), and spring (approximately April to May).
Reading comprehension
The FAIR Reading Comprehension screen is a computer-adaptive test consisting of literary and informational passages followed by seven to nine multiple-choice questions. Students read from one to three passages to receive an estimate of the probability that a student will pass the FCAT at the end of the year. Students are exposed to the passages until they approach a reliable estimate of their ability (i.e., at α = .90) or until they have taken three passages. Scores provided include a Lexile® score, percentile rank and standard scores, a developmental ability score, estimates of relative ability within the four cluster categories assessed by the FCAT (words and phrases in context; main idea, plot, and author’s purpose; comparison and cause/effect; and reference and research), and a probability score that reflects the student’s likelihood of passing the FCAT. This probability score type not only provides teachers with an estimate of future performance but also serves as a criterion for students to either stop receiving tasks or continue to move on to other diagnostic literacy assessments. Students with a probability greater than or equal to .85 are not required to take the Fluency and Spelling tasks, whereas those scoring less than .85 must take the two additional assessments.
Developmental ability scores for the reading comprehension task (M = 500, SD = 100, range = 200–800) were used in this study as they represent vertically scaled estimates that are appropriate for measuring growth (Kolen & Brennan, 2004). From the norming sample, the data were scaled using sixth grade as the referent group; thus, a score of 500 is an indication of performance at a sixth grade level, scores less than 500 represent ability at a lower grade level, and scores greater than 500 represent ability at a higher grade level. Generic estimates of reliability from item response theory (IRT) range from .90 to .92 in Grades 3–10. Correlations between the FAIR reading comprehension score and the FCAT range from .62 in ninth grade to .74 in fifth grade (Florida Department of Education, 2010).
Maze Task (fluency)
The Maze Task is a text reading efficiency (i.e., fluency) measure that is used in FAIR to determine whether students with a relatively low probability of success on the FCAT (i.e., < .85), as determined by the FAIR Reading Comprehension screen, have difficulties with fundamental reading skills such as accuracy and fluency or basic text comprehension. Students are presented with one informational and one narrative text and are asked to read a passage where every seventh word is deleted, and students are required to choose the correct word from three options to fit the sentence. The adjusted fluency score on the Maze Task is based on the average number of maze items answered correctly in 3 minutes on two passages and is prorated for time for students who complete the task in less than 3 minutes. Because this is a timed task, IRT scores were not generated; however, to control for passage difficulty effects, which are known to affect fluency scores (Francis et al., 2008; Petscher & Kim, 2011), scores were equated using equipercentile equating. Also provided are standard scores and percentile ranks based on a comparison of students’ performance to that of other students in Florida at the same grade level. In this study we used the adjusted fluency score, which is sensitive to growth.
Corrected parallel-form reliability ranges from .77 in Grade 7 to .90 in Grade 10. Construct validity of the scores from the Maze/Fluency task was estimated as a function of correlations between Maze/Fluency and Florida Oral Reading Fluency scores in Grades 6–10, with associations ranging from .51 in ninth grade to .64 in eighth grade (Florida Department of Education, 2009a).
Word Analysis Task (spelling)
The Word Analysis Task is a computer-adaptive test of spelling that assesses students’ knowledge of the phonological, orthographic, and morphological information necessary for accurate representations of English orthography. Students spell five words at their grade level before the system becomes adaptive, and they receive harder or easier words, depending on their ability, up to a maximum of 30 words. The average number of words spelled is approximately 12. Similar to the Maze/Fluency task, the only students required to take Word Analysis are those with a probability of success on the FCAT less than .85. This test can provide information about the strength of their knowledge of written words, which is fundamental to accurate identification of words in text. Because Word Analysis is part of a reading assessment system, we have focused on its relation to sentence-level and text-level reading comprehension rather than its relation to other spelling measures. Also, because we designed the Word Analysis Task to be computer adaptive to increase precision at the extremes of the spelling distribution, its correlation with group-administered, standardized spelling measures is less meaningful.
The Word Analysis Task provides developmental ability scores, percentile ranks, and standard scores. Similar to the reading comprehension ability score, the word analysis ability score was vertically scaled with the same range and distribution as the reading comprehension ability score. RT generic estimates of reliability range from .90 in Grades 4–6 to .95 in Grade 8. Correlations between Word Analysis and FCAT and between Word Analysis and Fluency/Maze are moderate, ranging between .44 and .54 in the first case and between .36 and .56 in the second. In an analysis of Word Analysis by FCAT levels, the differences at the two lowest levels of proficiency—Levels 1 and 2—were significant and greater than the differences between any other levels (Florida Department of Education, 2010). For the purposes of this article, we refer to the FAIR Word Analysis Task as the spelling task.
Data Analysis
Students’ change in reading comprehension, spelling, and text reading efficiency was measured using a five-level (time, students, classes, schools, and districts) linear growth model to correctly partition the variance in scores across the different clustering units. Because the sample size for the growth model analyses was rather large and was drawn from an extant database (i.e., the PMRN), it was important to account for the shared environments of classroom instruction, schools, and districts. Ignoring such potentially large sources of variances would result in biased estimates of the standard errors around the intercept and slope means, as well as in the values and standard errors of the random effects for lower level units (e.g., students and classes).
The data were first modeled in an unconditional random effects linear model to estimate the average achievement across all individuals as well as random effects (i.e., variances of intercept and slope and covariances between intercept and slope). This step was necessary to understand for which outcomes and grade levels students may significantly differ from the grand mean initial status and growth values. The data in this study were centered at the first time point of the year (i.e., September), meaning that the intercept in the model represented the grand mean estimate of performance at the beginning of the academic year. Intraclass correlations were estimated for the students, classes, schools, and districts to evaluate at relatively larger sources of variance. Intercept and slope parameters that demonstrated small levels of variance (i.e., < 2%) at any clustering unit were fixed and tested against the fully random effects model. Akaike’s information criterion (AIC) and sample-size-adjusted Bayesian information criterion (BIC) compared the two specifications, with lower values representing a better fitting model.
For the grades and parameters that demonstrated significant variance within a specific outcome, a conditional model was specified that included the dummy-coded indicators of ELL status, LD status, and LD-ELL status conditional on whether the student was FRL eligible. In total, eight different demographic subgroups were created: (a) general education students who were not eligible for FRL (GenEd-NFRL), (b) general education students eligible for FRL (GenEd-FRL), (c) LD students not eligible for FRL (LD-NFRL), (d) LD students eligible for FRL (LD-FRL), (e) ELL students not eligible for FRL (ELL-NFRL), (f) ELL students eligible for FRL (ELL-FRL), (g) LD students who were also identified as ELL not eligible for FRL (LD-ELL-NFRL), and (h) LD students who were also identified as ELL and eligible for ELL (LD-ELL-FRL). Nonsignificant variance components at any portion of the modeling process were fixed to achieve greater power in model estimation. Following the estimation of the conditional model, a pseudo-R2 statistic was estimated. This value provided information as to how much of the initial variance in the intercept and slope coefficients was explained by the demographic characteristics in the conditional model. Such evidence is often useful for understanding where predictors are most valuable for explaining individual differences in the selected outcome.
Results
Prevalence of Risk Status Groups Across Grade Levels
Our first research question simply calls for a descriptive investigation of the data to determine the prevalence of the risk status groups across grade level when FRL status is considered. Figure 1 illustrates these trends across the grades levels for our six risk status groups: ELL, LD, and LD-ELL, eligible for FRL and those who were not. The trends demonstrate a clear distinction between students FRL eligible and those who are not eligible across risk status groups. For example, across all grade levels, the average proportion of FRL eligible students identified as ELL was 9.25%, whereas it was only 2.06% on average for students who were not FRL eligible. Likewise, within the LD, FRL-eligible group, 9.49% of students were identified on average across grade levels as LD; this average drops for students not FRL eligible to 6.27%. It can be seen in Figure 1 that students identified as LD-ELL were a very small proportion of the students; however, differences remain between FRL-eligible students (1.0%) and those who were not FRL eligible (0.13%). An interesting finding from descriptively investigating the prevalence numbers is that the number of ELL student identified as LD did not increase across the grade levels, as has been found in previous research (Artiles et al., 2002).

Rates of identification for high-risk subgroups separated by free and reduced lunch status.
A chi-square analysis was run to test the relationship between FRL status and each of the other subgroups (i.e., LD, LD-ELL, and ELL) to examine the extent to which identification rates of the subgroups varied across FRL eligibility. In addition, a phi coefficient was calculated, which served as an effect size measure of the degree of association between each 2 × 2 comparison of FRL status (eligible or not) and subgroup identification (identified or not). Phi values of .10 are considered small effects, with .30 as moderate and .50 as large. The chi-square analyses across all grades were statistically significant (p < .001); however, Figure 2 displays that virtually no association existed between FRL and LD (phi range = .04 to .07) or FRL and LD-ELL status (phi range = .04 to .06). A small effect was observed between FRL and ELL status (phi range = .13 to .17).

Effect sizes for prevalence of risk status.
Analysis of Missing Data
Missing data were evaluated for each measure, assessment period, and grade to determine if any discernable pattern emerged. The missing data rate ranges for the measures by grade were as follows: Grade 3 (1.5% on fall reading comprehension to 8.1% on fall spelling), Grade 4 (1.8% on fall reading comprehension to 15.2% on winter spelling), Grade 5 (1.9% on fall reading comprehension to 17.1% on winter spelling), Grade 6 (3.9% on spring reading comprehension to 21.7% on fall spelling), Grade 7 (4.4% on spring reading comprehension to 21.8% on fall spelling), Grade 8 (4.5% on spring reading comprehension to 23.3% on fall spelling), Grade 9 (8.7% on spring reading comprehension to 29.4% on fall spelling), and Grade 10 (9.6% on spring reading comprehension to 29.9% on fall spelling). The relatively large ranges in missing data were largely indicative of a generalized pattern; that is, variables, that presented with large missing data (e.g., spelling) were tasks that were not required to be taken by all students in the FAIR system.
As previously described, only individuals with a lower probability passing of the FCAT were required to take the fluency and spelling tasks. Evidence from this FAIR sample demonstrated that, on average, 30% of students in Grades 3–10 were above the .85 probability target on the reading comprehension task. Of these students meeting the criteria, and thus not required to take spelling to fluency, the data revealed that 24% of the successful students contained the FAIR system to complete all three tasks. As such, the reason for missingness is the operational procedure of the FAIR and is unrelated to the scores. Through this mechanism, the data may be assumed to be partly missing by design, which is a qualifying component of data to be missing at random (Enders, 2010). To correct for the unbalanced design and the potential bias in the estimation of the parameters in the mixed model and hierarchical regression analyses, multiple imputation was conducted using SAS PROC MI with a Markov chain Monte Carlo estimation and 10 imputations. Imputations were conducted within grade level using the minimum and maximum values of the sample distribution for each variable by time point as an imputation constraint on the model. When comparing the original and imputed data files, the largest average standardized difference between scores across grades was d = 0.03, for fall spelling.
Reporting of Linear Mixed Model Results
Our second and third research questions were concerned with the extent to which variability in achievement existed across the measures of spelling, text reading efficiency, and reading comprehension. In advance of reporting the results for Research Questions 2 and 3, it is useful to orient the reader to the presentation of the forthcoming findings. First, we being by providing the intraclass correlations (ICCs) in Table 2 from unconditional linear mixed models, which were computed from the variance components of each grade and outcome-based model. The actual variances, covariances, and samples sizes at each level (i.e., student, classroom, school, and district) are provided in Appendices A–F so that the reader may appropriately contextualize the ICCs according to the number of units per level.
Initial Status and Slope Intraclass Correlations for Each Clustering Unit by Outcome and Grade Level.
Following this discussion, we report the fixed effects from the conditional linear mixed model where growth and initial status were modeled as a function of the dummy-coded subgroups. Because the conditional models of linear growth were modeled across eight different subgroups (GenEd, LD, LD-ELL, and ELL with and without FRL), we have opted to report the most parsimonious elements of the growth models for describing individual differences, namely, the fitted initial estimated fall score in September, the fitted estimated slope, and the model-predicted final status score in April. This information appears in Tables 3 (reading comprehension), 4 (spelling), and 5 (fluency). Specific statistics regarding the model estimated coefficient (i.e., the grand mean and fitted deflection estimates), standard errors, model degrees of freedom, and specific p values are available in Appendices A–F. Ancillary statistics in the way of a pseudo-R2 are provided in Table 6, which, as mentioned previously, describes the proportion of variance in the intercept and slope, which was explained by the student subgroup characteristics at each level and grade, across the three outcomes.
Fixed Effects for Reading Comprehension.
Note: ELL = English language learners; FRL = free and reduced lunch; GenEd = general education; NFRL = not eligible for FRL; SLD = specific learning disability.
Fixed Effects for Spelling.
Fixed Effects for Fluency.
Note: ELL = English language learners; FRL = free and reduced lunch; GenEd = general education; NFRL = not eligible for FRL; SLD = specific learning disability.
Proportion of Pseudo-Variance Explained by the Demographic Covariates for Initial Status and Slope by Outcome and Grade.
Finally, Tables 7–9 include standardized effect sizes for the comparison of subgroup differences in initial status (fall) and final status (spring), for a priori selected comparisons. We were primarily interested in comparing the following three classes of subgroups: (a) comparing all language/LD subgroups to each other who were not eligible for FRL, (b) comparing all language/LD subgroups to each other who were eligible for FRL, and (c) comparing each of the four language/LD subgroups across FRL statuses (e.g., GenEd-NFRL to GenEd-FRL). Such comparisons at the fall and spring allowed for an evaluation of how much standardized change between the specified groups occurred across the year.
Effect Size Comparisons for Reading Comprehension by Grade at the Fall and the Spring.
Note: Spring comparisons are in parentheses. Values in bold represent comparisons made within language/LD group, across the non-FRL and FRL subgroups. ELL = English language learners; FRL = free and reduced lunch; GenEd = general education; NFRL = not eligible for FRL; SLD = specific learning disability.
Effect Size Comparisons Spelling by Grade at the Fall and the Spring.
Note: Spring comparisons are in parentheses. Values in bold represent comparisons made within language/LD group, across the non-FRL and FRL subgroups. ELL = English language learners; FRL = free and reduced lunch; GenEd = general education; NFRL = not eligible for FRL; SLD = specific learning disability.
Effect Size Comparisons for Text Reading Efficiency by Grade at the Fall and the Spring.
Note: Spring comparisons are in parentheses. Values in bold represent comparisons made within language/LD group, across the non-FRL and FRL subgroups. ELL = English language learners; FRL = free and reduced lunch; GenEd = general education; NFRL = not eligible for FRL; SLD = specific learning disability.
Variation in Literacy Growth
The ICCs for intercepts and slopes in Table 2 highlighted that most of the variance were distributed across students, followed by classrooms. A significant portion still existed between schools and districts; however, these were relatively small compared to the lower level units. It can be seen from Table 2 that the mixed model treated all clustering units as random for the intercepts and slopes. Although the ICCs may seem small in some cases, a model that treated a district or school effect as fixed did not fit as well as a fully random effects model using traditional measures of AIC and BIC to evaluate differences. On evaluation of the models, it was apparent that significant variation existed in student performance in the outcomes at the fall as well as in slopes. Subsequently, the conditional linear mixed model was fit with the subgroup covariates to explain differences in performance.
Individual Differences in Initial Status, Growth, and Final Status
Across the three literacy outcomes (Tables 3–5), the LD-ELL-FRL students systematically performed lower than all other subgroups in Grades 3–10. Conversely, GenEd students who were not eligible for FRL (GenEd-NFRL) were the highest ability subgroups. The remaining six subgroups demonstrated differential ability across outcomes and grades, which are more fully explicated below.
Reading comprehension
The September status comparison of reading across grades (Table 3) highlighted that GenEd-NFRL students were the highest achieving subgroup, followed by the GenEd-FRL students across the grade levels. These two groups were statistically differentiated from each other (p < .001); however, this is largely a function of the large sample sizes reported Table 1. Notwithstanding this, the mean difference across the grades between these two groups corresponded to an average standardized effect size difference of d = 0.55, with a range of differences from d = 0.47 to d = 0.64 for Grades 3–10 (Table 7).
The fixed effects in Table 3 demonstrate that the mean growth in reading comprehension per month was approximately similar between GenEd-NFRL and GenEd-FRL students. For example, in Grade 3, the NFRL students grew at a rate of 8 points per month, compared to 8.43 per month for the FRL students. This trend was consistent across the grades, and, as a result, the standardized mean difference between the groups at the end of the year did not appreciably change. To illustrate, Table 7 highlights that the mean effect size difference between GenEd-NFRL and FRL students in Grade 3 was d = 0.64 at the first assessment and d = 0.62 as the end of the year. Such a finding shows that despite a statistical difference in estimated slopes between groups at this grade level, the distinction did not result in any meaningful change in the standardized difference at the end-of-year assessment.
The LD-NFRL students in Grades 3 and 4 performed similarly to the ELL-NFRL in terms of both their September performance (approximately 280 and 300, respectively) and growth across the year (9 and 10 points per month, respectively). However, in the remaining grades this subgroup was the highest achieving at-risk group in both September and April. This was indicative of relatively smaller individual differences in subgroup growth as the grades increased. ELL-NFRL students across the grades started the year significantly lower than GenEd-NFRL, but only marginally better than their ELL-FRL counterparts. In Grade 3, for example, a d = 1.01 difference in September scores existed between the ELL-NFRL and GenEd-NFRL students, but only a d = 0.37 difference occurred with the ELL-FRL. The pattern across the grade levels for these comparisons revealed that although the standardized effect size increased for comparisons between GenEd and ELL students without FRL (i.e., d = 1.01 in Grade 3 to 1.75 in Grade 10), the gap between ELL-NFRL and ELL-FRL shrank (d = 0.64 in Grade 3 to 0.47 in Grade 10).
The LD-ELL subgroups, both those eligible and not eligible for FRL, were consistently the two lowest performing subgroups. These two sets of individuals were generally very close to each other across the grade levels in their initial mean reading comprehension scores (average d = 0.30, Table 7); however, the LD-ELL-NFRL students made stronger gains than their FRL counterparts in Grades 3, 6, 8, 9, and 10, leading to a widening of the gap in comprehension by the end of the year (average d = 0.38, Table 7).
By accounting for the eight subgroups, Table 6 reports the proportion of pseudo-variance explained in reading comprehension. It can be readily seen that little of the individual differences in growth are captured by these demographic characteristics; however, a large proportion of the initial status variance is explained. Most noticeable is that large individual differences at the class and school levels are explained, followed by districts. It is interesting to note that even though the student level is where most of the variance lies (Table 2), relatively little of this variance is explained by the student characteristics above what is explained at the class, school, and district levels.
Spelling
Students’ developmental patterns in spelling (Table 4) mimicked those of reading comprehension for the subgroups. The GenEd-NFRL students were consistently the highest achieving spellers, followed by the GenEd-FRL individuals. The standardized difference between these two sets varied d = 0.31 (Grade 9) to d = 0.44 (Grade 10). ELL-NFRL students consistently outperformed their FRL counterparts by an average d = 0.25 across the grade levels, and both groups had similar growth rates at each grade level that were not statistically distinguished from each other (average p = .11). The relatively stronger growth rate of the ELL-NFRL compared to the GenEd-FRL led to a closing of the gap between these two groups in the April spelling scores, such that the fall average d = 0.93 changed to average d = 0.80 in the spring. A similar trend was observed for the ELL-FRL students who reduced the gap between the GenEd-FRL from average d = 0.84 in September to average d = 0.74 in April.
The LD subgroups with FRL and NFRL were generally separated by an average d = 0.29 across the year, but their growth led to a slight closing of the gap with some of the GenEd students. For example, the LD-NFRL group began the year with an average d = 1.22 below the GenEd-NFRL, but ended the year with a difference of d = 1.18. Student who were LD-ELL were consistently the lowest performers in spelling, as they were in reading comprehension, whether or not they were eligible for FRL. The gaps for these groups compared with GenEd students did not generally shrink, and either remained stable or increased. The September difference between the LD-ELL-NFRL and GenEd-NFRL group was d = 1.53 on average but shrunk to d = 1.20 in April. This was replicated with the FRL group, where the gap was average d = 1.46 between the GenEd and LD-ELL groups and stayed at a mean d = 1.42 in April.
The amount of variance explained at each level for spelling (Table 6) reflected the pattern observed for comprehension. Most of the explained variance was at the class, school, and district levels of the intercept. Virtually none of the slope variances were explained by subgroup covariates, and only a small proportion of the spelling variance was accounted for at the student level.
Fluency
Finally, when considering fluency development (Table 5), the development demonstrated that GenEd students (both FRL and NFRL) consistently began the year with higher scores than the other groups, and also grew at a faster rate. Most notably, unlike the reading comprehension and spelling tasks, the gap between the risk groups and GenEd performance grew from Grades 3–10. Much of the trends observed for the other outcomes extended to fluency development as well. Groups tended to be more strongly differentiated in initial status rather than growth, and the relative rank order of group performances was maintained across grade levels. The ELL-NFRL students were the third highest achievers in Grades 3 through 6; however, in the remaining grades their mean fluency scores fell below the LD-FRL group, though not significantly different (average p = .135). The LD-NFRL students were among the higher achieving subgroups for fluency above Grade 5, but were less fluent than ELL-FRL in the younger grades. The pseudo-variance explained for this outcome was largest for the classroom, school, and district levels. Dissimilar from the other outcomes, a larger proportion of the variance in slopes was explained across multiple levels and grades (Table 6).
Discussion
The growing number of ELL students coupled with the persistent low achievement of a large portion of this population raise many policy-level concerns and questions related to instructional practices at a national level. Using a large data set, this article sought to determine differences in initial status and growth across the school year between particular high-risk subgroups compared to GenEd students across grade levels (3–10) in different areas of literacy, paying particular attention to ELL students identified as LD. Because of the nature of the data set, we were able to compare eight different subgroups: ELL, LD-ELL, LD, and GenEd, both students who were FRL eligible and those who were not FRL eligible. In the past, it has been difficult to compare these subgroups because of the necessity of such a large data set and the way in which these risk groups have been identified in existing large data sets with the LD-ELL students typically being combined with the LD students. The large data set also allowed us to account for the shared environments of classroom instruction, schools, and districts, alleviating potential bias resulting from the clustered nature of the data. In general, our findings show that the student level is where most of the variance lies; however, a relatively small amount of this variance is explained by the student characteristics above what is explained at the class, school, and district levels. Our first research question addressed prevalence rates of the high-risk subgroups across grade levels. Our second and third questions were related to growth in literacy across a single school year. In particular we were interested in examining if the data suggested a trend toward Matthew effects or a widening of the gap between the high-risk groups and their GenEd peers. Or, if in contrast, a trend toward the compensatory growth model, which supports the achievement gap getting smaller or at a minimum remaining equivalent over time.
This article has four important findings. First, there are substantial differences in prevalence rates between students eligible for FRL and those are not. This difference is particularly evident in ELL students, with significantly more FRL-ELL students than NFRL-ELL. Second, all subgroups grew across literacy measures, however, in general, there was little to no narrowing of the achievement gap between the high-risk subgroups and the GenEd subgroups. Third, across all grade levels and literacy skills, the LD-ELL subgroups (both FRL and not FRL) consistently performed below all other subgroups, indicating that this subgroup is at significant risk of reading failure. Fourth, across all subgroups, students eligible for FRL performed below their similarly identified peers on all literacy measures, indicating that SES plays a significant and critical role in literacy achievement across grade levels and classification status.
The descriptive data illustrate differences in risk status identification for students eligible for FRL and those not eligible. The most glaring discrepancy was for students in the ELL category. In the 2009–2010 school year, Florida reported that 8.7% of the student population was identified as ELL (Florida Department of Education, 2010); the data in this study demonstrate that students who come from low-income backgrounds and who are learning English as a second language are more likely to be identified as ELL than their peers who are not eligible for FRL. Similar to this, the state of Florida reported that 5.6% of the total student population as LD in the 2009–2010 school year (Schull & Winkler, 2011); however, our data indicate that nearly 10% of students eligible for FRL were identified as LD, whereas only 6% of their peers not eligible for FRL were identified. These findings concur with extant data that show SES is related to identification rates of both ELL and LD (Artiles et al., 2005; Blair & Scott, 2002); Artiles and colleagues found that larger proportions of low-SES ELLs were identified as LD across grade levels.
We were surprised that the descriptive data did not concur with existing data that suggest ELL students are identified as LD at a higher rate in upper grade levels, especially beyond fifth grade (Artiles et al., 2002). Our data show that the number of ELL students identified as LD remained consistent across the grade levels for both the students who were FRL eligible and those who were not. It is important to note that because these data were within one school year, it is possible that students who were once ELLs were regrouped in the LD category only because once they exist from ELL status they are a part of the GenEd subgroup, not the ELL subgroup. A more thorough analysis of rates of LD-ELL students would include all students who were historically identified as ELL. We could then consider prevalence rates among language-minority students, rather than ELL classification, as was done in a recent study by Samson and Lesaux (2009). In addition, it is important to note that we did not have data for kindergarten, first, and second grade students in this particular data set. Typically students are not identified as LD until third or fourth grade; therefore, it is possible that we would have seen an increase in LD-ELL if we had the earlier grades as a comparison. Samson and Lesaux found that language-minority students were underrepresented in special education in kindergarten and first grade, but by third grade they were overrepresented as compared to the general population. It is extremely important to explore these questions as underidentification of LD-ELL students in early elementary school has ramifications that are persistent throughout a student’s educational career. The early grades are when intensive instruction in critical literacy skills that have been shown to ameliorate reading risk such as basic reading skills, alphabet knowledge, phonological awareness, and decoding is traditionally provided. In addition, data demonstrate that it is increasingly difficult to remediate reading difficulties after third grade (Fletcher & Foorman, 1994; Lyon, 1985); therefore, early identification and early intervention with LD-ELL students are critical.
In general, findings indicate that for all literacy subtests, all subgroups grew across the year. Across grade levels and literacy subtests, our high-risk groups (LD-ELL, LD, and ELL both FRL and NFRL), started the year with lower scores on all three literacy measures. However, growth trends varied depending on literacy measure, with reading fluency showing a trend toward an increase in the achievement gap between the high-risk subgroups and their GenEd peers, indicating a potential Matthew effect. For comprehension, in the early grades (3–5), there was a slight decrease in the achievement gap; however, in the upper grades (6–10) the gap widened between the high-risk subgroups and the GenEd group. For spelling, in general, we found parallel growth between the subgroups; however, the LD-ELL group began the year at every grade level at a significantly lower level than the other subgroups in spelling. Below, we pay particular attention the LD-ELL students, as our data set allowed for us to investigate the growth of this subgroup separately from ELL and LD students.
The LD-ELL students begin the year at a clear disadvantage, even more so than children who are monolingual LD and children who are solely identified as ELL. Furthermore, a more severe discrepancy between achievement is evident for the LD-ELL-FRL group. Differential growth was also evident between the LD-ELL-FRL and the LD-ELL-NFRL groups in reading comprehension specifically, favoring the LD-ELL-NFRL subgroup, who grew a faster rate than the LD-ELL-FRL subgroup in many grades. For reading comprehension, the LD-ELL students grew a lower rate than the ELL children, but at a comparable rate to the LD students in all grades except Grade 3. Although growth was similar for these two groups, because the LD-ELL group started the year significantly lower than the LD in most grades, the achievement gap between the two high-risk groups remained consistent. Comparisons of growth rates between the high-risk groups were not as straightforward. Comparing LD-ELL students to GenEd students revealed a trend toward Matthew effects for reading comprehension, with this group not only starting the year at a significantly lower level than the GenEd students, but growing across the year at a lower rate with the exception of fifth grade. In contrast, at the lower grade levels, a closing of the achievement gap was seen for the ELL children and the GenEd students, as the ELL students grew at a faster rate across the school year. Results for the upper grade levels indicate that although the all high-risk groups made growth slightly higher than their GenEd peers, it was not enough to make significant changes in the achievement gap between the high-risk groups and the GenEd students.
Spelling findings largely mimicked comprehension results with the GenEd students (both FRL and NFRL) performing significantly better than the high-risk groups across grade levels at the beginning of the year and the two LD-ELL subgroups performing the lowest of all the high-risk subgroups. Although the LD subgroups showed a trend toward closing the achievement gap between themselves and their GenEd peers at many grade levels, over the course of the school year the gap between LD-ELL subgroups and their GenEd peers did not shrink and either remained stable or increased. Fluency results show that the gap in achievement between the GenEd students and the high-risk students grew over the course of the school year. Again, consistent with other findings, the LD-ELL subgroups had the lowest initial performance and growth across the school year.
Growth rates for spelling and fluency show that at some grade levels, significant differences between the LD-ELL and other high-risk groups existed; however, large effect sizes were not noted. An increase in the achievement gap between the GenEd and LD-ELL students was seen at almost every grade level for spelling and fluency, showing support for the Matthew effect for the LD-ELL group. Overall, results indicate that, as suspected, all the high-risk subgroups show significantly lower levels of performance when compared to the GenEd students across literacy domains.
Perhaps the most concerning group is the LD-ELL students, who in many cases perform significantly worse than the other high-risk groups. This finding is not surprising given extant literature that demonstrates similar trajectories for high-risk subgroups; however, it is concerning given that students identified in these high-risk subgroups should be receiving some sort of special services to enhance their academic learning. Identification of students who have LD within the ELL populations is difficult. It is not clear if the identification process may be affecting the type of students who make up this subgroup. Because IDEIA (2004) is clear in that identification of LD cannot be based on cultural or language-related issues, school personnel may hold more stringent criteria when identifying ELL students as LD, therefore creating a group that is more impaired than students in the LD category. In addition, these findings highlight the importance of considering measurement tools and expected criteria when assessing students to determine if they are responding to treatment they are provided. It may not be appropriate for LD-ELL to be expected to meet similar criteria as their ELL-only or LD-only peers when such clear discrepancies exist between the groups, with little to no narrowing of the achievement gap occurring across the grade levels.
Of additional concern is the clear distinction across all subgroups between students eligible for FRL and those not eligible for FRL. Even within the GenEd subgroup, significant differences in initial status and growth was evident between the FLR and NRLF groups. Identification as a member of the high-risk subgroup and coming from a lower SES background puts students at a clear disadvantage initially and very little to no “catching up” occurring across the school year.
Limitations
One very clear limitation of this study is that we do not know the process by which each student was identified for our high-risk groups. We relied on school records and reports for the identification of all high-risk subgroups. It is common to use school records of disability (e.g., Hosp & Reschly, 2002); however, because of the nature of our data set, we were unable to independently identify high-risk status. Given the difficulty inherent in identifying LD-ELL students, this is our most glaring limitation. However, the data indicate that these subgroups were performing, on average, significantly lower than their GenEd peers, which demonstrates that collectively they were high-risk groups. The data also support that the LD-ELL students were performing lower than the students identified only as ELL and those identified only as LD, which supports that this group of students is at even a higher risk than students identified only as ELL.
A second limitation is that we were limited to only data within one school year for this study. A longitudinal data set would allow us to examine growth trends for the high-risk subgroups over a longer period of time, allowing for more conclusions to be drawn related to discrepancy in achievement between the subgroups. The majority of studies investigating Matthew effects in at-risk learners have utilized longitudinal data sets that span more than one academic school year. Because data for more than one school year were not available to the authors, we were not able to conduct an investigation of growth across more than one school year.
Methodologically, the results pertaining to spelling and fluency outcomes are limited by the design effect. Because students who are of high ability on reading comprehension are not required to take the other tasks, the GenEd sample does not include the highest ability students.
Implications for Practice and Conclusions
This exploratory study of prominent high-risk subgroups demonstrates that, in general, these groups are performing at lower levels than their GenEd peers on essential literacy skills across Grades 3 to 10. It also demonstrates that over the course of a school year, high-risk groups do not demonstrate a significantly higher level of growth, allowing for them to catch up or close the gap between themselves and their GenEd peers. These findings point to the need for more targeted, intensive instruction in literacy to better service these particular high-risk populations. Very often, the literature points to the need for early intervention; however, the results from this study indicate that discrepancies between GenEd students and high-risk groups are persistent across grade levels and need to be addressed through targeted intervention at all grade levels. Close attention needs to be paid toward students who are identified as LD-ELL. The development of targeted intervention for these students has not been prevalent, with very few empirical studies investigating intervention effects for this particular population. The few intervention studies that are available are largely for elementary-age students. A review of the What Works Clearinghouse shows that there are no reading interventions listed for upper-grade-level ELL children, and the interventions listed for LD children have been investigated only with emergent readers. These children may require significant intervention with targeted instruction that meets language targets as well as specific direct instruction in literacy development to avoid an increasing discrepancy between themselves and other high-risk groups and GenEd students.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
