Abstract
Helping students engage with complex texts has been a longstanding challenge, though teachers have received little guidance about practices that help students in engaging with texts. This paper provides a range of empirical evidence about a tool designed to provide formative insight into text-focused teaching, which we used to reliably score more than 500 reading lessons in a large district. We describe the structure of the tool, its relationship to other measures of teaching, teacher and school attributes, and student outcomes. We then provide guidance to practitioners and researchers seeking to employ such a tool for teacher development.
Helping students engage productively and proficiently with complex texts has been a longstanding challenge in the United States (Gamse et al., 2008; Terry, 2021). Researchers, educators, and policy makers alike have proposed a range of solutions for addressing this challenge including more ambitious standards for student learning, as well as aligned curricular resources for classrooms (Polikoff, 2012, 2015). The texts students are asked to read—whether they are sufficiently “complex,” “leveled,” or align with students’ interests and background knowledge—have been foregrounded in these conversations (Amendum et al., 2018; Goldman & Lee, 2014). Texts alone, however, will not promote positive student outcomes. Teaching is instrumental in these efforts, though there has been less attention paid to supports for teachers that could ultimately help students. Teachers need more clarity about practices that help students engage with texts, along with tools that provide formative information about their use of such practices. Our focus here is the development of such a tool and the provision of validity evidence about its use across hundreds of elementary reading lessons in a large urban district.
Standardized, reliable observational rubrics can provide teachers with a common language for describing and ultimately improving teaching. By naming key instructional practices and delineating different levels of quality instantiation of those practices, teachers can have a clearer and more coherent sense of what to focus on during instructional interactions and what “good” looks like (Allen et al., 2011; Cohen, 2015; Hill & Grossman, 2013). To date, however, there are relatively few observational tools that focus squarely on how teachers support students in making text-based arguments and claims. Existing measures frequently observe generic teaching practices (Danielson, 2007; Pianta & Hamre, 2009). Among content-specific tools, a majority identify disciplinary practices in science and mathematics, rather than reading (Hill et al., 2008; Schoenfeld, 2018; Walkington & Marder, 2018). Rare exceptions include the Protocol for Language Arts Teaching Observation (PLATO; Grossman et al., 2013), used to observe disciplinary practices in reading and writing instruction (Grossman et al., 2014; Kane & Staiger, 2012), the Argument Rating Tool (Reznitskaya & Wilknson, 2021), which focuses on teacher facilitation of student talk, and the Early Language & Literacy Classroom Observation Pre-K Tool (ELLCO Pre-K) which focuses on broad features of the literacy environment in early childhood classrooms, including the characteristics of the books, the organization of the book area, and the approaches to book reading (Smith, et al., 2008). The majority of other research tools developed to analyze reading comprehension instruction have focused on more micro-level classroom interactions, including the proportion of teachers’ and students’ fine-grained discourse moves used during text-based discussion (Connor et al., 2020; Dwyer et al., 2016; Gámez & Lesaux, 2015). While such tools are helpful for characterizing the discourse environment, they do not include rubrics with different levels of quality to help teachers in improvement efforts. Fine-grained, line by line coding of discourse may have value for research but is less practically useful for school or district improvement efforts.
In the absence of readily available, systematically researched tools for supporting text-focused instruction, schools have largely relied on open-source tools created by a range of educational support organizations. Among the most prominent among these are the Instructional Practice Guides (IPGs), coaching rubrics for elementary reading. Developed in 2012 by Student Achievement Partners (SAP), a non-profit organization founded by some of the authors of the Common Core State Standards and dedicated to helping teachers shift instruction in ways that are aligned with ambitious goals for students (achievethecore.org), the IPGs are available for free on SAP’s website and can be modified based on context-specific goals. The IPGs are downloaded approximately 100,000 times each year, and SAP has worked with scores of organizations, schools, and districts to support implementation of IPG coaching for teachers. Despite this widespread use, the schools and districts using the IPGs—or other tools like them—have no way of systematically tracking instructional improvements from IPG coaching because there is limited evidence about the tools’ measurement characteristics.
Rather than having researcher-developed observation instruments with robust psychometric properties exist at a distance from the practitioner-oriented tools used in schools across the country, we would be well-served to work together to develop a suite of tools that could be used in tandem to better support teachers in helping students engage with sophisticated texts. In this paper, we provide empirical validity evidence about a reading-focused observational tool, the Instructional Practice Research Tool for English Language Arts (IPRT-ELA) that we developed in concert with SAP to complement the practices featured in the IPGs. Over the course of several years, we worked together to substantially modify their coaching protocol to create a reliable observational protocol for analyzing reading teaching. After piloting and refining the tool, we used the IPRT-ELA rubrics, coupled with the content-generic CLASS rubrics, to score more than 500 reading lessons in a large, urban school district.
We describe the structure of the tool, the degree to which it can be used to measure text-focused practices consistently, its relationship to content-generic observational measures of teaching quality, as well as the relationship between scores and student achievement and observable teacher and school attributes (Steinberg & Garrett, 2016; Whitehurst et al., 2014). In doing so, we situate this tool in the broader landscape of supports for teachers and provide guidance to practitioners, researchers, and policymakers seeking to employ such a tool in conjunction with coaching protocols like the IPGs. We answer four research questions to help build a validity argument for the IPRT-ELA, which we list here along with our hypotheses for each: 1. What does text-focused instruction look like in upper elementary classrooms of a large school district? How much does instructional practice vary within a teacher’s lessons? How much does practice vary within and across schools?
We hypothesize that most of the variation in text-focused instruction will be among lessons within individual teachers. We also hypothesize that most of the remaining variation would be across teachers, within a school, and the smallest amount of variation would be across schools. 2. What is the relationship between text-focused measures of instructional quality and high and low stakes, content-generic measures of instructional quality?
We hypothesize that text-focused instruction will be modestly correlated with other observational measures of instructional quality, especially those focused on instructional support. We also hypothesize our tool would provide distinct, subject-specific insight into text-focused instructional quality. 3. What is the relationship between scores on text-focused instruction and teacher and student characteristics?
We hypothesize that text-focused instruction will vary with teacher characteristics, especially teacher experience. We hypothesize—or hope—that instructional quality will not systematically vary by student characteristics. 4. To what extent do scores on text-focused instruction predict teacher value-added estimates based on student achievement outcomes?
We hypothesize that text-focused instructional quality positively predicts teacher value-added estimates.
Background and Framework
Features of High-Quality Text-Focused Instruction
The IPRT, like the IPG coaching tools that preceded it, reflect Valencia and colleagues’ (2014) theory of instructional quality, which articulates that various, interconnected elements of comprehension instruction work together to support positive student outcomes. These include the texts teachers use, the tasks students are provided for engaging with those texts, the scaffolding teachers afford, the goals teachers have for students, and the clarity with which those goals are communicated to students (Cervetti et al., 2016; McKeown et al., 2009).
Teachers need to carefully select appropriately challenging texts, and there needs to be sufficient instructional time dedicated to engaging with those texts (Fang, 2016; Shanahan, 2013). Reading scholars emphasize the importance of combining both quantitative and qualitative dimensions to understand a text’s complexity, as each assesses distinct aspects of text quality (Mesmer et al., 2012; Valencia et al., 2014). Quantitative measures focus on text readability and are calculated in different ways. In determining vocabulary difficulty, Lexile levels assume that words with more frequency in the English language are easier for students, whereas Flesch-Kincaid defines difficult words as those with more syllables (Toyama et al., 2017). Standards highlight four qualitative text dimensions: levels of meaning, structure, language conventionality and clarity, and knowledge demands (CCSSO, 2010). Steinbeck’s The Grapes of Wrath has a low Lexile level but is primarily taught in high school because the historical context and themes are more appropriate for older readers (Glaus, 2014).
It is not just the sophistication or complexity of the text that generates disciplinary challenge for students. The tasks teachers ask students to complete working with the same text could be radically different. One teacher could ask students to recall facts made explicit in a text, while another could use the same text to push students to make synthetic arguments about an author’s use of language (Valencia et al., 2014). These two tasks with the same text would demand quite different reading skills, which are vital considerations in assessing and supporting the overall text-focused “instructional quality” (Goldman et al., 2016).
Scores of studies also underscore the vital role teachers play in facilitating meaningful student interactions with texts (Boardman et al., 2018; Dewitz & Graves, 2021; Duke et al., 2011). In particular, the questions teachers ask and how they respond to students’ questions and contributions can support active text engagement, help students revise textual misunderstandings, and make text-based inferences and arguments (Deshler et al., 2007; McKeown et al., 2009; Reznitskaya et al., 2009; Shanahan et al., 2010). In more productive interactions, teachers probe students’ contributions, pushing them to clarify and elaborate their text-based responses, using specific evidence from the text itself (Nystrand & Gamoran., 1991; Taylor et al., 2003). The questions teachers ask—coupled with their responses to student contributions—can support close reading (McKeown et al., 2009; Snow & O'Connor, 2016). Teachers who affirm the use of textual evidence have students who subsequently provide more text-based evidence (Gillies & Khan, 2009; Jadallah et al., 2011). It is thus important to analyze both teacher questions and student responses (Connor et al., 2020; Gámez & Lesaux, 2015).
Text-focusing questions can guide students to develop textual understandings that are more grounded in text evidence (Beck & McKeown, 1981; McElhone, 2012). When teachers support students in justifying their inferences with precise textual evidence, they help develop students’ argumentation skills (Hillocks, 2010; Reznitskaya et al., 2009). Some have argued that students can engage more deeply with texts when teachers begin lessons with a big picture “text-focusing question” (“What are Walt Whitman’s key ideas in this poem, and how does he express them?”)—followed by smaller-grained, aligned questions (“What are the narrator’s feelings in the first stanza?”) (Kaakinen & Hyona, 2005; McCrudden et al., 2010).
Research and standards also suggest the importance of students’ developing vocabulary knowledge during text-focused instruction. This can include instruction on identification of word parts (e.g., prefixes and suffixes and root words) and use of context clues to determine a word’s meaning (CCSSO, 2010; Lubliner & Smetana, 2005). Targeted vocabulary instruction can not only help students understand a particular text, but also build word knowledge and deduction skills for use with other texts (Beck et al., 1982; Lubliner & Smetana, 2005).
There are also teaching moves not exclusively related to literacy instruction that have been well-documented to support student engagement with texts. Illustrating how specific, more immediate academic tasks relate to broader, expansive learning goals can encourage student engagement (Grossman et al., 2014; Borko & Livingston, 1989). Research suggests the benefit of explicitly stating the purpose of the lesson, while also revisiting that purpose throughout a lesson, continually coming back to broader, disciplinary goals (McElhone, 2012; Nystrand & Gamoran, 1991). Finally, interactional scaffolding, or in-the-moment responses to students’ textual misunderstandings, has been shown to help develop students’ comprehension skills (Athanases & de Oliveira, 2014; Clark & Graves, 2005; Reynolds & Goodwin, 2016).
Building a Validity Argument for Classroom Observational Measures
In recent years, scholars have endorsed the use of subject-specific tools that provide teachers with feedback on specific facets of teaching and afford a common vocabulary for describing and supporting improvement on these facets (Hiebert & Stigler, 2017). Empirical evidence suggests teachers carefully attend to the language of a district’s observational rubrics, underscoring that teachers consider a tool’s dimensions not simply as a barometer, but also as a guidepost for quality (Phipps & Wiseman, 2021).While content-generic tools, like FFT (Danielson, 2007) or the CLASS (Pianta & Hamre, 2009), help characterize instruction across a variety of classrooms (Kane & Staiger, 2012), those tools’ macro-level ratings of broad constructs like the “Positive Climate” in the classroom often fall short of the nuanced, subject-specific information critical to helping teachers understand and improve disciplinary practice (Connor et al., 2009; Hill & Grossman, 2013).
Though the technical properties of a tool can seem removed from classroom interactions, such properties have direct implications on information afforded teachers. Does each rubric represent a distinct aspect of teaching, or do some practices “hang together” in empirical factors? If so, rather than focusing on each individual rubric, coaches or principals can provide teachers more parsimonious insight organized around broader factors. Teachers, like all learners, benefit from organizing structures to help them make sense of information gleaned from an observation.
Teachers would also benefit from evidence about whether raters consistently or reliably score instruction so that they know a score is indicative of the quality of their teaching, rather than sources of construct irrelevant variance such the idiosyncratic preferences of a rater (Hill et al., 2012; Zhai et al., 2021). Consistency across raters is needed for teachers to trust the “objectivity” of ratings. Teachers’ observational ratings are also often associated with students’ prior achievement (Whitehurst et al., 2014) or students’ demographics (Campbell & Ronfeldt, 2018). It’s unclear whether these observed differences are driven by “true” differences in instruction or rater bias, but regardless, such patterns can impede teachers’ trust in such ratings.
We also need evidence regarding the stability of instructional constructs from lesson to lesson and the extent to which this measure captures information distinct from other tools (Kane & Staiger, 2012). These kinds of evidence help inform both sampling decisions (e.g., how many lessons do I need to observe to get a reasonable sense of a teacher’s instruction?), as well as the inferences that a tool can support (e.g., does this tool provide different information about a teacher’s practice than other tools we might use?). These day-to-day or week-to-week fluctuations may depend in part on the goals of a lesson or sequence of the unit (Grossman, 1990; Grossman et al., 2014), but such variation can also complicate inferences about individual teachers needed for summative evaluations and personnel decisions (Cohen & Goldhaber, 2016). At the same time, descriptive information documenting temporal variation in instruction may provide invaluable formative information to teachers and school leaders. Similarly, if a content-specific tool is highly correlated with generic observation protocols, it would have little relative advantage.
Finally, a central component of recent classroom observation research has been the predictive relationship between teaching practices and student outcomes, which have tended to be relatively low to modest (Allen et al., 2011; Kane & Staiger, 2012). There are many factors that could mitigate a strong correlation. There may well be omitted, unobservable, or non-school related practices associated with student outcomes but are uncaptured by a particular tool (Condron, 2009). Teaching can also vary greatly across lessons, depending on the instructional practice being measured, which can complicate inferences about relationships with measures from a single time point, such as value-added models, which provide insight into the degree to which the achievement of a teacher’s students is better (a positive value-added coefficient) or worse (a negative value-added coefficient) than one would predict based on students’ prior achievement and demographic characteristics (Garrett & Steinberg, 2015; Grossman et al., 2014; Hill et al., 2012; Pianta & Hamre, 2009; Polikoff, 2012, 2015). Nevertheless, analyzing whether observational tools are correlated with VAMs is useful, as it can provide suggestive evidence of whether the instructional focus of a tool might also contribute to desired student outcomes. Our analysis is organized around these questions about observational measures based on these types of evidence that teachers, schools, and researchers alike might need.
Data and Methods
We leverage data from a larger study focused on the implementation and effects of a content-specific professional learning curriculum including administrative data on schools, teachers, and students, along with teacher-submitted videos of text-focused reading instruction.
Participants
In partnership with the district, the research team invited third-fifth grade ELA teachers clustered in a set of volunteer schools to participate. We selected teachers to maximize variation in demographics, experience, and prior performance (based on the district’s evaluation system). Sixty-six teachers in 24 elementary schools submitted a total of 580 videos of 30-min reading lessons over the course of two years. 1 We excluded the 64 lessons that did not involve a text. The average teacher submitted 8 scoreable lessons, with one teacher submitting 17 lessons.
Comparison of the Study Sample to the District’s Teacher and School Population.
Note. All data from 2018, the second year of our study with the exception of 10 study teachers whose data comes from 2017 as they did not continue with the study in 2018. One teacher in the study sample could not be linked to the administrative data and thus has missing values for all teacher characteristics. p-values in the last column reflect a simple difference-in-means t-test.
aDistrict sample N is 3870 for observation scores, is 269 for Value-Added Score, and is 3876 for Experience and the Study Sample N is 39 for Value-Added Score.
Measures
Below, we describe the measures we used in this study. First, we discuss the development and properties of the IPRT-ELA, the tool that we developed to better and more systematically describe text-focused instruction. We designed our tool with an explicit purpose: to provide teachers and school systems with discipline-specific, formative information about text-focused reading comprehension instruction. The tool was not designed for teacher-level, summative evaluation. To assess the convergent and discriminant validity of the IPRT-ELA, we leverage: 1) teacher-level observational scores from the district’s evaluation measure, rubrics derived from Danielson’s Framework for Teaching and 2) lesson-level scores on the Classroom Assessment Scoring System (CLASS), a widely used, content-generic observational tool, shown to have significant associations with students’ reading achievement in some contexts and grade levels (Kane & Staiger, 2012; Pianta et al., 2008).The predictive validity analysis of the IPRT-ELA uses teacher’s value-added scores based on district-provided student achievement data.
Instructional Practice Research Tool for English Language Arts
IPRT-ELA Items & Descriptions.
Over several rounds of refinement, we finalized items, modifying the original coaching protocol to emphasize lower-inference and more readily observable behaviors, while not sacrificing key aspects of reading instruction. For instance, we dropped an item from the coaching protocol that asked whether “teacher[s were] pos[ing] challenging questions…[that kept] all students persevering with them,” because we determined that perseverance was a construct that could not be consistently observed across raters (a crosswalk between the IPG and the IPRT-ELA, along with details on scoring are available upon request).
The resulting IPRT-ELA contains eleven items assessing teachers’ text-focused instructional practice. Raters score 30 minutes of consecutive instruction and also collect any supporting materials that teachers provided for that lesson (e.g., student worksheets and lesson plans). Nine of the eleven indicators are scored using 4-point ordinal scales, with 1 representing no opportunities to engage and a 4 indicating consistent engagement. The remaining two indicators are binary scales. For the Text Quality indicator where a 1 represents whether the text is of publishable quality (i.e., copyrighted) rather than adapted for adapted for classroom use. For example, Newsela provides several Lexile level options for a single newspaper article so teachers can select the appropriate level for their students, effectively modifying the news article from its original form. Though these adapted texts may be more accessible for students to read, the adapted version may also be stripped of the original’s unique structure and language, decreasing the text’s overall complexity and quality (Hoffman, 2017). The Quantitative Complexity indicator captures text complexity by mapping readability onto a grade band using at least two Lexile-level recommendations (i.e., ATOS, Degree of Reading Power ®, Flesch-Kincaid, The Lexile Framework ®, Reading Maturity, and Text Evaluator). Texts were coded as being at or above the complexity expected for the class’s grade band or below that expected.
CLASS Observational Tool
The Classroom Assessment Scoring System (CLASS) is undergirded by the idea that positive teacher-student interactions increase student engagement and learning (La Paro et al., 2004). CLASS assesses three domains—Emotional Support, Classroom Organization, and Instructional Support—in which raters score several dimensions on a 1 (low) to 7 (high) scale. The CLASS tool is designed to score 15-min lesson segments, whereas the IPRT-ELA tool requires a 30-min lesson. Therefore, we average scores across the two 15-min segments, so the unit of observation was consistent across tools. Dimensions were then averaged together at the domain level, as is typical with CLASS scoring (Pianta & Hamre, 2009). An overall CLASS score was generated by averaging the three domain scores.
Scoring Lessons
Members of our research team trained a team of five raters to score instruction using the IPRT-ELA and 10 raters to score both reading and mathematics lessons using CLASS (mathematics lessons were used in another component of the larger study from which this study’s data are drawn). For both tools, trainings lasted two days, followed by a certification process. Master raters from the research team scored a set of five videos of reading instruction. Raters were certified to begin scoring videos using the IPRT-ELA tool once they reached at least 80% agreement with the master scores for each IPRT-ELA item or CLASS dimension across the five videos. Throughout the two-year study, raters attended weekly calibration meetings.
IPRT-ELA and CLASS Reliability Statistics.
Note. CLASS lesson segments include mathematics lessons which were coded by the same group of coders at the same time as the ELA lessons.
District Observational Measures
The district provided us with scores from a content-generic observational measure to evaluate teachers’ instruction, which is based heavily on Danielson’s Framework for Teaching. The measure consists of five rubrics: (1) cultivate a responsive learning community, (2) challenge students with rigorous content, (3) lead a well-planned, purposeful learning experience, (4) maximize student ownership of learning, and (5) respond to evidence of student learning. The five scores, each on a four-point scale, are averaged to create an overall score. Teachers’ yearly observation score is an average of five, 30-min observations. School administrators conduct three, and two are completed by an “expert” who conducts observations at many schools. The final observation score is included as part the annual evaluation process that determines a teacher’s monetary incentive (for descriptive statistics, see Table 1).
Administrators engage in a three-day training about the general principles of the rubrics and receive monthly professional development to practice scoring instructional videos. Expert raters participate in a six-week intensive training, with bi-monthly norming sessions (Curtis, 2011). Though the district has employed multiple methods to increase reliability, there is little work about the reliability of these observational scores (Gitomer et al., 2014).
Individual Teacher Value-Added Student Achievement Data
Teachers’ VAM scores are intended to demonstrate their instructional impact on students’ achievement. The district provided us with the score they estimate as part of their evaluation system and produced by predicting students’ end-of-year test scores for each teacher based on students’ prior test scores and demographic variables. The district then compares the differences between predicted scores and actual test scores to generate a teacher-level VAM score. The tests used for calculating VAM in this study are the PARCC examinations in ELA.
Data Analysis
Multilevel Factor Analysis
The design of the IPRT-ELA tool is based on the theory that the 11 indicators collectively describe the degree to which a teacher’s text-focused practices align with the CCSS-ELA. We conducted multilevel exploratory and confirmatory factor analyses (EFA and CFA) to test this theory. The multilevel factor analyses account for nesting of lesson observations within teachers and, in so doing, allows us to examine the underlying structure of the IPRT-ELA both within and between teachers. Our analyses treat all indicators as categorical to preserve the ordinal nature of the response options. The search for the best fitting factor structure of text-focused instructional practice at each level informed our path through the EFA to the CFA. We estimate all multilevel factor analyses in Mplus 8 (Muthén & Muthén, 2017).
Exploratory Factor Analysis (EFA) Results.
*p < .05.
N = 683 scores within 66 teachers.
Confirmatory Factor Analysis (CFA) Results.
***p < .001, **p < .01, *p < .05.
N = 683 scores within 66 teachers.
aVariances set to 0 for these items to adjust for initial negative residual variances.
Summary of IPRT-ELA Instructional Practices.
aN = 382 scores, 305 lessons, 63 teachers, and 23 schools.
bN = 510 scores, 404 lessons, 66 teachers, and 23 schools.
Note. Some of the 516 lessons were coded by multiple coders, yielding 683 lesson scores. A cross-classified multi-level model was estimated for the variance decomposition in order to account for the two teachers who switched schools between the first and second year of the study.
Our Text-Focused Instructional Practices (TFIP) measure includes the four indicators about text-focused teacher questions and student responses (Text-specific Questions, Text-Dependent Questions, Question Sequence, and Text-based Responses), as well as the Text Focus indicator that captures how central the text is to the lesson. We theorize that Purpose, Teacher Probing, and Scaffolding indicators do not load on this factor because they are general practices and are not necessarily centered around supporting students to engage with the text.
Multilevel Modeling
With factor structure confirmed, we estimate a series of cross-classified multilevel linear models to explore the contextual variation of text-focused practices, and its association with other measures of instructional practice and value-added scores. Drawing on the factor structure analysis, we average the five text-focused IPRT-ELA indicators across raters of the same lesson and calculate the measure as a simple average (Cronbach alpha = .77).
2
The data have both a purely hierarchical nesting component with multiple lessons observed per teacher and a cross-classified structure with most teachers nested within a school, but with two teachers who taught in a different school in each of the study’s two years. The cross-classified multilevel linear model specification allows us to analyze associations between theorized predictors and the IPRT-ELA tool’s measure of text-focused practices. Using the classification notation of Browne et al. (2001), we fit the following model to the data
Association with Student Outcomes
In the version of this cross-classified multilevel linear model that explores the association between text-focused practices and teacher value-added scores, the
Findings
RQ#1: Instruction in Upper Elementary Classrooms
The average text-focused lesson here is somewhat moderately aligned with practices known to support students in making sense of complex texts (see Table 6). Across the IPRT-ELA items using a four-point scale, means fall between a 2 (lower frequency of text-focusing practices) and 3 (moderate frequency of text-focusing practices), with the exception of the Vocabulary Instruction (mean just below 2). The raters were unable to assess text quality in 21.7% of the lessons because they were unable to identify the text’s source. They were also unable to assess the text’s complexity in 40.9% of lessons because the text was not listed in any reference guides. Of the lessons in which text quality and complexity could be assessed, almost half (47.3%) used a quality text, and 44.7% used a complex text. Across all items, most of the score variation is within a teacher (80%–93%). That is, there is little consistency in text-focused instructional quality across an individual teacher’s reading lessons. A meaningful share is across teachers within a school (5%–19%). There is very little variation between schools (see Table 6).
Summary of CLASS Instructional Practices.
Note. Some of the 516 lessons were coded by multiple coders, yielding 683 lesson scores. A cross-classified multi-level model was estimated for the variance decomposition in order to account for the two teachers who switched schools between the first and second year of the study.
Teachers’ text-focused practices are associated with characteristics of the text used in the lesson (Table 8). Lessons that employ a quality text have a Text-Focused Instructional Practices score a third of a standard deviation higher than in lessons using a low-quality text (
RQ#2: Convergent and Discriminant Validity
Contextual Variation in Text-Focused Instructional Practices.
aWe exclude one teacher with five lessons who we were unable to link to DCPS’s teacher-level administrative data.
***p < .001, **p < .01, *p < .05.
Convergent and Divergent Validity of Text-Focused Instructional Practices.
***p < .001, **p < .01, *p < .05, + p < .1.
RQ#3: Contextual Variation
The use of text-focused practices is also associated with teacher experience and a school’s concentration of special education students (see Table 8). Teachers with more experience have higher Text-Focused Instructional Practices scores than teachers with fewer years of experience, although the association is small (.02 additional points on the rubric, or 3% of a standard deviation, with each additional year of teaching experience;
RQ#4: Predictive Validity
Predictive Validity of Text-Focused Instructional Practices.
Limitations
We recognize that the IPRT-ELA has its limitations. It is designed to only assess 30 consecutive minutes of text-based instruction and does not analyze how teachers might build on their instruction over several lessons, a pitfall of many observational measures (Cohen et al., 2020). Teachers selected the time and day for the research team to videotape, so it is possible that these lessons are not representative of a teacher’s overall instructional practice. In addition, the tool focuses only on whole group interactions around text. There may well be important small-group and/or individual-level instructional interactions that support students’ engagement with texts. These would not be captured in IPRT scores, but there are also other tools with a strong evidence base that focus on the reading classroom environment, from the perspective of an individual student, including the Individualizing Student Instruction (ISI; Connor et al., 2009).
The IPRT-ELA tool was designed to focus only on reading instruction. Other aspects of language arts teaching, including writing, speaking, and listening would not be well-aligned with the IPRT-ELA. It is worth noting that this tool is not designed to assess instruction on foundational reading skills (i.e., instruction on phonics or phonemic awareness), and thus would be less helpful for either formative or summative assessment of primary grades reading instruction.
We also acknowledge that our measures of text complexity and text quality are blunt and limited by what we could discern in video. In future research, we hope to collect texts used and develop additional methods for analyzing their quality and complexity. A central part of Valencia and colleagues’ (2014) framework is the nature of the instructional task, which we found impossible to meaningfully quantify. However, we asked raters to qualitatively note the nature of the task, and in future research, we plan to analyze these qualitative data. In addition, scholars have emphasized the importance of attending to the sociocultural contexts in which text-focused instruction is occurring (Goldman & Lee, 2014). While we, too, agree that context is crucial, we were also unable to quantify this in our observational measure. Future mixed-methods research should explore the degree to which and ways in which our findings may look different in distinct sociocultural contexts. No single observational tool is comprehensive, though we developed the IPRT-ELA to capture a wider range of reading focused practices than other existing tools. Our hope is that this tool may provide formative feedback for teachers about their text-focused teaching practices.
Discussion and Implications for Practice
Part of the challenge involved in improving students’ ability to engage with complex texts has been a lack of reliable observational tools that can be used to identify and support improvement on specific features of instruction (Grossman et al., 2014). In contrast to mathematics, where there are many observational tools (Hill et al., 2008; Schoenfeld et al., 2014; Walkington & Marder, 2018), there are very few tools to reliably assess reading instruction in elementary classrooms. Of those, none focus squarely on text-focused instruction that promotes students’ comprehension skills. Teachers would likely benefit from tools that capture the quality and quantity of practices known to support students (McDuffie et al., 2018; Porter et al., 2015). In particular, the Common Core and other ambitious standards foreground the need for students to engage deeply with authentic, challenging texts, but few observation tools feature text-focused teaching. Some tools like the PLATO include a broad scale for “Text-Based Instruction” but stop short of parsing different facets of such instruction (Grossman et al., 2014).
In this paper, we provide evidence about such a text-based observation tool, the IPRT-ELA. Over the course of several years, we used this tool to reliably assess hundreds of reading lessons in a large school district’s upper-elementary classrooms. We find encouraging evidence that many reading lessons in the district were indeed text-focused and frequently incorporated text-dependent questions. We also found considerable variability in other text-centric practices, such as whether teachers use scaffolded sequences of questions to support students’ engagement with texts. Districts will have to support teachers in continuing to develop their text-based instruction. The practices highlighted here and their corresponding rubrics, coupled with examples of high-level enactment, could be supportive tools for formative assessment and professional development, though more research is needed to explore this potential.
Variation Across Lessons and Formative Insight About Text-Focused Teaching
Like other subject specific tools, the IPRT-ELA scores fluctuated considerably more at the lesson level than at the school level. This finding resonates with Hill and Grossman’s (2013) argument that subject-specific tools measure more fine-grained practices that likely vary across lessons in ways that generic and broad measures of classroom climate and behavior management do not. Not every ELA lesson focuses on the same text-focused reading skills, and nor should it. However, even though this temporal variation may reflect natural—and true—variation in practices in reading, it does make a tool like the IPRT more challenging to use for consequential, summative evaluations of teachers. Our data suggest it would take many more observations than is typical to make appropriate inferences about a teacher’s overall practice. That said, we see the major benefit of the IPRT-ELA as the provision of more targeted, formative data for teachers, schools, and educational systems. In other words, this tool would be useful to offer teachers, instructional coaches, and districts information on where to improve teachers’ text-focused instruction. In particular, our factor analyses helped us identify a group of five text-focusing practices. This could provide teachers and schools a parsimonious set of practices for professional development and coaching, though more research is needed on the usability of the tool by school-based professionals in live observations, along with additional work on how teachers make sense of and respond to the information provided by the tool. We are in the midst of collecting those data.
Operationalizing Text Quality
We also found that Quantitative Complexity and Text Quality, measures of the classroom materials used, were associated with scores on text-focused practices. Indeed, it may well be difficult to engage in text-focused scaffolding absent rigorous texts. While prior literature theoretically highlights the importance of using complex texts in the classroom (Fang, 2016; Shanahan, 2013), our data provide early empirical evidence that quality reading instruction may indeed be contingent upon teachers using grade-level, appropriately challenging texts. In this way, there may be an important relationship between the quality of materials and supports for engaging with those materials. Though understanding the nuances of this relationship is beyond the scope of this study, we see this as an important direction for research (Cohen et al., 2020).
These findings also point to the need for more robust and validated measures of text complexity. We acknowledge that quantitative reading measures are only one way of examining the difficulty of a text. The extent to which a text is considered complex also involves readers’ prior background knowledge, whether the text has multiple levels of meanings, and the degree to which the author’s language is conventional (Fisher & Frey, 2014; Mesmer et al., 2012). We also need better insight into the ways in which text selection occurs based on the sociocultural context in which reading instruction takes place (Goldman & Lee, 2014). We need more robust evidence about such decision-making processes, and the implications of them, to better support literacy coaches and teachers in their selection of complex texts that will engage their students and cultivate their comprehension skills.
Is Text-Focused Instructional Quality Just “Good Teaching”?
We also set out to understand the degree to which the IPRT-ELA measured distinct facets of instruction from those captured by the CLASS and the district’s high-stakes observation measure. While we find statistically significant correlations between our measures and the district’s observational measure and the CLASS domain of Instructional Support, those correlations were low to moderate (i.e., r < .25). This suggests that the IPRT-ELA is assessing different aspects of reading instruction than those prioritized by content-generic tools. A tool like the IPRT-ELA could provide teachers with helpful data about their disciplinary teaching practices (Hill & Grossman, 2013), as well as illuminate for districts the relative strengths and weaknesses of reading instruction across classrooms. Such data could then inform the focus of district-wide professional development efforts, along with tailored coaching supports for individual teachers. In other work, we have argued for the value of using discipline-specific tools in concert with more general tools like the CLASS (Berlin & Cohen, 2020). Some teachers may well need support with generic practices—including the cultivation of warm and productive classroom environments—along with support in the text-focused practices captured here. Understanding the interplay between the two, as well as how schools use different tools to support teacher development, is also beyond the scope of this study but is a vital area for ongoing research.
We also found some unexpected correlations between IPRT scores and other CLASS domains. There was a significant, negative relationship between text-focused practices and the CLASS domain of Emotional Support, which focuses on teacher sensitivity, and the classroom climate. This is counter to what we found in mathematics lessons in the same district (Berlin & Cohen, 2020). Though our data preclude inferences about mechanisms underlying this relationship, the Emotional Support dimensions of CLASS do privilege student autonomy and peer-to-peer interactions. The teacher-focused instructional scaffolds highlighted in the IPRT-ELA may well occur less often in student-led lessons. Future research should investigate these relationships in greater depth, complemented by qualitative insight into mechanisms.
Variation by Context and Predicting Student Outcomes
Before using a given observational tool, it’s also important to understand the degree to which it captures the quality of instruction, distinct from the characteristics of students. Scholars have found that some observational tools may be biased, particularly in that teachers of predominantly students of color may systematically receive lower scores (Campbell & Ronfeldt, 2018; Chaplin et al., 2014). Fortunately, most of the student characteristics examined—including percentage of Black students, ELL students, and students receiving subsidized lunch—were not significantly associated with IPRT scores. However, we do see significant, negative associations between the percentage of students receiving special education services and scores IPRT-ELA text-focused scales. It is unclear whether this finding suggests an actual difference in teachers’ instruction in classrooms with these students or rater bias, but it indicates more research is needed to ensure students with disabilities in the district are receiving the same levels of text-focused instructional support as other students. Future studies should be conducted in a way that helps the field disentangle whether and how rater bias may influence scores. This, along with other threats to measurement reliability, should be examined with data collected on text-focused teaching in a range of contexts, to assess the generalizability of the measure.
Finally, we explored the relationship between text-focused practices and teacher value-added measures. Similar to previous studies analyzing the relationship between observational tools and student outcomes, we find small, not statistically significant correlations between the two (Bell et al., 2012; Grossman et al., 2013; Hill et al., 2011; Kane & Staiger, 2012; Strong et al., 2011). Many teachers in our sample did not have value-added data, so we could only test a smaller subset of participants, and hence our estimates are quite noisy. Examining these relationships with larger samples is an important next step in our research.
Conclusion
Despite the plethora of tools for assessing high-quality, standards-aligned teaching in mathematics, we have comparably few tools that measure reading instruction. Here we provide an array of evidence about such a tool, the IPRT-ELA, which we developed in partnership with Student Achievement Partners (SAP) to help districts, schools, and researchers both understand and support teachers in helping students engage with complex texts. Scores of schools and districts are already using SAP’s Instructional Practice Guides in informal walkthroughs and for coaching, though there is limited research about such efforts. A central goal of this work was to build an empirical base about these particular practices. Our efforts also provide a proof of concept that research teams can work productively with educational non-profits who provide much of the instructional support to districts around the country. Rather than treating practitioner-facing tools as distinct from research tools, we must find ways to bridge the two so that we develop rigorous empirical evidence about the tools being used in American schools.
Decades of research suggest that measuring core teaching practices and providing teachers with such information can help improve both instruction and student outcomes (Dee & Wyckoff, 2015; Dixit, 2002; Taylor & Tyler, 2012). However, in this district, as well as in most districts across the country, teachers receive only content-generic information about their teaching (Cohen & Goldhaber, 2016). These data suggest that such information would obscure content-focused aspects of text-based instruction. Though more research is needed in larger and more representative samples of teachers across districts, we see these findings as an exciting first step towards helping teachers support students in engaging with rigorous, grade-level texts.
Footnotes
Author’s Note
We could not have completed this work without our colleagues at Student Achievement Partners. In particular, Jessica Eadie, Lisa Goldschmidt, and Amanda Vitello worked alongside us throughout our work. Rachel Etienne spent countless hours discussing videos with us, as we worked together to develop the measures described here. We are also grateful to our district partners for supplying the data employed in this research and for answering our many questions.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Finally, the authors wish to thank the Charles and Lynn Schusterman Family Philanthropies and the Overdeck Family Foundation for their financial support of this work. The research reported here was also supported by the Institute of Education Sciences, U.S. Department of Education, through Grant #R305B140026 to the Rectors and Visitors of the University of Virginia. The opinions expressed are those of the authors and do not represent views of the Institute or the U.S. Department of Education. Errors are attributable to the authors.
