Abstract
Observational methods are increasingly being used in classrooms to evaluate the quality of teaching. Operational procedures for observing teachers are somewhat arbitrary in existing measures and vary across different instruments. To study the effect of different observation procedures on score reliability and validity, we conducted an experimental study that manipulated the length of observation and order of presentation of 40-minute videotaped lessons from secondary grade classrooms. Results indicate that two 20-minute observation segments presented in random order produce the most desirable effect on score reliability and validity. This suggests that 20-minute occasions may be sufficient time for a rater to observe true characteristics of teaching quality assessed by the measure used in the study, and randomizing the order in which segments were rated may reduce construct irrelevant variance arising from carry over effects and rater drift.
Teaching observation measures are growing in popularity as a method to evaluate the quality of teaching. The increased use of these measures is largely due to recent shifts in education policy. For example, Race to the Top encourages the use of teaching observation measures in conjunction with other measures of teacher performance such as value-added (U.S. Department of Education, 2009). State and district policies are also shifting toward formal observation systems, with 24 states and the District of Columbia requiring observations as components of yearly teacher evaluations (Heitin, 2011). Teaching observations have a long history of use in education research for purposes of identifying characteristics of classroom settings that are associated with student learning (Bell et al., 2012; Bill and Melinda Gates Foundation, 2012; Mashburn et al., 2008). Observational measures are also part of new models of teacher professional development that use video observations to provide feedback and support to teachers in ways that lead to improved student learning (Allen, Pianta, Gregory, Mikami, & Lun, 2011; Fritz & Chen, 2013; Mashburn, Downer, Hamre, Justice, & Pianta, 2010). Given the emphasis on teaching observations in policy, research, and professional development, it is no surprise that the psychometric characteristics of these measures are receiving increased scrutiny through large-scale studies such as the Measures of Effective Teaching project (see Bill and Melinda Gates Foundation, 2012).
Methods for conducting teaching observations vary across different instruments and across studies using the same instrument, and the methods that are adopted may affect the reliability and validity of scores. For example, methods of assessing the quality of teaching during a day may involve short observations with frequent ratings or long observations with infrequent ratings. Modes for collecting data may involve either live observations in the classroom or observations of videotaped lessons, and in the case of videotapes, observation segments may be presented either in sequential order as they occurred in real time or in random order. To examine the impact of length of observation and order of presentation on the reliability and validity of scores, we conducted an experimental study in which we manipulated the length and order of videotaped lessons of secondary grade classrooms to compare the reliability and validity of scores on the Classroom Assessment Scoring System–Secondary (CLASS-S; Pianta, Hamre, Hayes, Mintz, & LaParo, 2008). The goal was to identify observation procedures that maximize score reliability and the validity of score inferences.
Observational Measures in Education
Observational measures are available that assess various aspects of the quality of teaching; some of which assess the quality of teaching in a particular subject matter whereas others assess teaching more generally. For example, two content-specific measures are the Mathematical Quality of Instruction (MQI; Hill et al., 2008) instrument that assesses five domains of teaching mathematics to students in kindergarten through 8th grade classrooms, and the Protocol for Language Arts Teaching Observation (PLATO; Grossman et al., 2010) tool that measures four components of the quality of English language arts instruction. Other observation measures assess general dimensions of teaching quality regardless of the subject area. For example, the Framework for Teaching (FFT; Danielson, 2011) measures the general quality of instruction in four domains; two of which are measureable through teaching observations, and two require additional information about a teacher’s planning and professional behavior. The Classroom Assessment Scoring System (CLASS; Pianta et al., 2008) is a family of observational measures that tap into the quality of teacher–student interactions related to three domains—Emotional Support, Classroom Organization, and Instructional Support (Hamre et al., 2013). Interestingly, despite their differences, scores from subject specific and general observational measures tend to be highly correlated (Bill and Melinda Gates Foundation, 2012).
One characteristic shared by all these measures is that one or more trained raters observe a classroom or a videotaped lesson for a given period of time and provide scores reflecting the quality of teaching during that segment of time. For example, in the Measures of Effective Teaching project, raters judged videotaped classrooms using the CLASS-S and FFT (see Bill and Melinda Gates Foundation, 2012), and operational procedures for each of these measures during the project involved raters viewing and scoring the first 15 minutes of a lesson and then viewing and scoring a subsequent 15 minutes of the same lesson. However, in other studies using these measures, the protocols for dividing lessons into observation segments has varied. For example, prior use of the CLASS-S has involved dividing 40-minute lessons into two 20-minute segments (Mikami, Gregory, Allen, Pianta, & Lun, 2011), and standard protocols for the FFT involve a single rating of an entire 40- to 50-minute. Thus, the operational procedures for dividing a lesson into one or more segments vary across instruments and across studies using the same instrument. As we discuss in the subsequent section, these operational procedures regarding observation length can affect the score reliability and validity.
Framework for Reliability and Validity
Kane (1982, 2011) proposed a sampling model for validity that focuses on the accuracy of using an observed score to make an inference about a universe score. His framework derives from generalizability theory (Cronbach, Gleser, Nanda, & Rajaratnam, 1972) and makes an explicit connection between reliability and validity. He explains that large amounts of random variation in observed scores (i.e., measurement error) can be reduced in three possible ways, but each one entails a tradeoff. Measurement error can be reduced through (a) more complete sampling of the universe (e.g., longer tests, more raters), (b) restricting the target universe, or (c) standardizing the measurement procedure. Of these three approaches, the last two involve a tradeoff between reliability and validity. Restricting the target universe is an extreme solution that narrows to the universe to a specific set of conditions. It improves reliability by eliminating it altogether through a narrow definition of the universe. However, this narrow universe is far from the target universe, which ultimately leads to biased scores and inaccurate inferences. Standardization is a less extreme solution in which only a few facets of the universe are fixed to specific values. It improves reliability by eliminating some, but not all, sources of measurement error. Standardization still narrows the universe and biases inferences, but to a lesser extent than a complete restriction of the universe. Thus, reliability and validity are directly connected in Kane’s framework. Procedures that improve reliability often come at the expense of a loss of validity.
Kane (1982) viewed his sampling model and generalizability theory to be one method for providing evidence of construct validity. It is a complement to other methods and concepts in validity theory such as content and criterion-related validity, not a replacement of them. Indeed, Marcoulides (1989b) demonstrated the way convergent and discriminant validity can be explored through Kane’s framework. A benefit of Kane’s approach is that it makes an explicit connection between reliability and validity, and it provides a way of using generalizability theory to explore both. In the context of teaching observations, most researchers focus on reducing measurement error by sampling more of the universe. They fail to consider the way standardization improves reliability and biases inferences about the measured trait. Two features of teaching observations that are often standardized include the length of observation and the presentation order.
Observation Length and the Reliability and Validity of Scores
The length of time and frequency in which a rater observes a lesson before assigning scores can vary, and arguments can be made both for and against shorter, more frequent ratings and longer, less frequent ratings. In terms of validity, shorter ratings may be less subject to primacy and recency effects (Ebbinghaus, 1913) than longer ratings. For example, a single rating culminating at the end of a 40-minute observation will likely give undue weight to events that transpired in the first (primacy) and last (recency) 10 minutes of an observational period, and events that occurred during the middle 20 minutes are more likely to be forgotten and less likely to be incorporated into ratings. Primacy and recency effects can cause scores from a bad and subsequently good performance to be higher than scores from two good performances (Leventhal, Turcotte, Abrami, & Perry, 1983). Thus, shortening the observation from one 40-minute period to two 20-minute periods may improve validity by reducing construct irrelevant sources of variance such as primacy and recency effects.
A point of diminishing returns may be reached with respect to validity, however, if the observation time period is too short. This may be particularly salient when measuring complex, dynamic interactions between a teacher and students that are often the constructs of interest for teaching observations. In this case, an observational period may be too short for the desired classroom interaction to occur, and the resulting ratings will not permit a valid inference about the measured construct. As a result, excessively short observational periods can result in construct underrepresentation and compromise the measure’s validity.
Selecting an appropriate length of observation not only affects validity but also reliability. Dividing an event into more occasions may improve reliability by increasing the number observations that can occur during a set amount of time. If one focuses strictly on reliability and chooses excessively short, but very frequent observational periods, high reliability may be achieved at the expense of valid inferences. Conversely, if one focuses on validity, the optimal level of reliability may not be attained. In sum, there may be tradeoffs between reliability and validity whereby choices that improve validity may reduce reliability, and vice versa, and a challenge is developing observational procedures that maximize both reliability and validity.
Presentation Order and the Reliability and Validity of Scores
Teaching observation measures may be used during live lessons where the raters are present in the classroom while the lesson is taking place. Or, they may occur long after the lesson has taken place through videotaped footage of the lesson. Each mode has its benefits and limitations. Live observations allow raters to hear conversations and notice interactions that would otherwise be inaudible or hidden from view in a video recording, which benefit the validity of scores; however, there are limits to the number of raters who can be physically present in the classroom, which may compromise the reliability. In contrast, with videotaped lessons, there are no limits on the number of raters who can observe and rate each lesson. Error attributable to rater effects can be mitigated by increasing the number of raters; however, a video recording of the classroom may omit or not fully convey the quality of teaching within the classroom.
Another potential advantage of observing classrooms using videotaped lessons that is relevant to the current study is that the order in which segments are presented to raters can be manipulated. Whereas live observations require that lessons be viewed and rated in sequential order such that ratings for the first part of a lesson and followed by ratings for the second part, segments from video observations may be viewed in a random order whereby parts of lessons from any day and any teacher can be randomly presented to a rater. Manipulation of the order in which raters view lesson segments may reduce other sources of construct irrelevant variance including carry over effects (Ho & Kane, 2013) and rater drift (Casabianca & Lockwood, 2013). Carryover effects occur when the scores of one segment are not independent from the scores of another segment. In the case of sequential coding, segments from a specific lesson by a specific teacher are rated in back-to-back fashion; thus, the scores from a subsequent lesson are likely to be affected by the events from or impressions left by the occurrences during the prior segments. When the ordering of segments are presented randomly, such that raters view segments randomly drawn from the lessons of all teachers, the carryover effect from one segment to the next within a lesson for a given teacher can be mitigated.
Rater drift occurs when rater performance lacks invariance over time (Congdon & McQueen, 2000) such as raters changing their use/interpretation of a scoring rubric over the course of a rating period (Casabianca & Lockwood, 2013). It is a source of construct irrelevant variance because teacher scores are systematically affected by changes in rater behavior. Casabianca and Lockwood describe a statistical model for controlling rater effects, but random presentation of segments may also reduce the influence of rater drift as well as the effect of other extraneous variables.
Study Purposes
A major question for anyone wishing to implement observations for purposes of policy or research is, “How can I best allocate resources so as to minimize costs while producing reliable scores that permit valid inferences about teaching quality?” Manipulating the length of observations and presentation order of segments may affect the score reliability and the degree of validity evidence without affecting the financial costs of observing and rating the quality of instructional activities. Our study aims to experimentally test the effect of observation length and presentation order on the score reliability and the degree of validity evidence supporting teaching observations. Specifically, eight trained raters were randomly assigned to rate 40-minute videotaped lessons either in one 40-minute occasion, two sequential 20-minute occasions, four sequential 10-minute occasions, or two nonsequential 20-minute occasions. The purpose of this study was to compare the reliability and predictive validity of a teaching observation measure and explore other potential threats to validity using experimental conditions that represent different ways to fix observation length and presentation order.
Method
Videotaped Lessons
For the purpose of this study, we obtained a subset of data collected from teachers and students from Grades 6 through 11 as part of the efficacy study for the My Teaching Partner professional development program for secondary school teachers (Allen et al., 2011). The study involved eight schools from the southeastern United States with random assignment of teachers within schools to either a treatment or control condition. In total, the efficacy study involved 47 teachers in the treatment condition and 43 teachers in the control condition. During the course of the efficacy study, teachers videotaped 40-minute classroom lessons on multiple days throughout the academic year and submitted videotapes to researchers during predetermined windows of time. For this study, we retained 47 of these teachers who met the following two criteria: submitted at least one videotaped lesson during each of the following three time periods: (a) September to November, (b) December to February, and (c) March to May; and had at least one lesson during each time period that was 40 minutes in length or more and without audio or video problems. In cases when a teacher had two or more 40-minute lessons available during a time period, we randomly selected one of them. The resulting sample of videotaped lessons included a total of 141 (47 teachers, 3 lessons each) videotaped lessons.
In addition to videotaped lessons from teachers, demographic information and scores on state achievement tests were available for a total of 1,366 students enrolled in study teachers’ classes. Specifically, demographic characteristics of children included gender, minority status, and grade level. Achievement test scores in reading and math were collected during the prior school year and at the end of the school year during which the video lessons were collected.
Measure of Teaching Quality
The CLASS includes three observational measures that span pre-kindergarten through 12th grade. We used CLASS–Secondary (CLASS-S; Pianta et al., 2008) in this study to assess the quality of teacher–student interactions in middle and high school grades. The measure comprises 11 dimensions (i.e., items) that tap into three domains: (a) emotional support (EMSUP), (b) classroom organization (CLORG), and (c) instructional support (INSUP). Each dimension contributes to scores on one domain only and is rated on a 7-point scale. Anchor point descriptions for each dimension guide raters in selecting an appropriate score level. Factor analysis studies support the three domain structure of the measure (Bell et al., 2012; Malmberg, Hagger, Burn, Mutton, & Colls, 2010). However, the domains tend to be highly related, and models that take into account the nested structure of rating data suggest that a three factor or single factor model at the teacher level are plausible (McCaffrey, Yuan, Savitsky, Lockwood, & Edelsen, 2013; Savitsky & McCaffrey, 2013).
CLASS-S raters went through a formal training and certification period that required them to reach 80% agreement with master benchmark ratings. Certified raters return for additional training and calibration at a later date. The official protocol for CLASS-S requires raters to view a lesson for about 15 minutes, provide ratings, view the next 15 minutes of the same lesson, and provide another set of ratings. The ratings for each 15-minute segment are averaged to produce a score for the lesson. Rater agreement tends to be high in operational scoring (Bell et al., 2012) and generalizability studies indicate an index of dependability (i.e., phi-coefficient) that ranges between 0.5 and 0.63 when scores are averaged over four lessons (Bill and Melinda Gates Foundation, 2012).
Experimental Conditions
We designed an experiment to study the effect of observation length and order of presentation on the reliability and validity of CLASS-S scores. We randomly assigned two raters to one of four conditions. The first condition had raters judge a single 40-minute lesson. We call this condition the 1 × 40 condition. The second condition divided the lesson in half and raters observed the first 20 minutes, rated the quality of teacher–student interactions, observed the next 20 minutes of that lesson, and provided another rating. We refer to this condition as the 2 × 20 ordered condition. The third condition also involved two 20-minute segments of videotapes. However, in this case we randomized the order of presentation; raters watched 20-minute segments in random order and rated each segment. Randomization was at the segment level, not the teacher level. As such, it was unlikely for raters to observe all 40 minutes of a classroom in consecutive order. They may have watched the second 20 minutes of a lesson and then seen videotapes from other teachers before rating the first 20 minutes. We refer to this condition as the 2 × 20 random condition. Finally, we included a condition with four 10-minute segments assigned to raters in an ordered fashion. This fourth condition was like the 2 × 20 ordered condition, but it involved shorter and more frequent segments. We refer to this condition as the 4 × 10 ordered condition.
According to Kane’s (1982) sampling model of validity, these experimental conditions could be considered facets of the target universe and evaluated as variance components in a generalizability study. This would be the desired approach if these experimental conditions represented random facets. However, teaching observations typically treat observation length and presentation order as fixed facets. Therefore, we conducted a separate generalizability theory analysis for each condition. To determine if our experimental conditions represent important sources of variance, we compared results from each condition using bootstrap procedures.
Analysis
Generalizability theory (Cronbach et al., 1972) provides the analytic framework for examining reliability of scores for each of the four study conditions. Generalizability theory is well-suited to teaching observations that are influenced by multiple sources of variance (e.g., raters, segments, lessons, and teachers), and it has been applied to observational measures of education settings in order to specify and estimate salient sources of variance in observed scores and to identify observation procedures that minimize sources of error and optimize the reliability of the measure (e.g., Erlich & Borich, 1979; Hintze, 2005; Marcoulides, 1989a; Mashburn, Downer, Rivers, Brackett, & Martinez, 2013; Meyer, Henry, & Mashburn, 2011).
In applying generalizability theory, variance components are estimated in a generalizability study (G-study), and these components are combined to estimate error variance and reliability in a decision study (D-study). In the design of this G-study, we treat teachers as the object of measurement, with lessons (l) nested within teachers (
In the three study conditions for which we created multiple segments for each lesson (the two 2 × 20 conditions and the 4 × 10 condition), two raters (r) observed every segment of each lesson for every teacher. Therefore, raters were crossed with the other facets of the design. Considering the object of measurement (teachers) and all three facets, the universe of admissible observations corresponds to a design with raters crossed with segments nested within lessons within teacher or a
We estimated variance components for the universe of admissible observations using the MIVQUE estimation method as implemented in SAS version 9.2. In the D-Study, we then computed relative error 1 variance and the generalizability coefficient (i.e., a reliability estimate) for the universe of generalization.
The generalizability coefficient is
where the symbol n denotes the decision study sample size for the facet indicated by the subscript. This expressions reduces to
for the
We can determine the best approach for minimizing error variance in the
and
Assuming that all variance components are greater than zero, the following statements can be made about the effect of facet sample size on relative error variance. First, notice that Equation 3 contains the variance component in Equation 2 plus two additional components. Therefore, increasing the number of lessons will always result in lower relative error variance than increasing the number of segments. Second, increasing the number of lessons will produce lower relative error variance than increasing the number of raters when
Finally, increasing the number of segments will produce lower relative error variance than increasing the number of raters when
These inequalities are useful when interpreting the results of a decision study for a
To study the influence of our experimental conditions on validity, we analyzed the data in a variety of ways. First, we conducted two unplanned auxiliary analyses based on results from the G-study to explore reasons for observing changes in variance components across the conditions of our study. We focused on rater and carryover effects as possible sources for these changes. Next, we evaluated the impact of our conditions on CLASS-S domain scores through ANOVA methods, follow-up procedures, and correlations. Finally, we evaluated predictive validity through a multilevel model of general education student test scores that were standardized within subject and grade level. The first level of the model included a random intercept and nonrandom effects that accounted for gender, minority status, grade level, and prior achievement. The second level model for the intercept included a CLASS-S domain score, which was the main effect of interest. We conducted a separate multilevel analysis separately for each CLASS-S domain, given the high correlation among the three scales. We used the xtmixed command in Stata Version 12 to conduct the multilevel analysis.
Results
G-Study
Table 1 presents sources of variance estimated for each study condition for each domain of the CLASS. Across all conditions, teachers are the largest source of variance in CLORG scores. Teacher variance is also the largest source of variance in EMSUP scores in the 1 × 40 condition. This result is desirable as teacher variance in the G-study becomes universe score variance in the D-study. In all other conditions, the largest source of variance is due to raters, the rater by lesson within teacher component, or the residual. Rater variance is not much of a concern in the current study because it does not contribute to relative error variance in the D-study. However, the rater by lesson within teacher component and the residual component are problematic. They both contribute to relative error variance in the D-study as do all other variance components that are not teachers or raters.
Variance Component Estimates (Percentage of Total in Parentheses) for Each Condition.
For two of these variance components in Table 1, there is a discernible and interesting pattern across conditions. Specifically, in the 2 × 20 and 4 × 10 ordered conditions, the rater by lesson within teacher component accounts for a large portion of variance in EMSUP, INSUP, and CLORG scores, but the segment within lesson within teacher component does not. The opposite result is evident in the 2 × 20 random condition; segment within lesson within teacher accounts for a large portion of variance, but the rater by lesson within teacher only accounts for a small portion. We highlighted this pattern with bold font in Table 1. These results suggest that when segments are viewed and rated in immediate succession, raters’ scores on the second segment are rarely different from their scores in the first segment. Consequently, there is little segment variance but substantial variance in the rater by lesson within teacher component. In contrast, this effect diminishes when raters view segments in a random order that may be separated by days or months of time. In this case, raters are more likely to base their score on the actual segment being viewed and change their score accordingly when viewing segments in a random order. As a result, there is very little rater by lesson within teacher variance and a substantial amount of segment within lesson within teacher variance when segments are randomized.
As a follow-up test of the idea that raters rarely change their score when segments are viewed sequentially, we computed the correlation between segments in the 2 × 20 ordered condition and the correlation between segments in the 2 × 20 random condition. EMSUP, INSUP, and CLORG between segment correlations are 0.76, 0.64, and 0.74, respectively, in the 2 × 20 ordered condition, but they are only 0.47, 0.40, and 0.48 in the 2 × 20 random condition. Correlations in the random condition are all significantly lower than those in the ordered condition at a significance level of .001.
We also explored the idea that the large rater by lesson within teacher variance component in the 2 × 20 ordered condition was due to rater drift. To study this effect, we averaged segment scores to obtain a score for each lesson. We then computed the difference in days between the start of scoring for the study and the date a rater scored a lesson. We plotted the results for each domain and added a local linear regression line for each rater. Figure 1 shows a prominent decline in EMSUP and INSUP scores for both raters in the 2 × 20 ordered condition. Conversely, rater scores at the beginning and end of the scoring period in the 2 × 20 random condition are much more similar, although there are fluctuations over the time period. In both conditions, rater scores for CLORG are similar and stable over time. These results suggest that some of the rater by lesson within teacher variance in the 2 × 20 ordered condition could be due to rater trends. It seems that raters become more severe in their ratings as time goes on, perhaps due to fatigue. As a result, viewing lessons in order means that some teachers have both video tapes viewed when raters tend to give low ratings. The random condition counteracts this effect because teachers likely have at least one lesson viewed when raters are lenient and one rated when they are severe.

Rater trends over the duration of the scoring period. The x-axis of each plot is the rank of the number of days between the time scoring begins and when a video is scored. The top two panels are for the 2 × 20 ordered condition and the bottom two panels are for the 2 × 20 random condition.
To summarize the results of the G-study, our results provide evidence that support three findings. First, raters rarely change their scores from one segment to another when segments are viewed in sequential order and this contributes to rater by lesson within teacher variance. Second, raters also appear to trend downward as the rating period progresses when segments are viewed sequentially. Finally, randomizing segments seems to prevent the carryover from one segment to another and it also appears to mitigate rater trends over time.
D-Study
Variance components from the random condition and two ordered conditions (Table 1) mainly differed on the segment within lesson within teacher component and the rater by lesson within teacher component. As illustrated in the previous subsection, the way these two sources of variance changed in each condition yields insight into two rater processes—carryover effects and rater drift—that are not intended to influence scores. However, both these variance components are part of relative error and the change in magnitude across conditions may not make much of a difference in terms of relative error and the generalizability coefficient.
To investigate the impacts of these variance components on relative error and the generalizability coefficients, D-study results in Table 2 indicate that relative error variance for EMSUP and INSUP scores are lower in the 2 × 20 and 4 × 10 ordered conditions, but relative error variance for CLORG is lowest for the 2 × 20 random condition. However, universe score variance follows this same pattern, which leads the 2 × 20 random condition to have the highest generalizability coefficient estimates. Whereas the lowest generalizability coefficient was for the 4 × 10 ordered condition. Bootstrap confidence intervals 2 for the pairwise comparison of relative error variances between conditions indicate that the multiple segment conditions tend to have significantly lower relative error variance than the 1 × 40 condition. This result varies by domain, but it is consistent enough to suggest that considerable reductions in relative error variance are achieved by using multiple segments. However, large error variance in the 1 × 40 condition is also accompanied by large universe score variance leading the generalizability coefficients that were never the smallest in any condition. The results also suggest that length of observation is an important source of variance that could be included as a facet in the universe. Instead of narrowing the universe to inferences about 10- or 20-minute observations, observation length could be treated as random and incorporated into the measurement procedure.
Decision Study Results for all Conditions.
Note. Facet sample sizes are the same as those in the generalizability study.
Significantly lower than error variance for the 1 × 40 condition; 95% confidence interval is (.009, .104).
Significantly lower than error variance for the 1 × 40 condition; 95% confidence interval is (.002, .109).
Significantly lower than error variance for the 1 × 40 condition; 95% confidence interval is (.033, .197).
Significantly lower than error variance for the 2 × 20 random condition; 95% confidence interval is (.026, .170).
Significantly lower than error variance for the 1 × 40 random condition; 95% confidence interval is (.013, .137).
Significantly lower than error variance for the 1 × 40 random condition; 95% confidence interval is (.010, .162).
Although reliability estimates are highest in the 2 × 20 random condition, the question arises as to whether increasing the number of lessons, segments, or raters will produce reliability estimates that are more favorable for other conditions. We return to the inequalities discussed earlier to answer this question. We showed that increasing lessons is always better than increasing the number of segments. Using the inequality in Equation 5 and the variance components in Table 1, it is also always better to increase the number of lessons instead of the number of raters.
The inequality in Equation 6 is particularly important with respect to our experimental conditions because this inequality involves the rater by lesson within teacher component and the segment within lesson within teacher component. The relative magnitude of these two components differed in the 2 × 20 random and 2 × 20 and 4 × 10 ordered conditions. Thus, the decision to increase the number of segments or raters depends on the estimated variance components. Using the inequality in Equation 6 and the variance components in Table 1, increasing the number of segments in the random condition leads to lower relative error variance than increasing the number of raters. The opposite is true in the 2 × 20 ordered and 4 × 10 ordered conditions; increasing the number of raters leads to lower relative error variance. Thus, lower relative error variance in the 2 × 20 random condition is achieved by increasing the number of 20-minute segments, but in the 2 × 20 and 4 × 10 ordered conditions, lower relative error variance is achieved by increasing the number of raters. The implication is that the method for presenting segments to raters affects the choice of which facet sample size to increase.
To briefly recap the D-study findings, our analysis supports four important results. First, significantly lower relative error variances are frequently achieved in all domains by rating multiple segments instead of rating a single 40-minute lesson. Second, rating sequential 10-minute segments produced the lowest generalizability coefficient. This suggests that a 10-minute observation may not be sufficient for an observer to notice and rate true characteristics of teacher–student interactions. There simply may be inadequate time to see a complete and scorable interaction during a 10-minute period. Third, randomizing the order of segment presentation leads to higher, albeit not significantly different, reliability estimates. Finally, increasing the number of 20-minute segments is a better choice than increasing the number of raters when segments are presented in random order but the opposite is true when segments are presented sequentially.
Relationships Among CLASS-S Domain Scores
Table 3 shows that domain score means are rather variable across conditions. Indeed, significant and moderately sized differences exist for EMSUP
Correlations and Descriptive Statistics.
Bonferroni Adjusted Confidence Intervals for Pairwise Comparisons.
Note. Type I error adjustment was 0.05/6/2 = .0042.
Statistically different from zero.
For EMSUP and INSUP, the highest correlations are between the 1 × 40 and 2 × 20 random conditions (see Table 3). The correlation is also high (0.90) between these conditions for CLORG but not the highest. Correlations among the other conditions are also high for each domain. Indeed, no correlation is less than 0.65. In each domain, the lowest correlations typically involve the 4 × 10 condition. These results mean that rank ordering of teachers on the basis of observed scores is fairly similar for all domain scores between the 1 × 40 and 2 × 20 conditions, but slightly different for all domain scores in the 4 × 10 condition.
Predictive Validity Analysis
The predictive validity study focused on math and reading teachers that had student level data. Math teachers taught in Grades 6 through 11. Almost all math teachers taught a single class, but one teacher taught three. Among the 14 teachers with student-level math test data, class sizes ranged from 2 to 36, with 20 students being the typical class size. Table 5 lists descriptive statistics for math test scores by grade. Reading teachers with student-level data taught Grade 6, 7, 8, 11, and 12. Class sizes ranged from 1 to 36 students. The teacher with only one student also taught a second class with 17 students. Another teacher with only 3 twelfth-grade students also taught a class of 36 eleventh-grade students. Like math, the typical class size was about 20 students. Table 5 also lists descriptive statistics for reading scores.
Descriptive Statistics for Student Math and Reading Scores by Grade.
For the combined sample, a majority of students are female (55.6%) and most of them represented White (73.24%), Black (19.94%), Hispanic (3.89%), and Asian (1.76%) backgrounds. Native Americans, Hawaiians, and biracial students each represent less than 1% of the sample. The teacher sample consists of males (44.4%) and females (55.6%) who are White (91.1%), biracial (6.7%), or Black. They taught an average of 8.25 years. Sixty percent have a bachelor’s degree or 1 year beyond the bachelor’s level, 31% have a master’s degree, and the remaining teachers have an Educational Specialist of Doctor of Philosophy degree.
Table 6 lists results from the multilevel analysis of math test scores. The analysis takes into account demographic variables and prior achievement, but Table 6 only lists the effects of interest. Across all conditions, the 1 × 40 condition has the smallest coefficients, whereas the 4 × 10 always has the largest coefficients. For all but the 4 × 10 condition, the coefficients appear to be fairly similar. In terms of class domains, EMSUP and CLORG are significant predictors of math achievement in almost all conditions. EMSUP is not a significant predictor of achievement in the 1 × 40 condition, and CLORG is not a significant predictor of math achievement in the 1 × 40 or 4 × 10 condition. INSUP is never a significant predictor of math achievement.
Multilevel Model Estimates for CLASS Scales Predicting student Math Test Scores.
Note. Fixed effects for grade, prior achievement, minority status, study year, and gender are not listed.
Results for reading achievement are more variable across conditions (Table 7). The same patterns seen in the math results are not present in the reading coefficients. Moreover, the coefficients are more variable across conditions for reading scores than they were for math scores. CLASS-S domain scores are significant predictors of reading achievement in all but two conditions. INSUP is not a significant predictor of reading achievement in the 4 × 10 condition and CLORG is not a significant predictor in the 2 × 20 random condition.
Multilevel Model Estimates for CLASS Scales Predicting Student Reading Test Scores.
Note. Fixed effects for grade, prior achievement, minority status, study year, and gender are not listed.
Overall, presentation order and segment length do not appear to diminish or enhance predictive validity results. Coefficients do appear to be larger for math scores when lessons are broken into multiple segments, but multiple segments do not seem to make much of a difference in the coefficients for reading scores. Note that the observed lessons are not tied to a specific subject, but the outcomes of interest in the predictive validity study are test scores in a particular subject. It is possible that CLASS-S scores would be more predictive of student achievement when observations are limited to the subject of interest.
Discussion
Teaching observations are increasingly being used in education policy, research, and professional development. As a result, there is a need to understand and improve the psychometric properties of these measures, particularly if scores are to be used for high-stakes purposes, such as evaluating teachers’ performance. This study examined how various procedures for using the CLASS-S (Pianta et al., 2008)—a commonly used observational measure of the quality of teachers’ interactions with children in 6th grade to 12th grade classrooms—affected the reliability and validity of scores. Specifically, from three 40-minute videotaped lessons collected from 47 teachers, we manipulated the length of observation and order of presentation of the lessons in four different ways, and we randomized two raters to observe and rate all lessons from all teachers.
A generalizability study, decision study, and additional analyses of the validity of scores were conducted to contrast the reliability and validity for each study condition. The generalizability study estimated multiple sources of variance in scores related to rater, teacher, lesson, and for the three study conditions that decomposed lessons into multiple occasions, segment. Although there are no appreciable differences in the financial costs of implementing the four different operational procedures under study, there were notable differences in some aspects of the reliability and validity of scores related to segment length and/or order of presentation.
Specifically, results indicated that lessons rated in the shortest and most frequent manner (4 × 10 minute segments) produced the lowest generalizability coefficients for all three domains of the CLASS-S. Although this condition produces the highest number of occasions of measurement per lesson, which is favorable in reducing error attributable to segments that is part of the relative error estimate, results indicated that this condition was characterized by the lowest universe score variances and lowest generalizability coefficients. This suggests that 10 minutes may not be adequate time to observe the specific indicators of teaching quality that inform raters’ judgments about the quality of teaching. As a result, raters likely use extraneous information when judging the quality of a lesson, which subsequently reduces universe score variance and the generalizability coefficient.
Procedures involving single ratings made following 40-minute lessons resulted in domain scores that were highly correlated with scores from the other three observation conditions, and in generalizability coefficients that are similar to those from observations of 20-minute segments. However, scores from the 1 × 40 condition tended to suffer from large relative error variance and low predictive validity coefficients. Among the two conditions in which raters assigned scores following 20-minute segments, there was no evidence that the order in which the 20-minute segments were presented to raters significantly affected the reliability of scores. However, randomly presenting 20 minute segments to raters from the entire pool of 282 segments per teacher had the advantage of reducing sources of construct irrelevant variance by reducing carry over effects and rater drift.
It is important to consider some limitations with these results. Table 1 shows that a large source of variance is often attributed to the residual effect. This result suggests that conditions of measurement outside of those studied in this article are influencing scores in a notable way. We have heard raters report teaching methods and classroom management features that affect their ability to provide ratings. For example, they may view a classroom while students are completing a worksheet leaving little opportunity to observe and rate teaching. These characteristics are not easily classified by teaching observation measures and likely contribute to a large residual variance term. Unfortunately, our study was unable to address this possibility, but it is an area worth further study. Until more is known about conditions that contribute to large residual variance terms, practitioners should ask raters to provide annotations and notes about challenges encountered when using the rating scales. This could inform a revision to the measures or a standardization that reduces this source of variance.
In sum, results indicate that operational procedures related to length of observation and order of presentation can impact the reliability and validity of scores, while adding few financial costs for conducting teaching observations. Given the growing importance of teaching observations, further research is needed to understand the tradeoffs between reliability and validity related to the operational procedures under conditions of different instruments, frequencies of observations, ordering of presentation, and modality of collecting data (live and videotaped).
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by grants from the William T. Grant Foundation.
