Abstract
This study used generalizability theory to measure reliability on the Recognizing Effective Special Education Teachers (RESET) observation tool designed to evaluate special education teacher effectiveness. At the time of this study, the RESET tool included three evidence-based instructional practices (direct, explicit instruction; whole-group instruction; and discrete trial teaching) as the basis for special education teacher evaluation. Five raters participated in two sessions to evaluate special education classroom instruction collected from two school years, via the Teachscape 360-degree video system. Data collected from raters were analyzed in a two-facet “partially” nested design where occasions (o) were nested within teachers (t), o:t, and crossed with raters (r), {o:t} x r. Results from this study are in alignment with similar studies that found multiple observations and multiple raters are critical for ensuring acceptable levels of measurement score reliability. Considerations for the feasibility of practice should be observed in future reliability and validity studies on the RESET tool, and further work is needed to address the lack of research on rater reliability issues within special education teacher evaluation.
Keywords
Recent changes in federal requirements have compelled states to change their teacher evaluation practices (McGuinn, 2012; U.S. Department of Education, 2012). These changes have shifted teacher evaluation systems into a multiple methods approach that include the use of observations, measures based on student performance (e.g., value-added models), and other tools that are used to assess teacher “effectiveness” (Goe & Croft, 2009). This shift has generated new bodies of evidence regarding the reliability and validity of teacher effectiveness measurements; particularly those based on student outcomes and observation tools. However, current research has typically been restricted to teachers of tested subjects (i.e., standardized state assessments), such as the Measures of Effective Teachers (MET) study (Ho & Kane, 2013; Kane & Staiger, 2012), or in work specific to a content area, like math (Casabianca et al., 2013; Hill, Charalambous, & Kraft, 2012). This tendency to focus on teachers of tested subjects, while omitting what Prince et al. (2009) referred to as the “other 69%,” has left empirical gaps in teacher evaluation research for content areas like physical education, art, English language learners, and special education.
Observation tools and measurement score rater reliability issues and recommendations for teachers of nontested subjects continue to remain largely unaddressed (Holdheide, Browder, Warren, Buzick, & Jones, 2012; Jones, Buzick, & Turkan, 2013; Semmelroth, Johnson, & Allred, 2013; Sledge & Pazey, 2013). Recent research in teacher evaluation (e.g., MET) has focused on using a variety of different evaluation methods to measure teacher effectiveness, but these studies have been restricted to specific content areas (e.g., math, reading) and student ability level (e.g., general education; Kane & Staiger, 2012). The exclusion of other content areas and student populations is problematic because more than ever before, states are relying on evaluation tools to measure the effectiveness of all teachers (National Council on Teacher Quality, 2012). Despite the fact that most of the research on the reliability and validity of teacher evaluation tools in use is only specific to a certain group of teachers (Goe & Holdheide, 2011; Semmelroth et al., 2013; Sledge & Pazey, 2013), states have rapidly adopted new requirements for teacher evaluation systems using this evidence. For example, in 2009, only four states were using student achievement as an important criterion in how teacher performance was assessed, but in 2012, that number had increased to 22 states (National Council on Teacher Quality, 2012).
Current studies on observation tools to evaluate teacher effectiveness have examined rater reliability issues like the use of trained raters, and the ideal number of raters and observations (i.e., multiple raters and observations) to achieve measurement score reliability. For example, the results from the MET study inform issues related to rater reliability, including the use of multiple raters to ensure acceptable reliability, the use of “same school” versus “other school” administrators, and the use of certified peer raters (Ho & Kane, 2013; Kane & Staiger, 2012). Similarly, the results from Hill, Charalambous, and Kraft’s (2012) study on the Mathematical Quality of Instruction (MQI) examined rater reliability issues related to length of video observation, how the quantity of items on an observation tool might affect raters’ measurement reliability, and the number of raters and observations needed to maintain acceptable levels of measurement score reliability.
As a result, two of the most important findings from recent studies like the MET and on the MQI are (a) the use of multiple raters is critical for achieving acceptable levels of measurement score reliability, and (b) high-quality classroom observations require certified raters (Hill, Charalambous, & Kraft, 2012; Kane & Staiger, 2012). While these criteria for rater reliability compel change within our current context of teacher evaluation reform, as in the use of multiple raters and occasions, the complexities are intensified when considered within the context of special education. Special education teachers work in highly specific but diverse instructional environments, and it will be essential for any teacher evaluation system to have certified raters that are knowledgeable about evidence-based practices (EBPs), special education policies and procedures, and the wide range of roles and responsibilities a special education teacher must assume. Unfortunately, recent research on observation tools and rater reliability has been specific to teachers of tested subjects, and has not sufficiently considered special education. Thus, further research is needed on rater reliability issues that take into account the complexities found in the field of special education teacher evaluation.
The purpose of this study was to extend current research on rater reliability to the special education teacher evaluation context. We believe greater attention must be given to the complex measurement challenges involved in special education teacher evaluation. For teacher evaluation reform to be successful in ensuring teacher quality and promoting professional development (Danielson, 2011), a “one size fits all” approach is not practical nor effective for the field of special education (Benedict, Thomas, Kimerling, & Leko, 2013; Holdheide et al., 2012; Jones et al., 2013; Semmelroth et al., 2013).
In this study, measurement score rater reliability was examined using generalizability theory on the Recognizing Effective Special Education Teachers (RESET) observation tool. Five teacher coders (raters) participated in two sessions (October 2012 and April 2013) to evaluate special education classroom instruction collected from the 2011–2012 and 2012–2013 school years, via the Teachscape 360-degree video system. The raters were trained on the RESET observation tool, and participated in whole-group coding sessions to establish interrater agreement before evaluating assigned videos. The primary research question guiding the generalizability and decision study analyses in this study was “How many occasions and raters are needed for acceptable levels of measurement score reliability when using the RESET observation tool to evaluate special education teachers?” The results from this study will hopefully contribute to the growing body of research on special education teacher evaluation, issues related to rater reliability, and the use of observation tools to evaluate teacher effectiveness.
Special Education Teacher Evaluation Challenges
For special education, the stakes are especially high for developing what Danielson (2011) maintained are the two critical features of any teacher evaluation system: (a) ensuring teaching quality and (b) promoting professional development. Given the current state of the special education teacher profession, an effective special education teacher evaluation system must be able to recognize and address the unique systemic challenges that special education teachers face (Boe, Cook, & Sunderland, 2008; Holdheide et al., 2012; Spooner, Algozzine, Wood, & Hicks, 2010). Accordingly, to meet the needs of legislative policy (U.S. Department of Education, 2012), as well as the Council for Exceptional Children (2012) recommendations for teacher evaluation, an effective special education teacher evaluation system should be characterized with features that allow for (a) the evaluation of high-quality and evidence-based instructional techniques, (b) the measurement of teacher effectiveness using some measure of student growth or achievement, and (c) the flexibility to accommodate a variety of teaching contexts. Unfortunately, current teacher evaluation methods used for special education teachers do not support this theory of action (Council for Exceptional Children, 2012; Holdheide, 2012; Semmelroth et al., 2013) as there remains a gap in teacher evaluation research that is specific to special education teachers.
Special education teachers work under a variety of conditions (Stempien & Loeb, 2002), serve a heterogeneous population with diverse needs (Tyler, Yzquierdo, Lopez-Reyna, & Flippin, 2004), and historically have not entered the profession well prepared (Connelly & Graham, 2009; McLeskey & Billingsley, 2008). In addition to the professional challenges the field of special education faces, it has steadily seen a decrease in classroom instructional quality and time due to growing systemic demands (Russ, Chiang, Rylance, & Bongers, 2001; Vannest & Hagan-Burke, 2009). Special education teacher time has been compromised with additional duties like case management, testing, progress monitoring, paperwork, meetings, and management of support staff that most general education teachers are not required to complete, leaving as little as 20% of a special education teacher’s time dedicated to instruction (Vannest & Hagan-Burke, 2009). As a result, outcomes for students with disabilities (SWD) are not encouraging. Nationally, as few as 30% of SWD are able to meet performance standards (Odom, 2009), and postschool outcomes for SWD indicate SWD are less likely to attend college, earn as much, and live independently than their typical peers (Newman et al., 2011).
To improve the outcomes for SWD, the instructional practice of special education teachers must be improved (McLeskey, 2011; Morgan, Frisco, Farkas, & Hibel, 2008; Nougaret, Scruggs, & Mastropieri, 2005; Scruggs, Mastropieri, Berkeley, & Graetz, 2009). Promisingly, the field of special education research has a strong foundational knowledge base about evidence-based instructional practices that can be used to improve the current state of the profession (Baker, Chard, Ketterlin-Geller, Apichatabutra, & Doabler, 2009; Cook & Odom, 2013; Gersten et al., 2009; Odom, 2009; Odom et al., 2005; Smith, Schmidt, Edelen-Smith, & Cook, 2013). Despite this vast body of research on instructional practices that have been found to be highly effective for SWD, implementing these practices within the special education classroom has been problematic (Cook & Odom, 2013; Odom et al., 2005).
RESET Observation Tool
In Idaho, the RESET observation tool is being developed to evaluate special education teacher effectiveness through a teacher’s use of evidence-based instructional practices. The RESET observation tool is guided by the idea that increased use of effective evidence-based instructional practices will lead to increases in student outcomes (Cook & Odom, 2013; Odom et al., 2005; Odom, Collet-Klingenberg, Rogers, & Hatton, 2010), which is one of the primary goals of current teacher evaluation reform (Lewis & Young, 2013; National Council on Teacher Quality, 2012; Sledge & Pazey, 2013).
To measure special education teaching effectiveness, the RESET observation tool consists of descriptions and associated key characteristics of various evidence-based instructional practices (Johnson & Semmelroth, 2012). An evaluator using the tool does not have to be present in real time to evaluate an observed teacher. Instead, a special education teacher’s classroom instruction can be recorded using the Teachscape 360-degree video capture system. Trained raters can then observe the recorded session, and the special education teacher receives a score on how well the various instructional components have been implemented. In the same way, the RESET observation tool provides feedback to the observed teacher regarding his or her delivery of evidence-based instruction, and receives clear guidance on how to improve (rather than just receiving a score). The tool is flexible enough to be used across various special education settings, but specific enough in its focus on EBPs to provide targeted, specific feedback for teachers. The evaluative rubric for the RESET tool is aligned with Charlotte Danielson’s (2013) Framework for Teaching numeric rating scale (1–4). In this way, the RESET tool can be readily transferred for use at the district level in revised teacher evaluation systems that have shifted to a multiple methods approach.
Reliability
Of greatest concern in the development of a special education teacher evaluation system that relies on the observation of instructional practice is achieving acceptable levels of reliability in the ratings assigned. This is of significant concern if the evaluation will be used to make important decisions about special education teacher quality, such as driving professional development agendas, creating performance improvement plans, and/or high-stakes decisions related to compensation or tenure. Researchers studying teacher evaluation systems that include an observational component have indicated that different combinations of multiple raters, multiple observations, and extensive training of raters are required to reach acceptable levels of rater reliability, as in Hill, Charalambous, and Kraft’s (2012) study on the MQI and the MET study (Ho & Kane, 2013; Kane & Cantrell, 2013; Kane & Staiger, 2012). This is critical information for guiding the development of the evaluation system and has implications for how such a system could be practically implemented.
The reliability evidence required to ensure an observation evaluation system is appropriate for relative (e.g., performance improvement plans) and absolute (e.g., tenure) decision making must sufficiently establish rater reliability, determine the optimal number of observations, and examine the internal consistency of the ratings. However, traditional interrater agreement measures like kappa can be problematic for these types of decisions, because measures of reliability using interrater agreement only attend to only one source of variation (the rater) overlooking other sources of variation (e.g., teachers, occasions, items, etc.) that affect the consistency of evaluation scores within observations (Brennan, 2001; Cronbach, Gleser, Nanda, & Rajaratnam, 1972; Hill, Charalambous, & Kraft, 2012; Shavelson & Webb, 1991). For example, previous studies on the RESET observation tool found low to weak levels of agreements across raters using perfect agreement and kappa (see Landis & Koch, 1977, for a more detailed discussion on agreement measures) to measure observer agreement (Johnson & Semmelroth, 2012), but these agreement measures did not consider other sources of variance like number of occasions or teachers observed. Thus, a single score obtained on one occasion is not fully dependable (Shavelson & Webb, 1991), making the case for the use of generalizability theory to analyze multiple sources of variance in a rater agreement measurement. In fact, recent studies examining teacher evaluation instruments like the MET have used generalizability theory to estimate sources of error, and to improve the reliability of different “real-world” scenarios by varying the number and type of raters and the number and length of lessons (Ho & Kane, 2013; Kane & Cantrell, 2013).
Generalizability Theory
While perfect agreement and kappa analyses have been used to measure rater agreement and reliability for classroom and teacher observations, generalizability and decision studies are able to account for variability that traditional interrater agreement analyses cannot (Brennan, 2001; Cronbach et al., 1972; Hill, Charalambous, Blazar, et al., 2012; Tindal, Yovanoff, & Geller, 2010). The use of generalizability theory, or G theory, to inform the use of the RESET tool is important because it accounts for multiple sources of variance that cannot be measured in interrater agreement analyses like kappa, and reflects a much more realistic measurement of what happens in a classroom. The information produced in the generalizability study (G study) guides the exploration of optimal conditions for measurement score reliability in the decision (D study). The information from the D study provides facet conditions for increased measurement precision for future research, and guides practitioners regarding the basic conditions (i.e., number of raters and occasions) needed for acceptable levels of measurement score reliability.
In alignment with similar studies, we used a generalizability study design that included two facets (raters, occasions) and one unit of measurement (teachers), creating a two-facet, “partially-nested” design where occasions (o) (observations/lessons) were nested within teachers (t), o:t, and crossed with raters (r), ({o:t} x r) (Shavelson & Webb, 1991, pp. 52–54). Like the studies completed by Erlich and Shavelson (1978), and Hill, Charalambous, and Kraft (2012), this study aimed to identify the generalizability of measures of teacher behavior by systematically examining the effect of more than one facet (raters, occasions). Previous studies of teacher evaluation systems have indicated that sources of variance in scores obtained with observation tools typically come from the lessons (how much can we generalize a teacher’s ability from one lesson to the next), from raters (how much of the score is dependent on which rater evaluates the lesson), from teachers, and from error (Erlich & Shavelson, 1976; Hill, Charalambous, & Kraft, 2012; Ho & Kane, 2013; Shavelson & Dempsey-Atwood, 1976). This study adds to the current body of teacher evaluation and special education research because it is specific to (a) special education teacher evaluation, and (b) the reliability of a rater’s measurement score observations when evaluating a special education teacher based on his or her use of evidence-based instructional practices.
Purpose of the Current Study
The purpose of this study was to examine measurement score rater reliability using generalizability theory on the RESET observation tool. From the results of the generalizability studies, decision study analyses were also completed to identify optimal numbers of raters and teachers to maintain the highest levels of measurement score reliability when using the RESET tool. G theory was used to analyze sources of observed score variance because it decomposes variability into components (teachers, lessons, and raters), interactions, and measurement error (Brennan, 2001; Cronbach et al., 1972; Shavelson & Webb, 1991).
Method
Participants (Raters)
Five special education teachers were invited to participate as raters in two sessions (October 2012 and April 2013) to evaluate special education classroom instruction collected from the 2011–2012 and 2012–2013 school years via the Teachscape 360-degree video system. This study originally started with eight raters, but due to rater attrition caused by conflicting time commitments, only data from five raters over the course of two sessions were included for analysis. While not ideal, previous studies and generalizability theory explanations have established that smaller rater sample sizes are sufficient for research purposes (Erlich & Shavelson, 1978; Hill, Charalambous, & Kraft, 2012; Shavelson & Webb, 1991). For example, Erlich and Shavelson (1978) used a two-facet, nested measurement to study teacher behavior in which three raters evaluated five teachers who were videotaped on three occasions (i.e., consisting of different lessons). Hill, Charalambous, and Kraft (2012) similarly used a study design that included nine raters who evaluated eight teachers over three occasions. This study was conducted using five raters who evaluated nine teachers over three occasions. While acceptable for research purposes, the results of this study will hopefully contribute to futures studies on measurement score variance on the RESET observation tool. With that said, the results from this study should be interpreted with caution given the limited resources involved.
The raters were selected through communication with their district special education directors. Although the sample was one of relative convenience for this study, predetermined criteria were observed to ensure that invited raters represented a balanced sample of the range of content, placement, and grade found in special education, and that the invited raters had all completed a minimum of 5 years of certified teaching. Because this study details beginning work on a special education teacher observation tool based on the evaluation of evidence-based instructional practices, we believed it was important to include an experienced, well-balanced rater sample. This rationale is primarily due to the content knowledge and complex decision-making abilities required when evaluating special education teacher evidence-based instructional practices. In addition, a variety of teachers were included in this study to ensure that the rater sample was representative of the profession to avoid any possible biases that might have skewed the results (e.g., including only “elementary self-contained teachers” as raters). Table 1 provides rater demographics, including current teaching assignment, total years teaching, and highest level of education completed. All raters were female, and all raters worked in suburban districts (except Rater 1).
April 2013 and October 2012 Data Coding Rater Demographics (n = 5 Raters).
Rural.
Measures (RESET Observation Tool Subscales)
To analyze collected data, a two-facet, partially nested design ({o:t} x r) was used. All raters evaluated all videos, and all scores were initially aggregated at the lesson level by each rater (i.e., raters evaluated each occasion by each item in the RESET observation tool). However, following the work of Hill, Charalambous, and Kraft (2012), we used subscales from specific RESET observation tool items (questions) to examine sources of variance and measurement error to help identify levels of reliability on various components of the tool. To create the subscales used in the analyses in this study, items were grouped according to evaluative purposes, founded on evidence-based literature and rationales. Subscale 1 (Lesson Objective) was based on an evaluation of whether the lesson objective was evident to the students. Subscale 2 (EBP Implementation) was based on an evaluation of the EBP implementation, which consists of the scores received on the various criteria developed to evaluate the specific EBP. Subscale 3 (Whole Lesson Review) was developed as an overall evaluation of the lesson. The goal of Subscale 3 was to provide a broad evaluative score of the special education teacher’s performance throughout the lesson. All questions appearing on the RESET tool were scored using a numeric scale (i.e., a quantitatively defined rating scale from 0 to 3 that aligns with Danielson’s, 2013, 1–4 scaled rubric). Finally, similar to Hill, Charalambous, and Kraft’s (2012) averaging scores across dimensions, in this study, we created holistic scores for each subscale and used these subscale scores for the G study analysis.
Rater Training
The raters were trained on the RESET observation tool, and participated in whole-group coding sessions to establish interrater agreement before evaluating assigned videos. At the time of these data collection sessions, the RESET observation tool included three evidence-based instructional practices (direct, explicit instruction; whole-group instruction; discrete trial teaching) as the basis for special education teacher evaluation. Each rater was provided with two university-owned laptops for use: one to watch the assigned Teachscape videos and one to complete the observation tool via Qualtrics. For each session, raters were given a half-day training presentation, followed by each rater individually evaluating two videos for the purposes of calibration and measuring interrater reliability. A 45-page user manual was provided to explain the structure and features of the RESET tool. The manual also includes operationalized definitions and descriptions of the three evidence-based instructional practices, and the evaluation rubrics for all ratings on the RESET observation tool. Interrater agreement achieved during the training sessions ranged from .72 to .95, measured as a holistic score and by each subscale (see Table 2). After calibration videos were completed, the first author reviewed the videos and subsequent scores with the raters to reach consensus on the assigned evaluations.
Results of Interrater Agreement Compared Against Master Ratings (n = 5 Raters).
Setting
For the October 2012 and April 2013 data coding sessions, raters were hosted on the Boise State University campus. The sessions were designed to protect the confidentiality of the teachers appearing in the video observation data. Raters were seated away from one another, and were given headphones to wear throughout the sessions to prevent any sharing of rater visual or audio information. Raters evaluated video observation data that were collected via the Teachscape Reflect system, the same technology used by the MET study (Kane & Staiger, 2012). The Teachscape video capture system consists of two cameras: (a) a 360-degree camera that allows the observer to pan and zoom on various components of the classroom environment and (b) a fixed position camera, also referred to as a “board cam” because it is usually focused on a classroom board.
Video Data Files
Video data used for the rater sessions were collected across five school districts from 21 different teachers over the course of two school years (2011–2012 and 2012–2013). The mean time of each video was 25 min, with videos in the data set ranging from 17 to 72 min. Although the length of each video varied in length, the video observation represented what each observed teacher self-identified as one “lesson.” Due to limitations in the data set (i.e., some teachers had unusable observations, or all raters were unable to evaluate all assigned videos of one teacher), only 9 teachers (i.e., unit of measurement in the G study) were included in this study. As a result, the study design included five trained raters who evaluated 9 different teachers with three occasions per teacher.
Results
This generalizability study examined measurement score rater reliability using the RESET observation tool. Five raters were trained to use the RESET observation tool to evaluate video observations of special education classroom instruction captured via the Teachscape system. A two-facet, “partially nested” design was used and rater data were analyzed using the EduG v. 6.1 software program. Results are organized according to the three subscales: Lesson Objective (Subscale 1), EBP Implementation (Subscale 2), and Whole Lesson Review (Subscale 3).
Sources of Variance
Table 3 includes the results of the ANOVAs and is organized by each subscale, followed with a condensed table of just the variance decomposition for all three subscales for ease of comparison. The ANOVA table is organized by each facet or each facet interaction (the source of variation) and includes the sums of squares (SS); degrees of freedom (df); mean squares (MS); percentage contribution of each source to the total variance, that is, the sum of the corrected variance components (% of total variance); and the standard error associated with each variance component (SE). The variance component for teachers (
ANOVA for Subscales 1 to 3.
Note. SS = sums of squares; MS = mean squares; EBP = evidence-based practice.
The variance component for raters (
Similar to Erlich and Shavelson’s (1978) study on teacher behavior, in our study, there were multiple occasions for each teacher, and the occasions were different from teacher to teacher (Shavelson & Webb, 1991). However, because the variance component for occasions is nested within teachers (
The variance component for the interaction between teachers and raters (
Finally, the interaction between raters and occasions; the three-way interaction between teachers, raters, and occasions; and unaccounted/unmeasured variation are confounded in this two-facet, partially nested design. The residual component (
Variance Decomposition for RESET Subscales.
Note. RESET = Recognizing Effective Special Education Teachers; EBP = evidence-based practice.
While the
G Study Results
The G study was completed to determine the variance components attributable to teachers (t), occasions (o), and raters (r); their two-way interactions; and the combination of the three-way interaction and the measurement error. As with the ANOVAs, items from the RESET observation tool were collapsed into three subscales. Table 5 reports the results of the G study, including the source of variation (% absolute), total differentiation variance, standard deviation, total relative error variance, relative G coefficient, and absolute G coefficient. The % absolute source of variation reports how the absolute error variance is distributed among the other sources; the information from this result indicates the sources of variance that have the greatest negative effect on the precision of the RESET observation tool (Cardinet, Johnson, & Pini, 2010). In addition, these results also inform the design of the follow-up D study as it indicates which facet contributes the most to measurement error.
Generalizability Study Error Variance and G Coefficients for the RESET Observation Tool.
Note. RESET = Recognizing Effective Special Education Teachers; EBP = evidence-based practice.
For Subscale 1, the occasions (31.5%) and residual (38.5%) facets are the two largest contributors to measurement error. In addition, the raters (r) facet (4.8%) has the lowest value of all values reported. This suggests that questions related to the lesson objective might be easier for raters to identify, but is affected by unknown (residual) sources of error. For the other two subscales, however, the residual components are much lower: Subscale 2 (11.7%) and Subscale 3 (10.5%). Again, these lower residual values indicate that Subscales 2 and 3 have sources of error more evenly distributed among facets than Subscale 1.
The differentiation and relative error variances provide insight into whether a weak G coefficient (relative or absolute) is due to high measurement error, or just to minimal differences between the objects measured (Cardinet et al., 2010). These measurements provide a holistic indication of the reliability of the measurement procedure and give a general indication of each of the measurements’ precision. There is no agreed upon “cutoff” score for what might be considered a strong level of reliability versus a weak level of reliability. For example, Ho and Kane (2013) described a range of different scenarios to achieve reliabilities of .65 or higher in classroom observations, while Cardinet et al. (2010) considered a sample measurement of .78 as “not entirely satisfactory” (p. 53). Thus, as Brennan (2013) maintained, to really understand the value of a G coefficient, one must know the level of variance, what is most contributing to error, and to what extent these influences have in a given sample size.
Across all three subscales, the relative and absolute G coefficients follow the pattern found in the sources of variance and G study analyses. In addition, the coefficients for Subscales 2 and 3 might be affected by the differentiation variance (t), both having high reported values. Because of the difference between the differentiation variance and the relative error variance, the lower G coefficient values might be attributable to either measurement error or minimal differences between the objects measured (Cardinet et al., 2010). Although the reported G coefficients might initially be interpreted as less than desirable, the values do suggest that the measurement was not entirely inadequate, and that with a few modifications to facet sample sizes, more desirable levels of reliability might be obtained in the decision studies appearing in the next section (Brennan, 2001; Cardinet et al., 2010; Hill, Charalambous, & Kraft, 2012; Ho & Kane, 2013; Shavelson & Dempsey, 1975; Shavelson & Webb, 1991; Webb, Shavelson, & Haertel, 2006).
Decision Study Results
The decision study, or D study, procedure allows for the “what if?” analyses that develop through the interpretation ANOVA and G study results (Cardinet et al., 2010). D studies use information from a G study to design a measurement to reduce error for a particular purpose. For the D study procedures conducted in this article, the relative G coefficient and standard error of measurement (SEM) scores were recorded throughout the process when changing the rater (r) and occasion (o) facet size characteristics. The relative G coefficient generally corresponds to higher scores and is recommended for use in relative decision making (e.g., rewarding teachers for rated excellence), while the absolute G coefficient generally reports lower values and should be used for absolute decisions (e.g., firing teachers for rated unsatisfactory performance).
Table 6 shows the relative G coefficient and SEM scores across the three subscales. The rater and occasion facets were “optimized” using different sample sizes to obtain “optimal” levels of measurement score reliability (e.g., Cardinet et al., 2010). Figures 1 to 3 are the graphical representations of the relative SEM and measurement score reliability across raters and occasions by each subscale.
Relative G Coefficient and SEM for Decision Studies Comparing Occasions and Raters.
Note. SEM = standard error of measurement; EBP = evidence-based practice.
≥.65 relative G coefficient score.

Lesson objective D Study, raters and lessons, SEM and G coefficient.

EBP implementation D Study, raters and lessons, SEM and G coefficient.

Whole lesson review D Study, raters and lessons, SEM and G coefficient.
From Figures 1 to 3 it can be seen that as raters and lessons (occasions) increase, so too does the relative G coefficient, while the SEM steadily decreases. For all three subscales, while there are significant differences in reported values between the Lesson 1 and Lesson 2, and somewhat between the Lesson 2 and Lesson 3, the gaps between measurements are smaller between Lessons 3 and 6. This suggests that there might be a “happy medium” between empirical reliability and practical application somewhere between multiple raters observing two to four lessons. Similarly, while there are significant differences for all three subscales from Rater 1 to Rater 3, the increase seems to flatten out from Rater 3 to Rater 5. Like the differences between lessons, there seems to be a practical middle ground somewhere between two and four raters. This finding suggests that real-life applications of the RESET observation tool would not require ideal, research-like settings, but instead be able to more practically consider finite resources.
Implications for Practice
The findings in this study are in alignment with similar generalizability theory studies on observation tools to measure teacher behavior: To achieve acceptable levels of relative reliability and error when evaluating special education teachers, multiple raters and occasions must be used (Bell et al., 2012; Erlich & Shavelson, 1978; Hill, Charalambous, & Kraft, 2012; Ho & Kane, 2013; Kane & Cantrell, 2013; Medley & Mitzel, 1958; Shavelson & Dempsey, 1975). Facets reported as the highest and lowest sources of variance for Subscales 1, 2, and 3 suggest that there might be substantive differences between what each subscale is able to measure at this time. In the decision study analyses, as raters and occasions increase, levels of reliability correspondingly do as well, while the relative SEM decreases.
In answering the primary research question of this study, “How many occasions and raters are needed for acceptable levels of measurement score reliability when using the RESET observation tool to evaluate special education teachers?” conclusions must be considered with caution. Given the small rater size used in the G study, and the ongoing development of the RESET tool, inferences must be made with the understanding that further work is needed. Interpreting the relative G coefficient scores from this study depends first on which “cutoff” score one subscribes to. For example, Brennan (2013) does not provide specific benchmarks to determine measurement score reliability as he maintains that variance and error must be included in the determination of results, while Cardinet et al. (2010) lean toward the more traditional .80. Because this study contributes to a developing body of research on rater reliability using an observation tool designed for use in special education, we chose to use a less stringent cutoff score, and followed Ho and Kane’s (2013) use of .65 (to describe results from the MET study). Thus, using .65 as the cutoff relative G coefficient score, acceptable levels of measurement score reliability were found in Subscales 1 and 3, but not for Subscale 2 (see Table 6). In practical consideration of time and resources, the least amount of raters and occasions needed for acceptable levels of measurement score reliability for Subscale 1 are three occasions, four raters (.67), and for Subscale 3, four occasions and five raters (.67).
In addition, the use of at least four raters seems to be optimal (with the number of observations varying across subscales). While this finding is consistent with other generalizability theory studies on teacher observations that indicate more than one rater and more than one observation are needed for reliable evaluations (Hill, Charalambous, & Kraft, 2012; Ho & Kane, 2013; Kane & Cantrell, 2013), the use of four, trained raters may be impractical for districts. Across all subscales, low levels of measurement score reliability and high levels of error were reported when using just one to two of these conditions (raters and occasions). This empirically consistent finding suggests that future development on the RESET tool (and other observation-based systems) must plan for the use of multiple observations and raters to obtain acceptable levels of measurement score reliability. In consideration of limited resources, four observations per teacher per school year might be too resource intensive, and additional research is needed to determine ways to minimize error and increase measurement precision. One anticipated line of research to do this is to systematize the link between teachers and evaluators (raters). This type of study would require observed teachers to identify which instructional practice will be used before collecting the video observation data, and would improve overall levels of rater reliability.
The findings from this study seem to indicate that an overall evaluative judgment of special education teacher performance (Subscale 3) is more reliable than ratings on individual lesson components (Subscale 2). However, the lower levels of reliability reported in Subscale 2 might be due to the fact that instructional practice is an extremely complex activity that takes place over the course of time. As a teacher receives feedback from an evaluator using the RESET observation tool (i.e., evaluating the teacher based on his or her use of evidence-based instructional practices), it will be critical to measure the level of change within observations. While our study design assumes “trait invariance” within each occasion, future research should define the number of time periods needed to study changes in the state over time and the number of occasions within a time period using the RESET observation tool (see, for example, Meyer, Cash, and Mashburn’s, 2011, study that used multivariate generalizability theory to study changes over time within occasions).
Finally, because instructional practice can be a very complex activity, and because Subscale 2 is comprised of the essential building blocks of instructional practice, it lends itself most vulnerable to issues that influence rater disagreement. The RESET observation tool was developed in alignment with Danielson’s (2011) assertions that an effective evaluation system should ensure teacher quality and promote professional development. Even though the overall judgment of a teacher’s practice was found to be more reliable in this study (Subscale 3), it does not really address specific components of instructional practice. The higher levels of reliability found in Subscale 3 might be useful in assisting schools and districts with relative decisions, but it is the feedback found in Subscale 2 that will provide a teacher with targeted, specific feedback to improve components of evidence-based instructional practice.
Limitations
There were several limitations to this study. Although the use of five raters is acceptable for research purposes, it would be beneficial to include more raters in future studies. As the body of research in this area continues to grow, it will be important to explore other rater conditions (e.g., trained administrators vs. trained special education teachers) that can impact measurement score reliability on special education observation tools.
Also, there are some issues overall that might have affected the results of the generalizability and decisions study analyses completed in this study. First, the use of two data sets over a 6-month period may have led to a range of unaccounted sources of variance (e.g., differences in training sessions, different data sets, etc.). Second, the collapse of the evidence-based instructional components into one holistic score might have affected the results of the G study analyses. Because each component is defined through the review of literature specific to the instructional practice, the nuances of differences within specific scores might have been lost in the holistic score used in the G studies. Third, the rubric (0–3) used on the RESET observation tool might be too restrictive in the determination of a teacher’s ability to implement very specific evidence-based components within one lesson. Further research is needed to develop a deeper understanding of how to evaluate a range of special education EBPs and their components using one tool.
Conclusion
The purpose of this study was to identify how many occasions and raters are needed for acceptable levels of measurement score reliability when using the RESET observation tool to evaluate special education teachers. Generalizability and decision study analyses were completed to identify and measure sources of variance from rater data collected from two separate data coding sessions. The results from this study are not only supported by similar results found in previous research but also suggest that additional work is needed to refine and develop the optimal use of the tool. Future studies should focus on further explorations of reliability, preliminary work on validity, and measurements of improved teacher instructional practice after using feedback from the tool.
Overall, further research on the RESET observation tool, as well as rater reliability issues specific to the use of observation tools to evaluate special education teachers is needed. Although previous studies suggest that lower facet sample sizes are acceptable, the results from the D study suggest that the more the raters, the more accurate the results. Future studies should focus on minimizing the influence of other sources of variance to maximize the accuracy of results with the least amount of raters. Nevertheless, the use of an observation tool that strives to ensure teacher quality and promote professional development through the evaluation of a teacher’s use of EBPs holds promise as a meaningful way to evaluate performance.
Also, additional research and refinement on the RESET observation tool has the potential to serve states and districts faced with limited resources. Because the RESET observation tool utilizes video observations to evaluate teachers, highly trained raters can be remotely located while observing a teacher’s instructional practice. Trained raters will also be able to serve a population of special education teachers scattered across districts, instead of being geographically restricted. Future studies should also investigate the use of different raters that might eventually be tasked with evaluating special education teachers using the RESET observation tool, that is, principals, special education teachers with specific expertise, mentor teachers, district personnel, and university faculty, and how these different roles affect measurement score reliability.
Attending to the existing empirical gaps in teacher evaluation research to ensure the topic of special education is included will be crucial as states continue to reform systems and practices. When researchers begin to fully address the complex measurement challenges found within special education teacher evaluation, the body of teacher evaluation research will only be further developed and positively impacted.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Idaho State Department of Education (Grant 6FT84XXXX0046).
