Abstract
Teaching and learning are part of a complex interaction between teachers and students. Educational leaders cannot improve the teaching and learning process without quality measurement of effective teaching. One way to capture this complex interaction is by using structured observations. However, the extant literature on classroom observation instruments in the field of gifted education is limited. For that reason, a systematic search was undertaken to identify the observation instruments for assessing instructional practices used with gifted and talented students. In this article, eight observation instruments were identified: (a) Rating Scale of Significant Behaviors in Teachers of the Gifted, (b) Kulieke’s adaptation of the Rating Scale of Significant Behaviors in Teachers of the Gifted, (c) Teaching Observation Form (TOF; also known as Purdue Observation Form), (d) Classroom Practices Record (CPR), (e) Classroom Practices Record–Form VA (CPR-Form VA), (f) Classroom Instructional Practices Scale (CIPS), (g) Classroom Observation Scales–Revised (COS-R), and (h) Differentiated Classroom Observation Scale (DCOS). The instruments are described in terms of developmental process, purpose, and any reliability and validity evidence reported. This systematic search has shown the need for a new observation instrument that is comprehensive and closely tied to professional standards.
Teaching and learning are part of a complex interaction between teachers and students (Karnes & Bean, 2005). For continuous improvement of this interaction within a classroom, educators and administrators need to employ accurate assessments of instruction measured against clear standards, which must be aligned with effective practices. Kane, Kerr, and Pianta (2014) noted that “real improvement requires quality measurement” (p. 1). It would be inappropriate for educational leaders to support improvement in the teaching and learning process without quality measurement of effective teaching (Kane et al., 2014). One way to capture this complex interaction is by using structured observations (Cohen, Manion, & Morrison, 2013). A structured observation involves systematically investigating and gathering information about people, settings, routines, products, performances, and other measureable events and their interactions with one another (Cohen et al., 2013; Marshall & Rossman, 2010; Simpson & Tuson, 2003). The distinctive feature of such an observation is that it offers the opportunity to gather data in a naturally occurring environment (Cohen et al., 2013). Using structured observations, researchers and practitioners can collect and analyze data to determine the status of instructional practices and to identify instructional problems (Kane et al., 2014). This type of measurement tends to yield more direct and authentic data than might be possible with mediated or inferential methods.
By using structured observations, educational leaders can answer questions related to teacher quality and alignment with best practices in gifted education (Peters & Gates, 2010). Such information provides leaders, evaluators, and supervisors with information about the strengths and weaknesses of instructional practices. Such information can also highlight areas that require improvement. VanTassel-Baska (2004) explained that “structured forms, [classroom observations] provide benchmarks against which the teaching process can be assessed based on expectations derived from best practices in the field of gifted education” (p. 92).
By using an observation instrument that is tied closely to professional standards, educational leaders can be confident that data from the instruments can show the quality of teaching taking place (Peters & Gates, 2010). Within the field of gifted and talented education, the widely used Teacher Preparation Standards in Gifted and Talented Education (Kitano, 2008; National Association for Gifted Children, & Council for Exceptional Children, The Association for the Gifted [NAGC & CEC-TAG], 2006; Roberts, 2008) serves as the primary document for teacher preparation. These standards, which have been validated in the professional literature of the field, should be used not only in university programs in gifted education, but also by school and district personnel (Kitano, Montgomery, VanTassel-Baska, & Johnsen, 2008; Roberts, 2008). The revised standards all have elements that may be directly observed by use of a structured observation instrument (NAGC & CEC-TAG, 2013).
The extant literature on classroom observation instruments in the field of gifted education is limited (VanTassel-Baska, Quek, & Feng, 2006). To our knowledge, no one has compared psychometric qualities of the observation instruments available in our field. Furthermore, practitioners and researchers who are looking for observation instruments to collect information about teachers’ practices will need to do an extensive search. Our goal is to provide all necessary background information to make an informed decision when selecting an observation instrument that is relevant to teacher practices in the field of gifted education. In this article, we identified instruments by describing them in terms of developmental process, purpose, and any reliability and validity evidence reported.
Search Strategy
To identify observation instruments assessing instructional practices used with gifted and talented students, we completed a search of databases as well as specific journals. The following databases were used: Academic OneFile, Academic Search Complete, EBSCO, ProQuest, JSTOR, PsycArticles, and PsycInfo. These educational databases cover scholarly research related to all areas of education, thus providing comprehensive search results. Given that the aim of this literature review relates to observation instruments in the field of gifted education, the search also included the five top established peer-reviewed journals within the field of gifted education in the United States using their online websites: Gifted Child Today, Gifted Child Quarterly, Journal of Advanced Academics, Journal for the Education of the Gifted, and Roeper Review. In each database and journal’s website, the search terms used were as follows: observation*, scale*, and form*. The asterisks after a term indicate that all terms beginning with that root were included in the search.
First, the articles’ abstracts were reviewed and selected for full-text review if the article had an instrument that was (a) included in a review of instruments, (b) included in an empirical study of classroom practices, or (c) reviewed as an instrument by itself. This process yielded 34 articles. Then, full-text articles were reviewed and any instrument that met the following criteria were noted: (a) the data were collected through observations, (b) the data were collected about teachers, and (c) the evaluation of the instrument related to the field of gifted and talented education. In addition, as the full-text articles were being reviewed, reference tracking was used to identify additional articles, which were also full-text reviewed for any instrument meeting the three criteria. This process yielded 20 articles meeting the search criteria.
Once the instruments from the articles were identified, their names were entered in Google to identify articles, reports, or manuals discussing the development of the instrument, the technical features of the instrument, and any revisions or updates to the instrument. From this systematic review, several observation scales that were developed to assess classroom practices for use with gifted students were identified. The sources of information about the instruments came from 24 different manuscripts. In addition to the 20 articles mentioned earlier, the four additional sources were manuscripts that were not peer-reviewed but were included for a comprehensive understanding of the instruments.
Structured Observation Instruments in the Field of Gifted Education
In total, eight observation instruments were found and are reviewed in this section. Some instruments are described in much greater detail than others. This is because the level of specificity about each instrument within the identified articles varies considerably. Table 1 provides the purpose and any available technical data for each instrument, as well as the source of information.
Classroom Observation Instruments.
Rating Scale of Significant Behaviors in Teachers of the Gifted
As part of a study, the Rating Scale of Significant Behaviors in Teachers of the Gifted (Martinson & Wiener, 1968) was developed as a screening instrument for prospective teacher participants who would mentor other teaching colleagues to work with gifted students effectively. Prior to the development of the scale, the two co-principal investigators of the study independently created a list of essential strategies and behaviors of teachers working with gifted children. Through a comparison of the two lists, as well as exploration of sources such as the rating scales in the Handbook of Research on Teaching, edited by N. L. Gage, the co-directors agreed on 27 types of behavior that were included in a preliminary rating scale, Scale of Flexibility and Innovation in Teachers with Descriptive Behaviors of Students and Teachers.
Ten experts in writing, program development, and research in the area of gifted education collaborated to develop the final instrument (Martinson & Wiener, 1968). As part of the preliminary process, organized by the co-directors of this study, the experts attended an orientation on the use of the preliminary scale, used the scale to judge teachers, and reviewed the scale item by item. On the basis of their suggestions and item analysis, the 10 experts developed the Rating Scale of Significant Behaviors in Teachers of the Gifted. The scale was developed to rate 37 behavioral statements as seldom, occasionally, or frequently occurring after a minimum of a 40-minute observation of teaching and examining portfolio materials, files of children’s products, and other evidence to confirm their impressions.
The scale was then subjected to a statistical analysis after the experts independently observed, in groups of two, 75 teachers for 40 to 45 minutes (Martinson & Wiener, 1968). Using the Kuder-Richardson formula, a reliability coefficient of .95 was obtained among the experts. In addition, four distinct factors emerged during the factor analysis. The first factor, labeled as Individualized Materials and Instruction, included 9 items with the lowest loading being .32 and highest .80. The second factor, labeled as The Art of Questioning, included 6 items with the lowest loading being .57 and highest .72. The third factor, labeled as The Encouragement of Higher Level Learning, included 14 items with the lowest loading being .30 and highest .82. The fourth factor, labeled as Communication-Interaction, included 8 items with the lowest loading being .41 and highest .70 (Martinson & Wiener, 1968).
Kulieke (1986) offered an adaptation of the Rating Scale of Significant Behaviors in Teachers of the Gifted (Martinson & Wiener, 1968). According to Kulieke (1986), the adaptation was scaled to make more consistent comparisons between each aspect of the classroom being observed. The scale included 22 behavioral statements divided into 7 categories: conducts group discussions, selects questions that stimulate higher-level thinking, uses varied teaching strategies effectively, utilizes critical thinking skills in appropriate contexts, encourages independent thinking and open inquiry, understands and encourages student ideas and student-directed work, and demonstrates understanding of the educational implications of giftedness. The scale should be used after two 30-minute observation periods and each behavior statement is rated on 5-point Likert-type scale (i.e., Excellent [5], Good [4], Fair [3], Poor [2], Very poor [1]; Kulieke, 1986). However, no reliability or validity data were reported for this scale.
Teaching Observation Form (Purdue Observation Form)
Another scale, the Teaching Observation Form (TOF; also known as the Purdue Observation Form), was originally developed to evaluate teachers working in a Saturday enrichment program specifically designed for gifted and talented students at Purdue University (Feldhusen & Hansen, 1987). Later, the instrument was used to evaluate the teaching of preservice teachers as part of a three-credit practicum in teaching the gifted offered by Purdue University (Feldhusen & Huffman, 1988). The development of the instrument started with a review of the literature to determine the characteristics and competencies that are considered essential for teachers of gifted education (Feldhusen & Hansen, 1988). Accordingly, the instrument was developed to rate a teacher’s performance on a 5-point scale (i.e., Outstanding [5], High [4], Average [3], Needs Improvement [2], Not Satisfactory [1]) and included a “not observed” option in 10 categories: (a) subject matter coverage, (b) clarity of teaching, (c) motivational techniques, (d) pace of instruction, (e) opportunity for self-determination of activities by students, (f) student involvement in a variety of experiences, (g) interaction between teacher and student, student and peers appropriate to course objective, (h) opportunity for student follow-through of activities outside class (homework), (i) emphasis on higher level thinking skills, and (j) use of teaching and learning aids (Feldhusen & Huffman, 1988). Each category included at least three criteria expected to be observed, and accordingly the rating reflects the number and quality of criteria observed (Feldhusen & Huffman, 1988). In addition, the instrument included a comment section on the method, activities, materials, or skills. The observation time required for this instrument is one full-class period, which is usually 45 to 60 minutes (Feldhusen & Huffman, 1988).
A year later, 10 gifted education content experts evaluated the items and added two categories: (a) emphasis on creativity and (b) lesson plans designed to meet program, course, and daily objectives. Also, some statistical analyses were completed as part of a doctoral dissertation (Hansen, 1988). Hansen (1988) found an overall interrater reliability of .87 when rating a 15-minute video segment, an overall alpha reliability coefficient of .86, and different item-total correlations ranging from .64 to .84.
In 2007, the TOF was revised (Peters & Gates, 2010). First, the 12 categories were evaluated by 14 experts who were directors and faculty at universities with gifted education programs, as well as coordinators and veteran teachers of gifted and talented programs. These experts rated each item on a 5-point Likert-type scale for importance and clarity of language, as well as any additional comments to improve the instrument (Peters & Gates, 2010). In addition, scholarly literature was reviewed for current best practices in gifted education pedagogy. Accordingly, the 12 categories and the criteria within each were revised to update language, reflect the current literature, and increase clarity (Peters & Gates, 2010). The response scale was also changed from a 5-point to a 7-point scale. The current 12 categories are the following: (a) content coverage; (b) clarity of teaching; (c) motivational techniques; (d) pedagogy/instructional techniques; (e) opportunity for self-determination of activities by students; (f) student involvement in a variety of experiences; (g) interaction between teacher and student, student and peers; (h) opportunity for student follow-up on activities and topics on their own; (i) emphasis on higher-level thinking skills; (j) emphasis on creativity; (k) lesson plans designed to meet program, course, and daily objectives; and (l) appropriate use of classroom technology (Peters & Gates, 2010).
Once the revisions were completed, the instrument was used to observe and evaluate instructors in one university’s enrichment program courses (Peters & Gates, 2010). Data collection between 2007 and 2009 was completed by eight observers, yielding 217 separate observations; the data were used for a statistical analysis of the instrument. The new TOF achieved an overall alpha reliability estimate of .95, different item-total correlations ranging from .54 to .84, and an intraclass correlation coefficient value for the observer-level effect of .44 (Peters & Gates, 2010).
Classroom Practices Record
The three scales discussed above focus primarily on teacher behaviors and skills. The scales do not provide detailed information regarding classroom environment, specific differentiated strategies, or student engagement. The Classroom Practices Record (CPR) was developed to document the extent to which gifted and talented students receive differentiated instruction through modifications in curricular activities and materials, as well as the verbal interaction between the teacher and students (Westberg et al., 1990). The instrument was developed by adapting two instruments used to document teacher practices in the general education classroom: (a) Classroom Observation Instrument (Giesen & Sirotnik, 1979) and (b) Classroom Activity Record (Evertson & Burry, 1989).
The CPR was not designed to document information about the instructional practices of the teacher with the entire classroom, but to provide descriptive information on the instructional practices with two students: Target Student #1 is a student who has been identified as a gifted and talented student or a student with high ability, and Target Student #2 is an average ability student (Westberg et al., 1990). The instrument contains six sections:
Identification Information: Observer records information on the school, the teacher, and the target students observed.
Physical Environment Inventory: Observer records information on the availability and types of learning/interest centers, the seating pattern, and the location of the two target students in the classroom.
Curricular Activities: Observer records information about the academic subject, the instructional activity, the grouping practices, and the differentiation experiences by Target Student #1. The information on instructional activities includes reporting the beginning and ending time of the activities and selecting among 14 types of curricular activities. The information on grouping practices includes selecting among four types of group size and two types of group composition. The information on differentiation includes only the curricular experiences of Target Student #1 that are different from those experienced by Target Student #2. The observer selects among five types of differentiation or writes a description of the differentiation experience.
Verbal Interactions: Observer records information on all verbal interactions that occur between the teacher and target students. Codes are used to record who is involved in the verbal interaction (i.e., teaching adult, Target Student #1, Target Student #2, students-at-large), the type of interaction (i.e., knowledge/comprehension question, higher order thinking skills question, request or command, explanation or statement, response, no verbal response), and the existence of 3 or more seconds of wait time associated with questions.
Teacher Interview Record: Observer conducts a semi-structured interview with the teacher at the end of the school day to clarify or elaborate on information recorded in the curricular activities section.
Daily Summary: Observer writes a summary report that describes the instructional situation and summarizes the observed differentiated experiences observed in the classroom during the day (Westberg et al., 1990; Westberg et al., 1992).
According to the CPR manual, the instrument was piloted four separate times for the purpose of making revisions in the instrument, the training manual, and the observation procedures (Westberg et al., 1990). Following Frick and Semmel’s (1978) recommendation, the developers of the instrument used a criterion-related procedure, which refers to consistency between ratings of an observer and a criterion rating, to establish the reliability of the CPR (Westberg et al., 1992). The standard criterion recordings were grouped into four events and calculated with a total of 15 observers after being provided with training through the manual (Westberg et al., 1992). All observers demonstrated at least 80% criterion-related agreement on the four event categories and the total training exercise (Westberg et al., 1992).
Classroom Practice Record–Form VA
In a 3-year study aimed at studying how pre-service teachers develop awareness of the needs of academically diverse students and implement modified instruction to meet those needs, Tomlinson et al. (1995) used a modified version of the CPR; their purpose was to describe the environment and instructional activities in the project’s classrooms. This new version named CPR-Form VA was designed to collect information about the classroom student composition and the type of instructional activities taking place. This instrument contains only the first three sections of CPR:
Identification Information
Physical Environment Inventory
Curricular Activities
In addition, rather than merely providing descriptive information on the instructional practices with two students (i.e., Target Student #1 is a student who has been identified as a gifted and talented student or a student with high ability, and Target Student #2 is an average ability student), which is in the case of the CPR, CPR-Form VA was designed to provide descriptive information on the instructional practices with various groups of targeted students. In fact, the observer notes, within the Identification Information section, the number of students in each of five categories (i.e., Limited English Proficient, handicapping condition(s), economically disadvantaged, student(s) accelerated one grade, and student(s) accelerated more than one grade) and how these students were identified (e.g., achievement tests, group IQ tests, individual IQ tests, or teacher rating scales). Another modification is that one type of curricular activity, labs, was added to the list within Curricular Activities. Reliability and validity data were not reported for this scale.
Classroom Instructional Practices Scale
The Classroom Instructional Practices Scale (CIPS; Johnsen, 1992) is another instrument designed to determine the degree to which the classroom teacher is adapting for learner differences in the general education setting. More specifically, the instrument measures how teachers organize the curriculum and manage the classroom in order to adapt for individual differences in content, rate, preference, and environment (Johnsen et al., 2002). The description of each area is hierarchical, beginning with the least adaptive classroom practice for individual differences and progressing to the most adaptive classroom practice, allowing for the measurement of teacher growth in the identified areas of differentiation (Johnsen, 1992). After observing approximately 200 teachers who were involved in changing their classroom practices, Johnsen et al. (2002) developed descriptors and a hierarchy of steps to adapt instructional practices for gifted learners. The four major areas for adapting for individual differences can be defined as follows:
Content: Describes the way the teacher organizes and sequences skills, concepts, strategies, and generalizations within and across disciplines. For example, the lowest rating (C1) describes content that is organized around the book’s scope and sequence, while C7 describes content organized around individual student interest.
Rate: Describes how the teacher uses assessment to vary the amount of time needed by the students in learning new content. For example, a teacher who receives an R1 rating provides the same amount of time for every student in the classroom, while a teacher receiving an R9 rating uses a preassessment to identify a student who needs or may choose in-depth study, enrichment, or acceleration.
Environment: Describes the way the teacher arranges the physical environment to facilitate interaction and learning among students. For example, the lowest rating (E1) describes a classroom in which the teacher limits interaction between students and with learning materials. An E6 rating describes a classroom where students learn from one another and use the community and the school as learning centers.
Preference: Describes how the teacher aligns activities with the content and provides for individual student choice. For example, at the lowest rating (P1), the student has no choice of learning materials and uses materials that have a similar format such as paper-pencil. At P5, the student may select or create learning activities. At the highest level, these activities also vary the task (e.g., visual, auditory, kinesthetic) and the response (e.g., written, oral, physical; Johnsen et al., 2002).
To triangulate the data collected from the classroom observations and confirm the differentiation ratings, the teacher and two students are interviewed regarding typical practices related to content, rate, preference, and environmental adaptations. Reliability for the instrument was established among research assistants through training sessions that produced an interrater reliability of .92 (Ryser & Johnsen, 1996).
Classroom Observation Scales–Revised
Another observation instrument that documents differentiated instruction with gifted and talented students is the William & Mary Classroom Observation Scales–Revised (COS-R). While the CIPS focuses on ways teachers organize their curriculum, use assessments, vary activities, and arrange the learning environment in order to adapt for individual differences, the COS-R focuses on differentiation in terms of classroom-based instructional behavior (VanTassel-Baska et al., 2003). The development of this instrument went through several stages that occurred over more than a decade (VanTassel-Baska et al., 2006). The original form was developed as part of a study by graduate students and Center for Gifted Education staff to evaluate the level of differentiation teachers employed in the general education setting (VanTassel-Baska et al., 1997). The instrument was called the Classroom Observation Form (COF) and included a total of 40 items describing observable behaviors with a dichotomous rating of yes or no. The items were developed by reviewing the literature about effective teaching practices with particular attention to gifted learners (VanTassel-Baska et al., 1997). The items represented seven clusters of observable content and instructional strategies and two categories of elements of educational reform (VanTassel-Baska et al., 1997).
The COF was piloted with 50 teachers working in an enrichment program at William & Mary (VanTassel-Baska et al., 2005, 2006). Based on the data gathered, the number of clusters and items was reduced, the rating scale was modified, and the instrument was renamed the Classroom Observation Scales–Revised (VanTassel-Baska et al., 2005, 2006). The COS-R includes the following clusters of teaching behaviors:
Curriculum Planning and Delivery
Accommodations for Individual Differences
Problem Solving
Critical Thinking Strategies
Creative Thinking Strategies
Research Strategies
In addition, to complement the six clusters of teaching behaviors, sets of domain-specific indicators were developed, including illustrative examples pertaining to the different subject areas (i.e., math, science, literature, social studies, and foreign/second language; VanTassel-Baska et al., 2005).
According to the manual, the COS-R should be completed based on a 30- to 50-minute lesson (VanTassel-Baska et al., 2005). To reduce subjectivity, two observers should conduct the observation in each classroom and complete the COS-R. First, the observer determines whether the behavior is observed or not. If the behavior is observed, then the observer should rate its effectiveness on a 3-point scale (i.e., Effective [3], Somewhat Effective [2], Ineffective [1]; VanTassel-Baska et al., 2005, 2006).
Six experts in the field of gifted education, including professors, scholars, practitioners, and administrators, were identified to review the instrument for content validity (VanTassel-Baska et al., 2005). However, only four experts returned their ratings. The experts were asked to rate two dimensions on a 3-point scale: (a) the importance of each behavioral item (Dimension 1; Very Important [3], Important [2], Not Important [1]), and (b) the accuracy of the language describing the behavior (Dimension 2; Very Clear [3], Clear [2], Lack of Clarity [1]). An intraclass coefficient alpha of .86 was found for rater agreement on the instrument on the first dimension, .99 on the second dimension, and .98 for the entire COS-R content validity (VanTassel-Baska et al., 2005).
The COS-R was also subjected to a statistical analysis during two waves of data collection from a research study funded by a United States Department of Education Javits grant (VanTassel-Baska et al., 2005, 2006). Twenty-three teams of observers collected 72 observations during the first wave and 62 observations during the second wave (VanTassel-Baska et al., 2005). According to the item reliability analysis, overall the instrument has an alpha of .91 from the first wave of data collection, and .93 from the second (VanTassel-Baska et al., 2005, 2006). The interrater reliability was .87 for the first observation period and .89 for the second (VanTassel-Baska et al., 2005, 2006). The reliability of the COS-R was replicated a few years later, yielding similar results (VanTassel-Baska et al., 2005).
Differentiated Classroom Observation Scale
Another instrument evaluating differentiated instruction is the Differentiated Classroom Observation Scale (DCOS; Cassady et al., 2004). Although the instrument was initially developed to examine the strategies employed to meet the needs of gifted children receiving instruction in cluster grouped classrooms, the instrument could be used to evaluate any classroom setting that includes differentiated instruction to meet the needs of groups of students (Cassady et al., 2004). The instrument was developed to include three main components used at different periods:
Pre-observation interview: This is the initial contact with the teacher being observed to gather information before the actual observation period. According to Cassady et al. (2004), this information includes “contextual factors that may be atypical to the standard classroom experience” (p. 140). The teacher provides the observer with a copy of the lesson plan, and they agree on an arrangement to provide differentiated nametags or identifiers so the observer can distinguish among the different groups of students. In addition, teachers are asked questions that will help classify the lesson to be reviewed. However, even though the instrument already includes specific questions, which were developed in accordance with the author’s project focusing on differentiation and professional development training, Cassady et al. recommend modifying them according to the purpose of the data collection.
Observation period: In this period, the observer collects data on the classroom environment and learning experiences for the different groups of students during the lesson. First, the observer notes the visible features of the classroom. Then, using 5-minute segments, the observer documents the instructional activities, level of student engagement, levels of represented cognitive activity, and the learning director for identified and non-identified groups. Finally, at the conclusion of the observation, the observer rates 12 items related to the methods employed for grouping, the differentiated instruction, the classroom environment, and the interaction between students and the teacher, using a 5-point Likert-type scale (i.e., Strongly Disagree, Disagree, Neutral, Agree, Strongly Agree). These judgments are made separately for each group of students.
Postobservation debriefing and reflection: Immediately following the observation period, the observer provides the teacher with a moment to debrief and provide any final comments. This gives the teacher the opportunity to share information that may be helpful to understanding the observation. Following this step, outside the confines of the classroom, the observer records any additional comments or statements to provide his or her impression that may not have been represented in the standard protocol (Cassady et al., 2004).
Through live and video-based training, the authors found two primary challenges: (a) it takes time for an individual to be able to use the instrument as intended, and (b) some observers found it challenging to monitor all the factors targeted in this observation (Cassady et al., 2004). However, the authors explained that they were successful in training experienced gifted education teachers and graduate students in using the DCOS (Cassady et al., 2004).
The DCOS includes a unique material that has not been used with any of the other observation instruments: a data management system (Cassady et al., 2004). Along with the protocol that includes the instructions and the Scoring Form, a DCOS Database was developed by the authors to track the three different phases of the observation while maintaining individual observation data (Cassady et al., 2004). According to Cassady et al. (2004),
The advantages of this database application include: (a) simplified data entry; (b) sorting capabilities on any identified variable; (c) disaggregation of data, allowing the examination of within-segment effects and trends; and (d) ability to move the data entry process to the web, enabling multi-site data sharing. (p. 142)
The DCOS includes rigorous methodologies in terms of systematic documentation and evaluation of data. However, no statistical analysis has been reported for this instrument.
Discussion
In spite of decades of advocacy and efforts to meet the needs of gifted and talented students, only eight classroom observation instruments have been published since the late 1960s within the field of gifted education. This is of concern, since the knowledge base as it is related to gifted education has grown significantly with research findings that provide the field with a better understanding of meeting the needs of gifted students (Plucker & Callahan, 2014). Although new research-based instructional practices are being recommended, the instruments are not being updated. Additionally, there has been considerable evidence collected over the years suggesting that teachers’ instructional practices and behaviors influence students’ learning (Hunsaker, Nielsen, & Barlett, 2010; Sanders & Rivers, 1996; Wenglinsky, 2000).
Among these few available instruments, limited evidence regarding their technical adequacy is available. As stated at the beginning of this article, “real improvement requires quality measurement” (Kane et al., 2014, p. 1). From a psychometric perspective, classroom observation instruments should be designed, tested, and used with careful attention to the methodological quality. According to the American Psychological Association (APA, 2014), three criteria should be considered when using an observation instrument. The following is a comparison of the eight instruments using the criteria proposed by APA.
Is the observation instrument well standardized in terms of its administration procedures? Does it offer clear directions for conducting observations and assigning scores?
Four of the instruments include some type of assistance for the users: (a) Rating Scale of Significant Behaviors in Teachers of the Gifted, (b) CPR, (c) COS-R, and (d) DCOS.
An orientation on the use of the Rating Scale of Significant Behaviors in Teachers of the Gifted was available. The CPR has a manual which provides the reader with background information, instructions related to observation arrangements and procedures, a description of each section within the instrument, and training exercises to be completed before using the instrument. The COS-R also includes a user’s manual which provides the reader with background information, description of the scale development, instructions related to implementation and scoring, information about the technical adequacy of the instrument, and a list of references used during the development of the instrument. The DCOS requires either live or video-based training before being used. In addition, three documents are provided to assist in the data collection and the synthesis process: (a) DCOS protocol, (b) DCOS scoring form, and (c) a data management system.
2. Does the observation instrument include interrater agreement (reliability)?
Evidence of interrater agreement among raters using the instrument were reported for five of the instruments: (a) Rating Scale of Significant Behaviors in Teachers of the Gifted, (b) TOF, (c) CPR, (d) CIPS, and (e) COS-R.
Interrater agreement was calculated for the Rating Scale of Significant Behaviors in Teachers of the Gifted, using the Kuder-Richardson formula. A reliability coefficient of .95 was obtained between 10 experts.
For the TOF, interrater agreement was calculated during the development of the instrument and after it was revised. The establishment of the first interrater agreement was part of a doctoral dissertation (Hansen, 1988). Hansen found an overall interrater reliability of .88 when rating a 15-minute video segment, an overall alpha reliability coefficient of .86, and different item-total correlations ranging from .64 to .84. Once the revisions were completed, the instrument was used to observe and evaluate instructors in one university’s enrichment program courses (Peters & Gates, 2010). Data collection was conducted between 2007 and 2009 by eight observers, yielding 217 separate observations: the data were used for a statistical analysis of the instrument. The new TOF achieved an overall alpha reliability estimate of .95, different item-total correlations ranging from .54 to .84, and an intraclass correlation coefficient value for the observer-level effect of .44 (Peters & Gates, 2010).
For the CPR, a criterion-related procedure was used to establish the reliability. The standard criterion recordings were grouped into four events and calculated with a total of 15 observers who were provided with training through the manual (Westberg et al., 1992). All observers demonstrated at least 80% criterion-related agreement on the four event categories and the total training exercise (Westberg et al., 1992).
Reliability for the CIPS was established among research assistants through training sessions that produced an interrater reliability of .92 (Ryser & Johnsen, 1996).
The interrater agreement for COS-R was done during two waves of data collection from a research study funded by a United States Department of Education Javits grant. Twenty-three teams of observers collected 72 observations during the first wave and 62 observations during the second wave (VanTassel-Baska et al., 2005). The interrater reliability was .87 for the first observation period and .89 for the second (VanTassel-Baska et al., 2005, 2006). The reliability analysis of the COS-R was replicated few years later, yielding similar results (VanTassel-Baska et al., 2005).
3. Is there at least some preliminary research demonstrating the validity of the instrument that is more than just reliability information?
Validity was only demonstrated for four of the eight instruments: (a) Rating Scale of Significant Behaviors in Teachers of the Gifted, (b) TOF, (c) CIPS, and (d) COS-R.
For the Rating Scale of Significant Behaviors in Teachers of the Gifted, validity was addressed during several phases of development of the instrument. First, content validity was judged by ten experts in writing, program development, and research in the area of gifted education. In fact, as part of their preliminary process, the experts developed and used the Scale of Flexibility and Innovation in Teachers with Descriptive Behaviors of Students and Teachers. They attended an orientation on the use of the scale, used the scale to judge teachers, and reviewed the scale item by item. On the basis of their suggestions and item analysis, the 10 experts developed the Rating Scale of Significant Behaviors in Teachers of the Gifted.
In addition, construct validity was investigated for the Rating Scale of Significant Behaviors in Teachers of the Gifted through factor analysis. The results showed four distinct factors: (a) Individualized Materials and Instruction, included 9 items with the lowest loading being .32 and highest .80; (b) The Art of Questioning, included 6 items with the lowest loading being .57 and highest .72; (c) The Encouragement of Higher Level Learning, included 14 items with the lowest loading being .30 and highest .82; (d) Communication-Interaction, included 8 items with the lowest loading being .41 and highest .70 (Martinson & Wiener, 1968).
During the development of both the TOF and CIPS, content validity was addressed. In the process of designing the TOF, the researchers conducted a review of the literature to determine the characteristics and competencies that are considered essential for teachers of gifted education (Feldhusen & Hansen, 1988). A year later, 10 gifted education content experts evaluated the items and added two categories. In 2007, the 12 categories in the TOF were evaluated by 14 experts who were directors and faculty at universities with gifted education programs, as well as coordinators and veteran teachers of gifted and talented programs (Peters & Gates, 2010). These experts rated each item on a 5-point Likert-type scale for importance and clarity of language, as well as any additional comments to improve the instrument (Peters & Gates, 2010). In addition, scholarly literature was reviewed for current best practices in gifted education pedagogy (Peters & Gates, 2010). Accordingly, the 12 categories and the criteria within each were revised to update language, reflect the current literature, and increase clarity (Peters & Gates, 2010).
For the CIPS, the descriptors and hierarchy of steps were developed after observing approximately 200 teachers who were involved in changing their classroom practices to adapt for gifted learners (Johnsen et al., 2002).
The development of the COS-R went through several stages that occurred over more than a decade during which validity was demonstrated. At first, content validity was addressed when the items were developed by reviewing the literature about effective teaching practices with particular attention to gifted learners (VanTassel-Baska et al., 1997). Then, six experts in the field of gifted education, including professors, scholars, practitioners, and administrators, were identified to review the instrument for content validity (VanTassel-Baska et al., 2005). The experts were asked to rate two dimensions on a 3-point scale. An intraclass coefficient alpha of .86 was found for rater agreement on the instrument on the first dimension, .99 on the second dimension, and .98 for the entire COS-R content validity (VanTassel-Baska et al., 2005).
For the teaching and learning process to improve within a classroom, in addition to the APA criteria used for determining the quality of an observation instrument, accurate assessment of instruction measured against clear standards for what is known to be effective should be used (APA, 2014). The instruments identified were developed for different purposes, with the majority being as part of a research study. Therefore, the available instruments measure different aspects of teachers’ instructional practices with gifted students.
The Rating Scale of Significant Behaviors in Teachers of the Gifted and the TOF focus primarily on teacher behavior and skills but do not provide detailed information regarding classroom environment, specific differentiated strategies, or student engagement. The CIPS, COS-R, CPR, and DCOS document the extent to which gifted and talented students receive differentiated instruction through modifications in curricular activities and materials. Although the CIPS focuses on ways teachers organize their curriculum, use assessments, vary activities, and arrange the learning environment in order to adapt for individual difference (Johnsen, 1992), the COS-R focuses on differentiation in terms of classroom-based instructional behaviors (VanTassel-Baska et al., 2003). The CIPS and COS-R have been used together in a dissertation study to obtain a more comprehensive evaluation of teachers (Ochoa, 2013).
On the other hand, the CPR and DCOS provide comparative information with regard to different instructional practices with different groups of students. Although the CPR focuses on comparing instructional practices provided to only two students (Westberg et al., 1992), the DCOS captures the instructional practice with the entire classroom. However, the DCOS was initially developed to examine the strategies employed to meet the needs of gifted children receiving instruction in a cluster-grouped classroom. The authors suggest that, with minimal changes, the instrument could be used to evaluate any classroom setting that involves differentiated instruction to meet the needs of groups of students (Cassady et al., 2004).
Need for the Development of a New Observation Instrument in Gifted Education
Hertberg-Davis (2009) stated that
recommendations for differentiating learning experiences for gifted students include principles of providing not only challenges generally considered beneficial for gifted students (e.g., greater depth and complexity, adjusted pace, greater independence) but also curricular and instructional modifications geared toward individual student need. (p. 251)
To operationalize such a recommendation and provide formative assessment for those implementing differentiation, building administrators and others who observe gifted education teachers need a tool for establishing baseline information and making suggestions for improvement. In our systematic search, we identified eight observation instruments for assessing instructional practices used with gifted and talented students. Our findings suggest that a new instrument could meet some needs that are not currently being addressed for those who work with preservice and in-service teachers in gifted education.
First, a few of the instruments reviewed address some of the criteria provided by APA. However, none are comprehensive in this respect, nor have they been updated to address those criteria. Deliberate and systematic use of the APA criteria when designing a new classroom observation instrument will ensure greater methodological quality.
Second, these instruments were developed over a wide timespan, from the late 1960s to 2010. As such, there is a concern about whether the instruments reflect expectations about contemporary pedagogy as it relates to working with gifted students. The influence of content standards also creates a significant influence on the requirements in any given discipline. This thus calls for an examination of how those curriculum elements and the associated pedagogical frameworks should be framed when working with gifted students.
Third, within the field of gifted and talented education, collaboration between NAGC and CEC-TAG has produced specific gifted education standards that provide guidelines for teacher preparation and effective teaching practices with gifted students (Johnsen, 2011). The observation instruments described in this article were created prior to the development of these standards. Thus, none are aligned with the standards that have been adopted by the field as best instructional practices for working with gifted students. Accordingly, we encourage scholars and practitioners to develop new observation instruments, designed to reflect the professional standards and tested with careful attention to methodological quality. This is a task we are pursuing as a result of this review of the literature, and we encourage others in gifted education to do the same.
We conclude with a cautionary note. Because a single instrument will not be viable in all situations, scholars and practitioners conducting teacher observations in varying contexts must develop instruments specific to these settings. Even the most psychometrically sound instruments will not be effective if the instrument is not useful for helping observers evaluate teachers for continuous improvement. Therefore, an emphasis on continuous improvement depends on both the use of valid and reliable instruments and on the development of close relationships with school personnel to ensure that instruments help answer useful questions.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
