Abstract
This study investigated teacher and parent ratings of teacher-nominated gifted elementary school students’ verbal abilities, mathematical abilities, deductive reasoning, creative thinking, and engagement, and connected these ratings to school grades. Teacher and parent ratings were compared with regard to accuracy levels and halo effects. Furthermore, this study explored the correlations between teacher and parent ratings and how they are related to school grades. The study was based on data from 572 elementary school students participating in an enrichment program. The results indicated the same accuracy levels for teachers and parents. However, teacher ratings were more strongly affected by halo effects than parent ratings. The correlations between teacher and parent ratings were small to medium. Both raters’ ratings were independently and positively associated with German grades when controlling for each other. Positive teacher or parent ratings of mathematical abilities and engagement buffered the relation between the other rater’s ratings and math grades.
Keywords
Both teachers and parents are important to students. For example, their agreement or disagreement on learning-related values are relevant for students’ school performance (for reviews, see Christenson, 1999; Glueck & Reschly, 2014). Research also indicates that considering teachers’ and parents’ ratings of students’ abilities jointly provides a better understanding of students’ school achievement (Benner & Mistry, 2007; Peet, Powell, & O’Donnel, 1997). One of the most prominent questions in gifted education is how gifted students can transform their potential into high achievement (see Shavinina, 2009). Hence, teachers’ and parents’ views of gifted children and how they relate to each other (as well as to school grades) should receive attention in the analysis of factors associated with this transformation.
Although a number of studies have compared teacher and parent ratings of students’ cognitive abilities, creativity, and motivation (e.g., Geiser, Mandelman, Tan, & Grigorenko, 2016; Sommer, Fink, & Neubauer, 2008), all of which are seen as central facets of giftedness (see Sternberg & Davidson, 2005), few have concentrated on gifted students in particular (see Chan, 2000, for an exception). Furthermore, empirical studies of the relation between teacher and parent ratings on the one hand and (gifted) students’ school achievement on the other are sparse (Glueck & Reschly, 2014).
The contribution of the present study is thus twofold: first, it advances the literature on ratings of gifted students by comparing teachers’ and parents’ ratings of German elementary school students who were nominated for a statewide enrichment program for gifted students by their teachers. We considered the following facets of giftedness: cognitive abilities (i.e., verbal abilities, mathematical abilities, and deductive reasoning), creative thinking, and engagement. Second, the study explores how teacher and parent ratings are jointly related to students’ school grades in German and mathematics.
Gifted Students
Modern scientific conceptions of giftedness (e.g., Heller, Perleth, & Lim, 2005; Subotnik, Olszewski-Kubilius, & Worrell, 2011, 2012) consider high general or domain-specific cognitive abilities to be the main aspects of giftedness, but also include high creative, motivational, and beneficial environmental characteristics. Similarly, teachers’ and parents’ beliefs about giftedness typically comprise the notion that gifted students have high cognitive abilities, are creative, and are motivated to learn (Buckley, 1994; Endepohls-Ulpe & Ruf, 2005; Schack & Starko, 1990).
A variety of different methods such as tests, teacher nominations, or parent ratings as well as different criteria such as high intelligence or high achievement are used in different combinations to identify gifted students (e.g., Carman, 2013; National Association for Gifted Children, 2013). As a consequence, different groups of students can be considered gifted in different contexts. For example, a recent meta-analysis (Acar, Sen, & Cayirdag, 2016) distinguished nonperformance methods such as nominations and ratings by teachers and parents from performance methods such as tests of academic achievement, cognitive ability, and creativity. The nonperformance methods were only moderately correlated with the performance methods (r = .30). Thus, teachers and parents see a somewhat different group of students as gifted than one would expect on the basis of ability and achievement tests.
Comparing Teacher and Parent Ratings of Facets of Students’ Giftedness
Teachers and parents have different perspectives on children (Harder, Trottler, Vialle, & Ziegler, 2015; Petscher & Li, 2008): typically, parents have known their children for longer and know them better than teachers, whereas teachers typically have more professional knowledge in the areas of education and academic development. Furthermore, teachers usually see children in a larger number and wider variety of academically relevant situations than parents, and can better compare a given child with a larger group of other children.
Teachers’ and parents’ ratings of student characteristics that are seen as facets of giftedness (e.g., cognitive abilities, creativity, and motivation) tend to be investigated separately (e.g., Baudson, Fischbach, & Preckel, 2014; Genser, Strasser, & Garbe, 1981; Gralewski & Karwowski, 2013; Rennen-Allhoff, 1991; Skinner, Kindermann, & Furrer, 2009). A few studies have examined teacher and parent ratings together (e.g., Baudson & Preckel, 2013; Geiser et al., 2016; Miller & Davis, 1992; Peet et al., 1997; Petscher & Li, 2008; Sommer et al., 2008) but did not concentrate—with a few exceptions (e.g., Chan, 2000)—on gifted students. In the following sections, we will summarize the main results of these studies with regard to the accuracy of teacher and parent ratings, halo effects in teacher and parent ratings, and the correlation between teachers’ and parents’ ratings.
Comparing the Accuracy of Teacher and Parent Ratings
Teachers and parents are often asked to rate student characteristics (e.g., mathematical abilities or engagement) on a Likert-type scale. The accuracy of these ratings is typically determined by comparing them to students’ test scores or self-reports (Machts, Kaiser, Schmidt, & Möller, 2016). Teachers and parents seem to be more accurate in rating students’ cognitive abilities than in rating students’ creativity (Machts et al., 2016; Schrader, 2010; Sommer et al., 2008; Urhahne, 2011). Teacher and parent ratings of cognitive ability have been found to have similar average associations with test scores of general cognitive ability (rteachers = .56, rparents = .50; Sommer et al., 2008), but Geiser et al. (2016) reported higher convergent validity concerning analytical abilities for teachers than parents. The correlations between ratings and test scores for students’ verbal (rteachers = .57, rparents = .19), mathematical (rteachers = .70, rparents = .43), and figural abilities (rteachers = .53, rparents = .35) were higher for teachers than parents in a study by Miller and Davis (1992). However, comparing teachers’ and parents’ difference scores (i.e., the difference between a rating and a student’s test score) led the authors to conclude that teachers’ and parents’ judgments were equally accurate. The association between creativity ratings and student test scores was higher for teachers than parents (rteachers = .34, rparents = .24) in Sommer et al.’s (2008) study, and Geiser et al. (2016) again reported a higher convergent validity for teacher ratings than parent ratings. Concerning students’ motivation, correlations between ratings and students’ self-reports have been reported to be small 1 to medium for teachers (rteachers ≤ .30, as summarized by Spinath, 2005) and small or large for parents (rparents = .25, Helmke & Schrader, 1989; r parents = .69, Genser et al., 1981). To the best of our knowledge, no studies directly comparing the accuracy of the two types of ratings of students’ motivation have been conducted.
Comparing Halo Effects in Teacher and Parent Ratings
The correlations between different characteristics rated by teachers or parents have been frequently found to be higher than the correlations between corresponding student data such as tests or self-reports (e.g., Li, Lee, Pfeiffer, & Petscher, 2008; Sommer et al., 2008). For example, in a study by Urhahne (2011), correlations between teachers’ ratings of students’ mathematical abilities, creativity, and task commitment ranged from r = .54 (mathematical abilities and creativity) to r = .67 (mathematical abilities and task commitment). However, correlations between students’ test scores for mathematical abilities and creativity and self-reports for task commitment ranged from r = −.07 (mathematical abilities and task commitment) to r = .25 (mathematical abilities and creativity). This phenomenon has been termed the halo effect (e.g., Babad, Bernieri, & Rosenthal, 1989; Babad, Inbar, & Rosenthal, 1982). Fisicaro and Lance (1990) offered three explanations for halo effects: first, a broad general impression of the person; second, one highly salient factor influencing all ratings; or third, raters might be unable to conceptually discriminate between the different dimensions. On a descriptive level, the correlations between characteristics were higher for teacher ratings than for parent ratings (Chan, 2000; Petscher & Li, 2008; Pfeiffer, Petscher, & Kumtepe, 2008), indicating greater halo effects among teachers. To the best of our knowledge, however, differences in the strengths of the correlations have not yet been tested for statistical significance.
Correlations Between Teachers’ and Parents’ Ratings of the Same Characteristic
Studies directly comparing teacher and parent ratings of the same characteristic have reported medium to large correlations for students’ cognitive abilities (Geiser et al., 2016; Miller & Davis, 1992; Sommer et al., 2008; Spinath & Spinath, 2005). Furthermore, small correlations between teacher and parent ratings of students’ creativity (Geiser et al., 2016; Runco, 1989; Sommer et al., 2008) and medium correlations for students’ motivational characteristics (Chan, 2000; Peet et al., 1997) have been found.
Overall, teachers appear to be similarly or more accurate than parents in rating students’ cognitive abilities, more accurate in rating creativity, and similarly or less accurate in rating motivation. The literature indicates halo effects for both raters, but the effects might be stronger for teacher ratings. The correlations between teacher and parent ratings seem to fluctuate greatly in accordance with the characteristic in question. In this study, we inspected these facets of giftedness with regard to teacher-nominated gifted elementary school students.
Joint Link Between Teacher and Parent Ratings and Students’ School Achievement
Teacher and parent ratings of students are each separately important for students’ academic development. For example, more positive ratings of achievement-related characteristics are linked to higher academic performance (e.g., Fischbach, Baudson, Preckel, Martin, & Brunner, 2013; Froiland & Davison, 2014; Poropat, 2009). However, although scholars have stressed the importance of considering both teachers, parents, and their interaction in understanding students’ academic development (Christenson, 1999; Glueck & Reschly, 2014), only a few empirical studies have explored whether teachers’ and parents’ perception of student characteristics jointly affect students’ academic achievement.
A cross-sectional study by Peet et al. (1997), for instance, showed that fourth-grade students’ average school grades were better if teachers and mothers agreed in their ratings of students’ competence and engagement. Peet et al. used a difference score that did not capture the mean level on which teachers and parents agreed in their ratings (e.g., both positive, both medium, or both negative ratings). However, different relations between teacher–parent agreement and school achievement are plausible. For example, agreement in positive ratings should be associated with high student achievement, but agreement in negative ratings should be associated with lower student achievement. Benner and Mistry’s (2007) study supports this reasoning. They focused on the specific constellations of teachers’ and mothers’ agreement or disagreement in having high or low expectations about students finishing high school and going to college. In their cross-sectional study with teachers and mothers of 9- to 16-year-olds, they found that students’ combined test score for reading and mathematics was highest if both adults’ expectations about students were high and lowest when both raters’ expectations were low. In cases of disagreement, the link between low teacher expectations and low student achievement was cushioned when mothers had high expectations. However, whether the relationship between low mother expectations and low student achievement is also buffered by high teacher expectations could not be determined in their study due to a too low cell frequency.
Both studies (Benner & Mistry, 2007; Peet et al., 1997) indicate a joint association between teacher and parent ratings and students’ school achievement. However, what kind of association this is has still not been sufficiently researched. In Cohen, Cohen, West, and Aiken’s (2003) terminology, the connection might be additive, synergetic, or compensative. If the connection is additive, both ratings would be uniquely and independently related to student achievement in a summative manner after controlling for one another. If the connection is synergistic, the effects would be more than the sum of both ratings. An agreement in positive ratings would be related to even better school achievement than in the case of an additive effect. If the connection is compensative, positive ratings by either teachers or parents would diminish the relation between the respective other rating and school achievement. Positive teacher ratings would compensate for more negative parent ratings, and positive parent ratings for more negative teacher ratings. Benner and Mistry’s (2007) study might indicate either additive or compensative connections. In the present study, we aimed to determine the character of the joint relations between teacher and parent ratings of five facets of giftedness and school grades.
The Present Study
In this study, we aimed to extend the research on how teachers and parents view gifted students, and how their ratings are connected to students’ school grades. To this end, we investigated teacher and parent ratings of elementary school students who were nominated as gifted by their teachers. We concentrated on ratings of the following facets of giftedness: cognitive abilities (i.e., verbal and mathematical abilities, deductive reasoning), creative thinking, and engagement. Specifically, we pursued four research questions falling under two objectives.
Objective 1 involved a comparison of teacher and parent ratings with regard to three research questions:
We expected the accuracy of ratings on cognitive abilities to be either similar or higher for teachers (Geiser et al., 2016; Miller & Davis, 1992; Sommer et al., 2008), higher for teachers with regard to creative thinking (Geiser et al., 2016; Sommer et al., 2008), and not different or lower for teachers with regard to engagement (Genser et al., 1981; Helmke & Schrader, 1989; Spinath, 2005).
We hypothesized that we would find halo effects in both teacher and parent ratings (Li et al., 2008; Sommer et al., 2008) and that the halo effects would be stronger for teacher ratings than parent ratings (Chan, 2000; Petscher & Li, 2008; Pfeiffer et al., 2008).
We expected a medium to large correlation between teacher and parent ratings of cognitive abilities (Miller & Davis, 1992; Sommer et al., 2008), a small correlation between their ratings of creative thinking (Geiser et al., 2016), and a medium correlation between their ratings of engagement (Chan, 2000).
Objective 2 concerned the fourth research question of how teacher and parent ratings are jointly linked to school grades. We investigated whether teacher ratings and parent ratings on each of the facets of giftedness were connected additively, synergistically, or compensatively to students’ school grades in German and mathematics. According to Benner and Mistry’s (2007) results, the connection might be additive or compensative.
Method
Participants
In 2010, an enrichment program for gifted elementary school students known as the Hector Children’s Academy Program (HCAP) was established in the German state of Baden-Württemberg. Sixty-one academies belong to the HCAP. These academies are located at elementary schools and offer enrichment classes for the 10% most gifted students in the state. In each academy, about 60% of classes cover STEM-related (Science, Technology, Engineering, and Mathematics) themes, while the remaining classes address other topics (e.g., art, languages). To participate in the HCAP, students have to be nominated by their teachers, and parents must give their permission.
The data used in our study stem from a larger intervention study on the process of identifying gifted students. Here, we used a subsample of this larger intervention study for whom we had information from at least two sources out of (a) teacher ratings, (b) parent ratings, and (c) student data (i.e., test scores and self-reports). This resulted in data on 572 HCAP students, who attended 189 schools (M = 3.03, SD = 2.74, Min = 1, Max = 21). However, the combinations of the three sources differed. See Table 1 for the sample sizes and descriptive information. We received one parent rating per child. Mothers and fathers rated the child together in 32.42% of cases, only mothers in 60.81% of cases, and only fathers in 5.49% of cases (1.28% missing).
Descriptive Statistics for Teacher Ratings, Parent Ratings, and Student Data Based on All Available Data and the Intersections Between the Three Data Sources.
Estimations based on the unweighted items that loaded ≥.50 on the factors. Min = 1 (disagree), Max = 5 (agree). bEstimations based on the unweighted items that loaded ≥.50 on the factors. Value labels: among the 5% (rated 5), 10% (rated 4), 25% (rated 3), 50% (rated 2) best students of their age, or among the remaining 50% (rated 1). Treated as interval-scaled. If treated as ordinal-scaled, Mdn of all five factors = 4.00, 25th percentile for all five factors = 3.00, 75th percentile for teacher ratings of verbal abilities and creative thinking = 4.50, and for teacher ratings of mathematical abilities, deductive thinking, and engagement = 5.00. cTests for verbal abilities, mathematical abilities, and deductive reasoning: percent of correct answers, Min = 0%, Max = 100%. Test for creative thinking: Min = 1 (not at all creative), Max = 5 (highly creative). Self-report of engagement: Min = 1 (strongly disagree), Max = 4 (strongly agree). dSchool grades could range from 1 (excellent) to 6 (unsatisfactory). Report card grades below 4 are seldom assigned at the primary school level in Germany (e.g., Stubbe, Bos, & Euen, 2012).
Procedure
The data for the present investigation came from a study that took place in the first half of the 2013-2014 school year. The study was voluntary for all participants (i.e., teachers, parents, and students). To ensure participants’ anonymity, all data were collected and kept safe by the academy until the study was over. After the study, the data were coded, sorted, anonymized, and sent to the research team.
After attending information sessions about the study, 12 of 61 academies agreed to participate. The study involved two phases. First, before the HCAP classes started, we collected rating scale data from the teachers who nominated the students for the HCAP. The nomination was a global, undifferentiated judgment of students, and was conducted by registering students at an academy with their parents’ agreement. Teachers used the rating scales to rate the students they had nominated and handed the rating scales in with the students’ registration. Teachers were informed that the HCAP’s goal was to cater to the 10% most gifted students, but were otherwise given no systematic information about giftedness. Teachers received no feedback from the academy about the students they had nominated. A teacher nomination was sufficient for a student to participate in the HCAP. For example, neither an intelligence diagnostic nor school grades were a selection criterion. Most nominated students participated in the HCAP (e.g., r = .91 between nomination and participation in a different study on the HCAP; Rothenbusch, Zettler, Voss, Lösch, & Trautwein, 2016).
Second, students were tested and surveyed by trained test administrators during one of their HCAP classes. Students received an envelope from the test administrators with a questionnaire for their parents and were asked to pass it on to them. It was sent back to the academy in a sealed envelope (either through their children, who passed it on to their HCAP instructors, or via post in an enclosed stamped envelope).
Measures
We assessed students’ verbal abilities, mathematical abilities, deductive reasoning, creative thinking, and engagement on the basis of teacher ratings (TR), parent ratings (PR), and students’ tests and self-reports (S). The acronyms TR and PR will be used in the text for both the singular and plural forms of teacher and parent ratings. Furthermore, we obtained students’ most recent report card grades in German and mathematics (r = .38, p < .001) through the parent questionnaire. Descriptive statistics can be found in Table 1.
TR and PR of Student Characteristics
TR and PR were examined with newly developed rating scales for assessing verbal and mathematical abilities, deductive reasoning, creative thinking, and engagement. Decisions about what aspects to measure were based on current definitions of giftedness, such as those by Renzulli (2005) and Subotnik et al. (2011, 2012). The items for TR and PR had the same wording. However, the anchors of the 5-point scales varied for teachers and parents. Teachers rated whether students were among the 5%, 10%, 25%, 50% best students of their age (rated 5, 4, 3, or 2, respectively), or among the remaining 50% (rated 1). Since parents usually have smaller reference groups available than teachers, the parents’ scale ranged from 1 (disagree) to 5 (agree). Further information about the factor structure is presented in the Results section. The wording of the items can be found in Table 2.
Factor Loadings from the Exploratory Structural Equation Modeling With Invariant Factor Loadings Between Teacher and Parent Ratings.
Note. TR = teacher ratings; PR = parent ratings. Results from Model TR-PR2. The intended main factor loadings are in boldface.
Cronbach’s alpha based on the items that loaded ≥.50 on the corresponding factor.
Student Tests of Cognitive Abilities
Students’ cognitive abilities were measured with three subtests from a version of the Cognitive Ability Test for the Gifted (Kognitiver Fähigkeitstest für Hochbegabte, KFT-HB 3; Heller & Perleth, 2007) for gifted third graders. Because of time restrictions, we administered only one out of two existing subtests for each of the three abilities measured by the test: verbal, mathematical, and nonverbal abilities. We then generated planned missingness with a three-form design (Little & Rhemtulla, 2013). Each student took two subtests and was thus measured on two of the three abilities. We used three combinations of subtests that were block randomized across HCAP classes: (a) verbal–mathematical (N = 138), (b) verbal–nonverbal (N = 145), and (c) nonverbal–mathematical (N = 125).
On the subtest for verbal abilities, students had to select words with similar meanings (α = .60 in our study). On the subtest for mathematical abilities, students had to form mathematical equations (α = .80 in our study). On the subtest for nonverbal abilities, students had to select the correct figure in relation to the presented ones (α = .91 in our study). We used this subtest as a proxy for deductive reasoning.
Student Test of Creative Thinking
To estimate students’ potential for creative thinking, we applied a verbal open-ended unusual uses task in which students had to produce creative examples of uses for a common object. Students had 2 minutes to generate creative answers to the question, “What can you do with a wooden board?” Students were instructed to “be creative” because research has indicated that such instructions increase the validity of divergent thinking scores (O’Hara & Sternberg, 2001). We used the average creativity index as an indicator of creativity (Silvia et al., 2008). Three raters rated students’ answers with respect to the criteria of uncommonness, remoteness of associations, and cleverness on a 5-point Likert-type scale ranging from 1 (not at all creative) to 5 (highly creative). For every student and rater, the mean was calculated for the ratings of the students’ answers. The resulting three scores for each student were averaged. The intraclass correlation (ICC) between the three raters was .80. 2
Student Self-Reports of Engagement
We measured students’ engagement with the Willingness to Make an Effort (Anstrengungsbereitschaft) subscale from the Questionnaire for Assessing Emotional and Social School Experiences of 3rd and 4th Grade Primary School Children (Fragebogen zur Erfassung emotionaler und sozialer Schulerfahrungen von Grundschulkindern dritter und vierter Klassen, FEESS 3-4; Rauer & Schuck, 2003). The subscale consists of 13 items (α = .79 in our study) rated on a 4-point Likert-type scale ranging from 1 (strongly disagree) to 4 (strongly agree). An example item is “I try to solve even very difficult tasks.”
Analyses
We conducted structural equation modeling to address our research questions. The manifest indicators for the TR and PR factors were treated as interval-scaled. 3 Test and questionnaire scores were standardized separately for Grades 3 and 4, respectively, and then combined into one manifest variable for each student characteristic.
We conducted the analyses in Mplus 7.4 (Muthén & Muthén, 1998-2015) using maximum likelihood estimation with robust standard errors. We handled missing data with the full information maximum likelihood algorithm, which uses all of the information from the covariance matrices (Enders, 2010). The levels of missing data within the three combinations of two data sources were as follows: (a) TR and S: M = 7.25%, SD = 11.82%, Min = 0.75% for five teacher-rated items, Max = 36.84% for the deductive reasoning test score; (b) PR and S: M = 6.75%, SD = 9.95%, Min = 0.52% for the engagement self-report, Max = 34.55% for the deductive reasoning test score; (c) TR and PR: M = 2.77%, SD = 1.36%, Min = 0.74% for a teacher-rated item, Max = 7.75% for students’ math grade.
Our data had a nested structure because the students in the sample belonged to different school classes (M = 1.96 students per school class, SD = 1.49). Each class had a different teacher. Hence, the students’ test and questionnaire data, grades, and parent ratings could be divided into different classes or teachers. Furthermore, students in the same class were rated by the same teacher. Because this dependence can lead to a violation of the assumption of traditional analysis approaches (e.g., ordinary least squares regression) that residuals are uncorrelated (Snijders & Bosker, 2012), we inspected ICC, which can range from 0 (total independence of observations from the cluster variable) to 1 (maximum dependence of observations on the cluster variable). In the present study, ICCs varied between ICC < .01 for an item from the PR scale for creative thinking and ICC = .57 for an item from the TR scale for engagement (M = .21, SD = .19). Even small ICCs of .05 or .01 can bias estimations in conventional ordinary least squares regression (Cohen et al., 2003). Therefore, we applied the “type = complex” procedure in Mplus 7.4 to adjust the standard errors of the correlation and regression coefficients (for more information, see Muthén & Satorra, 1995).
The specific steps we undertook to test our hypotheses were as follows: First, we analyzed the factorial structures of TR and PR. Second, we compared TR, PR, and S by inspecting their correlations (Objective 1). Third, to investigate the relations between TR, PR, and S and students’ German and math grades (Objective 2), we inspected correlations and ran separate models regressing the two grades (dependent variables) on TR, PR, and S (independent variables) for each student characteristic. A more thorough description of the three steps of the data analysis follows.
Factor Structure of TR and PR
To examine the factor structure of TR and PR, we ran two exploratory structural equation modeling (ESEM) models with five factors each (Model TR-5F for TR and Model PR-5F for PR). ESEM integrates exploratory factor analysis, confirmatory factor analysis (CFA), and structural equation modeling. This analysis approach has several advantages. For example, even small cross-loadings can inflate the correlations of CFA factors (Marsh et al., 2010), while ESEM does not have the restrictive assumption required in CFA that there be no cross-loadings (Marsh, Morin, Parker, & Kaur, 2014). Furthermore, ESEM allows measurement invariance testing. As recommended by Marsh et al. (2010) and Marsh, Nagengast, and Morin (2012), we used an oblique geomin rotation with an epsilon value of 0.5 (the default in Mplus 7.4).
We further inspected models with a second-order factor and five first-order factors for TR (Model TR-2nd) and PR (Model PR-2nd). For this, we changed to the ESEM-within-CFA (EWC) framework because specifying a second-order factor is not possible with ESEM factors in Mplus 7.4. More information on EWC can be found in an overview on ESEM by Marsh et al. (2014). We assessed the fit of the models (see the passage on “goodness of fit” at the end of the Analyses section) and the factor loadings, following the rule of thumb that items should have loadings of at least .30 on the target factor and should not have cross-loadings above .30 on any other factor (Cudeck, 2000; Tinsley & Tinsley, 1987).
To test how TR and PR were correlated, we combined the sets of ratings for TR and PR with the best model fit (see Results section) into a multi-trait (e.g., verbal abilities and creative thinking) multi-method (i.e., teacher and parent ratings; multitrait multimethod [MTMM]) model (Model TR-PR1). In Model TR-PR1, the factor loadings were allowed to vary freely, and we investigated whether the items still had high loadings on the target factors and low cross-loadings on other factors. To ensure the comparability of raters, we tested for metric invariance. Thus, we restricted the factor loadings between raters to invariance in Model TR-PR2 and tested whether the Model TR-PR2 had a worse fit than Model TR-PR1 (see the passage on “goodness of fit”; Meredith, 1993).
Objective 1: Comparison of TR and PR
Objective 1 was divided into three research questions. First, to examine the rating accuracy, we investigated the correlations between TR and the corresponding S on the one hand, and the correlations between PR and the corresponding student data on the other. Second, we analyzed halo effects by comparing the correlations within TR and PR, respectively, with the corresponding correlations within S. Third, we studied the correlations between TR and PR for the same student characteristic.
The correlations inspected to answer the three research questions all stemmed from the MTMM Model TR-PR-S (see Figure 1) that combined the sets for TR and PR (from Model TR-PR2, under the condition that this model did not have a worse fit than Model TR-PR1; see the passage on “goodness of fit”) with the manifest variables for S. The covariance estimates were standardized to obtain correlations. To test whether the correlations differed statistically significantly from each other, we applied Fisher’s z transformations to the correlations, calculated the differences between them, and tested whether these differences were statistically significantly different from 0.

Multitrait multimethod (MTMM) Model TR-PR-S of teacher ratings, parent ratings, and student data concerning the following student characteristics: verbal ability (VA), mathematical ability (MA), deductive reasoning (DR), creative thinking (CT), and engagement (E).
Objective 2: Connections Between TR and PR and Students’ School Grades
To investigate our second objective, we added German and math grades as manifest variables to Model TR-PR-S (resulting in Model TR-PR-S-G). This allowed us to inspect how TR, PR, and S are related to German and math grades. In two further steps, we investigated the joint associations between TR and PR and school grades using regression analyses. We applied the EWC framework to test whether TR and PR were related to students’ school grades additively, synergistically, or compensatively. In our case, we investigated the specific effects of two corresponding TR and PR factors per model and aimed to specify the latent interaction between the TR and PR of the same characteristic. This is not possible with ESEM factors in Mplus 7.4. We used the parameter estimates from Model TR-PR-S-G to estimate the TR and PR factors in the EWC models. Two models per characteristic were estimated (10 models in total). The first model for each characteristic (Models VA1, MA1, DR1, CT1, and E1) included school grades as the dependent variable and ratings and students’ tests or self-reports as predictors to test for an additive connection between TR, PR, and school grades. To test for interaction effects, latent interaction terms for TR and PR were added to the analysis in the second model for each characteristic (Models VA2, MA2, DR2, CT2, and E2). A statistically significant positive interaction coefficient would indicate a synergistic effect, while a statistically significant negative interaction coefficient would point to a compensative effect (Cohen et al., 2003). See Figure 2 for an example of the models.

Example of the regression analyses for Objective 2. Depicted is Model E2.
Goodness of Fit
We evaluated the model fit and used the χ2 goodness-of-fit statistic and the root mean square error of approximation (RMSEA), for which values below .08, .05, and .01 indicate a mediocre, good, and excellent fit to the data (MacCallum, Browne, & Sugawara, 1996), respectively. Furthermore, we inspected the comparative fit index (CFI), the Tucker–Lewis Index (TLI), and the standardized root mean square residual (SRMR). For the TLI and CFI, values above .90 indicate an acceptable fit to the data, and values above .95 an excellent fit, while for the SRMR, values below .08 are considered to indicate good model fit (Hu & Bentler, 1999).
We conducted Δχ2 tests to compare the relative fit of nested models in the analysis of the factor structure of TR and PR and the invariance of factor loadings between TR and PR, which should be statistically nonsignificant if nested models fit the data comparably well. No model fit estimates were available in Mplus 7.4 for latent moderated structural equations (Muthén & Muthén, 1998-2015). Therefore, to address Objective 2, log-likelihood ratio tests were used after inspecting the model fit indices in the models without latent interaction terms to determine whether the more parsimonious models (without a latent interaction term) exhibited significantly worse fit than the more complex models (including the latent interaction term). The log-likelihood ratio test should be statistically significant in the case of a significant loss in fit. Otherwise, the more parsimonious model had no significant loss in fit compared with the more complex model (Maslowsky, Jager, & Hemken, 2015). Finally, if the log-likelihood ratio test was statistically significant, we inspected the estimates from the model with the latent interaction term.
Because multiple significance testing increases the probability of falsely rejecting the null hypothesis (i.e., inflated Type I error rate), we applied the Benjamini–Hochberg procedure (Benjamini & Hochberg, 1995). By controlling for the false discovery rate instead of limiting the family-wise error rate, the Benjamini–Hochberg procedure yields more power than the Bonferroni technique (Williams, Jones, & Tukey, 1999). Consequently, p values ≤.029 were considered statistically significant at an overall level of α = .05 in this study.
Results
Factorial Structure of Teacher and Parent Ratings
We tested the five-factor structure (i.e., one factor for each rated characteristic) of the parent and teacher ratings separately with an ESEM model and specified a EWC model with a second-order factor and five first-order factors. Model fit indices are given in Table 3. The five-factor models (i.e., Model TR-5F and Model PR-5F) had acceptable to good fits to the data. 4 All items had loadings of at least .55 (mean loading = .84) for TR and .41 (mean loading = .73) for PR on the intended factors, while cross-loadings did not exceed .17 for TR and .27 for PR. The correlations between the dimensions were large for TR (mean r = .64, SD = .09) and medium for PR (mean r = .34, SD = .15). Both second-order factor models (Model TR-2nd for TR and Model PR-2nd for PR) exhibited a statistically significant loss in model fit in comparison to Model TR-5F, Δχ2(5) = 27.596, p < .001, and Model PR-5F, Δχ2(5) = 13.014, p < .02.
Structure of Teacher and Parent Ratings (Models TR-5F, TR-2nd, PR-5F, PR-2nd, TR-PR1, and TR-PR2).
Note. CFI = comparative fit index; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; FL = factor loading.
Hence, we combined the ESEM sets for TR and PR with five factors each into one model (see Table 3). In Model TR-PR1, all parameters were allowed to vary freely to test for configural invariance. The model’s fit indices indicated a good fit to the data. In Model TR-PR2, we set the factor loadings to invariance between raters to test for metric invariance. Model fit did not worsen, Δχ2(65) = 77.226, p > .05, indicating comparable factor structures for TR and PR. Details on the items and factor loadings for Model TR-PR2 are presented in Table 2.
Objective 1: Comparing TR and PR
To reach Objective 1, we compared TR and PR in relation to their accuracy levels, explored whether they were affected by halo effects, and investigated the correlations between TR and PR with regard to the same characteristics. Therefore, we added students’ test scores and self-report (S) to Model TR-PR2 in order to obtain correlations (i.e., Model TR-PR-S). Model fit was good (see Table 4). The correlations can also be found in Table 4. All results for Objective 1 stem from Model TR-PR-S.
Correlations Between Teacher Ratings, Parent Ratings, and Students’ Test Scores and Self-Reports (Model TR-PR-S).
Note. RMSEA = root mean square error of approximation; CFI = comparative fit index; TLI = Tucker–Lewis index; SRMR = standardized root mean square residual. Results from Model TR-PR-S, χ2(640) = 938.456, RMSEA = .029, CFI = .973, TLI = .965, SRMR = .034. Based on the adjustment of significance tests with the Benjamini–Hochberg procedure (1995), p values ≤.029 are considered statistically significant at an overall level of α = .05.
p ≤ .029. **p < .010. ***p < .001.
Differences in the Accuracy of TR and PR
To investigate the accuracy of TR and PR, we inspected the correlations between TR or PR and the corresponding student data. We expected TR to be either similarly or more accurate than PR with regard to verbal abilities, mathematical abilities, and deductive reasoning, more accurate for creative thinking, and either similarly or less accurate for engagement.
Table 4 presents the correlations between the ratings and the student data in Sections 3 and 5. TR (rTR = .35, p < .01) and PR (rPR = .31, p < .01) were moderately associated with the student data for verbal abilities. Both ratings were weakly associated with test scores for mathematical abilities (rTR = .26, p < .01; rPR = .22, p < .01) and deductive reasoning (rTR = .24, p = .01; rPR = .18, p < .01), and not statistically significantly associated with test scores for creative thinking (rTR = -.05, ns; rPR = .04, ns). TR were not statistically significantly associated (rTR = .15, ns) with students’ self-reported engagement, while PR were moderately correlated (rPR = .37, p < .01). TR and PR did not differ statistically significantly from each other in the strength of their associations with any of the students’ data (all ps > .030). These statistically nonsignificant results concerning the accuracy of TR and PR were as expected for verbal abilities, mathematical abilities, deductive reasoning, and engagement but not for creative thinking, for which we had expected TR to be more accurate.
Difference in Halo Effects in TR and PR
We hypothesized that ratings of different characteristics would be more highly intercorrelated in TR and PR than in the corresponding student data (i.e., halo effects). Furthermore, we expected that TR of different characteristics would be more strongly intercorrelated than PR of different characteristics.
The correlations for each of the three methods (i.e., teachers, parents, and students) can be found in Table 4 in Sections 1, 4, and 6. Figure 3 displays the differences in correlations graphically. A comparison of the correlations for TR, PR, and S revealed that all ten correlations were statistically significantly higher for TR than for S (see TR–S in Figure 3). Nine out of 10 correlations were statistically significantly higher for PR than for S (see PR–S in Figure 3). Furthermore, correlations for TR were statistically significantly higher than for PR in all cases (see TR–PR in Figure 3). Therefore, the results indicate that the ratings were affected by halo effects, but TR more strongly than PR.

Differences between corresponding heterotrait-monomethod correlations of teacher ratings and parent ratings (TR–PR), teacher ratings and student data (TR–S), and parent ratings and student data (PR–S) are depicted.
Correlations Between TR and PR
We expected to find medium to large correlations between TR and PR for verbal abilities, mathematical abilities, and deductive reasoning, a small correlation for creative thinking, and a moderate correlation for engagement. The correlations are depicted in Table 4 on the diagonal of Section 2. TR and PR were moderately correlated for verbal abilities and mathematical abilities (rVA = .31, p < .01; rMA = .40, p < .01), not statistically significantly correlated for deductive reasoning and creative thinking (rDR = .12, ns; rCT = .10, ns), and weakly correlated for engagement (rE = .28, p < .01). Hence, the correlations for ratings of verbal abilities, mathematical abilities, and creative thinking were as expected, but the correlations for ratings of deductive reasoning and engagement were lower than anticipated.
Objective 2: Connections Between TR and PR and Students’ School Grades
We examined whether TR and PR were related additively, synergistically, or compensatively to students’ German and math grades. Their bivariate correlations with school grades can be found in the first columns of Table 5. We ran two models for each characteristic (i.e., 10 models in total; see Table 5): one model to inspect whether TR and PR were statistically significant additive predictors of students’ school grades, and one model to test for a synergistic or compensative relation between TR and PR and school grades with a latent interaction term. For ease of interpretation, we recoded the grading scale so that higher values indicate better school grades. All model fits were acceptable or good (see Table 5).
Associations Between Teacher and Parent Ratings and Students’ School Grades, Controlling for Students’ Test Scores and Self-Reports (Models TR-PR-S-G, VA1, VA2, MA1, MA2, DR1, DR2, CT1, CT2, E1, and E2).
Note. RMSEA = root mean square error of approximation; CFI = comparative fit index; TLI = Tucker–Lewis index; SRMR = standardized root mean square residual. Based on the adjustment of significance tests with the Benjamini–Hochberg procedure (1995), p values ≤.029 are considered statistically significant at an overall level of α = .05.
Correlations based on Model TR-PR-S-G, χ2(692) = 999.220, RMSEA = .028, CFI = .973, TLI = .965, SRMR = .033. bModel fit indices for Models VA1, MA1, DR1, CT1, and E1: RMSEA = .028-.037, CFI = .957-.977, TLI = .947-.972, SRMR = .032-.059. cGrading scales inverted for ease of interpretation; higher values indicate better school grades.
p ≤ .029. **p < .010. ***p < .001.
Model VA1 showed that both TR and PR of verbal abilities were statistically significantly positively associated with German grades after controlling for students’ test scores for verbal abilities. Test scores were not statistically significantly related to grades. Based on a log-likelihood ratio test, adding the interaction term (Model VA2) did not increase the model fit relative to Model VA1, DVA1 vs.VA2 (1) = 1.280, p > .05. Hence, the results point to an additive rather than a synergistic or compensative connection between TR and PR and German grades.
Model MA1 showed that both TR and PR of mathematical abilities were statistically significantly positively associated with math grades after controlling for students’ test scores for mathematical abilities. The log-likelihood ratio test comparing the two models was statistically significant, DMA1 vs. MA2 (1) = 5.010, p = .025, indicating that Model MA2 fit better the data than Model MA1. In Model MA2, aside from the statistically significant positive relations between students’ test scores for mathematical abilities and math grades, the interaction term for TR and PR was statistically significantly negative. These results indicate a compensative connection between TR and PR and math grades. The interaction effect is illustrated in Figure 4. Parents’ positive ratings buffered the association between TR and math grades, while teachers’ positive ratings buffered the relation between PR and math grades. Hence, the grades of students with negative TR and either positive or negative PR differed more strongly than those of students with positive TR who had either positive or negative PR.

Effects of interactions between teacher ratings and parent ratings on students’ math grades.
We examined the associations between TR, PR, and student data on deductive reasoning, creative thinking, and engagement on the one hand and German and math grades on the other. Each of the regression models had both grades as dependent variables. The models with interaction terms for deductive reasoning (Model DR2) and creative thinking (Model CT2) did not have statistically significantly better model fit than the corresponding models without interaction terms, DDR1 vs. DR2(2) = 3.684, p > .05; DCT1 vs. CT2(2) = 2.321, p > .05. However, the log-likelihood ratio test comparing the two models was statistically significant for engagement, DE1 vs. E2(2) = 27.063, p < .001.
A closer examination of the results concerning the connection between TR and PR of deductive reasoning (Model DR1), creative thinking (Model CT1), and engagement (Model E2) and German grades revealed similar patterns to the one reported for verbal abilities. Both TR and PR were statistically significant positive predictors of German grades in an additive manner after controlling for student data, which were not statistically significantly related to German grades. Although Model E2 had better model fit than Model E1, the latent interaction term was not statistically significant, thus speaking against a synergistic or compensative connection between TR and PR of engagement and German grades.
Concerning math grades, the models for deductive reasoning (Model DR1) and creative thinking (Model CT1) showed that only TR but not PR were statistically significantly positively associated with students’ math grades after controlling for student data. Students’ test scores for deductive reasoning were statistically significantly related to math grades, but their test scores for creative thinking were not. However, the interaction term between TR and PR in the model connecting engagement to math grades (Model E2) was statistically significantly negative, indicating a similar pattern to the one found for ratings of mathematical abilities (see Figure 4).
In conclusion, TR and PR were additively connected to German grades. TR and PR were either both compensatively connected to math grades, indicating that positive PR or TR buffered the relation between the respective other rating and math grades (i.e., for ratings of mathematical abilities and engagement) or only TR were statistically significantly connected to math grades (i.e., for ratings of deductive reasoning and creative thinking).
Discussion
In this study, we examined teacher and parent ratings of facets of giftedness in elementary school students who have been nominated by their teachers for a statewide enrichment program. Our study pursued two objectives: first, we compared teacher and parent ratings of students’ cognitive abilities (i.e., verbal and mathematical abilities and deductive reasoning), creative thinking, and engagement with each other with regard to accuracy levels and halo effects. We further examined the correlations between teacher and parent ratings of the same characteristics. Second, we analyzed how teacher and parent ratings of the same characteristic are related to students’ German and math grades.
Comparing Teacher and Parent Ratings of Teacher-Nominated Gifted Elementary School Students
Difference in the Accuracy of Teacher and Parent Ratings
Most previous work on teacher and parent ratings of students’ cognitive abilities, creativity, and motivation has taken place in general education settings (e.g., Geiser et al., 2016; Sommer et al., 2008) and has seldom focused specifically on gifted students (see Chan, 2000, for an exception with gifted secondary school students). In our study of teacher-nominated gifted elementary school children, the accuracy of teacher ratings on all five facets of giftedness did not differ statistically significantly from the accuracy of parent ratings. For example, both teacher and parent ratings of students’ verbal abilities were moderately correlated with students’ test scores for verbal ability, and the difference between the two correlations was not statistically significant. Therefore, although teachers and parents have different perspectives on students (Harder et al., 2015; Petscher & Li, 2008), they seem to be equally able to rate facets of students’ giftedness. However, the students in our sample had already been rated as gifted by their teachers, and this range restriction might have underestimated or otherwise distorted the accuracy differences between teacher and parent ratings.
The statistically nonsignificant differences between parent and teacher ratings of cognitive abilities was similar to the results reported by Sommer et al. (2008), although Miller and Davis (1992) showed higher correlations between ratings and test scores for verbal, mathematical, and figural abilities for teachers than for parents. The statistically nonsignificant difference in creativity ratings was in contrast to previous research, which reported higher accuracy levels for teacher ratings than parent ratings (Geiser et al., 2016; Sommer et al., 2008).
One reason why teachers were not more accurate than parents in rating teacher-nominated gifted students might be that parents’ but not teachers’ accuracy in rating students’ cognitive abilities have been found to be higher if the children had high achievement levels and school grades (Miller & Davis, 1992). The authors offered several explanations, such as that parents who can rate their children more accurately might be providing their children with more appropriate supports, resulting in more competent children (i.e., match hypothesis; Hunt & Paraskevopoulos, 1980) or that it might be easier for parents to detect higher-achieving children. Furthermore, parents’ knowledge of their children’s status as gifted (as determined by teachers) might have led them to pay greater attention to characteristics that might be connected to their children’s status as gifted. Methodologically, our study inspected correlations indicating the similarity of rank orders between two variables. Our study focused on a student sample with only a few students from any given class, whereas other studies (e.g., Miller & Davis, 1992; Sommer et al., 2008) used entire classes. Hence, teachers’ tendency to use their specific class instead of all same-aged students (as instructed) as a reference for their ratings (Rothenbusch et al., 2016) might have reduced the correlations more strongly in our study than in others.
However, teachers and parents’ accuracy levels when rating cognitive abilities and creative thinking tended to be lower in this study than in other studies that did not focus exclusively on teacher-nominated gifted students (e.g., Machts et al., 2016; Miller & Davis, 1992; Sommer et al., 2008; Urhahne, 2011). The teacher nominations might have led to a sample with limited variance in these characteristics in comparison to samples from general classrooms, artificially reducing the correlations for statistical reasons (Wild, 1993). Furthermore, teachers’ decision to nominate a student as gifted and parents’ knowledge of their children’s nomination status might have induced beliefs about how these students should be (e.g., Baudson, 2016; Baudson & Preckel, 2016) and therefore distorted teachers and parents’ ratings. Studies of teacher and parent ratings of students with and without gifted status are needed to explore the reasons for the levels of accuracy and the (non)differences between teacher and parent ratings found in this study.
In line with previous studies, parent ratings of engagement were moderately connected to students’ self-reports (Helmke & Schrader, 1989), whereas teacher ratings were correlated only weakly and not statistically significantly with them (Spinath, 2005). Although the difference between teacher and parent ratings was not statistically significant in our study, this result should be investigated further. It might indicate that parents can rate aspects of students’ engagement better than teachers can, perhaps because parents have the opportunity to observe their children in learning-relevant situations outside of school, such as completing homework, preparing for tests, or talking about school and learning content.
Difference in Halo Effects in Teacher and Parent Ratings
Our results are in agreement with previous research indicating the presence of halo effects (e.g., Li et al., 2008; Petscher & Li, 2008; Urhahne, 2011). Cognitive abilities, creative thinking, and engagement were more strongly correlated when rated by teachers and—with the exception of the correlation between ratings of verbal and mathematical abilities—by parents as well than they were in the student data (i.e., test scores and self-reports). Furthermore, our study found differences in the strength of halo effects in line with the descriptive results reported by Chan (2000), Petscher and Li (2008), and Pfeiffer et al. (2008): teacher ratings were more strongly intercorrelated than parent ratings. Hence, both teacher and parent ratings indicated that the teacher-nominated gifted students in our study had rather homogeneous profiles, whereas the student data revealed a greater diversity of combinations of strengths and weaknesses on the characteristics under investigation.
Fisicaro and Lance (1990) offered three explanations for halo effects: (a) a general impression, (b) a salient characteristic, or (c) an inability to discriminate between factors. Our analyses found no support for a second-order factor across all teacher ratings or all parent ratings. This might indicate that neither a general global impression nor a single salient factor such as the students’ nomination for the enrichment program inflated all the investigated correlations per se. However, factors like a giftedness label or an academic achievement factor like the one found by Anders, McElvany, and Baumert (2010) for teachers including cognitive abilities and motivation might have influenced some but not all correlations between the rated characteristics. Our analyses did not support the third explanation either, as both groups of raters were able to discriminate between the rated characteristics.
The reason for the higher correlations between teacher-rated characteristics than parent-rated characteristics might be connected to Gräf and Unkelbach’s (2016) finding that positive information leads to stronger halo effects than negative information. As the rated characteristics were all important variables for school performance, they might be especially valued by teachers and thereby have a more positive valence for teachers than parents. Having a child with high academic potential is likely to have a positive valence for parents as well, but other characteristics might be present to a greater extent, more relevant for day-to-day life, or more valued. Overall, our study could demonstrate halo effects of different strengths for teachers and parents, but additional research is necessary to explore the mechanism behind this difference.
Correlations Between Teacher and Parent Ratings
Teachers and parents agreed that the students in this study should participate in a statewide enrichment program. Although this agreement might have led to higher correlations between their ratings of facets of giftedness, the correlations found were in fact not higher than those reported in studies with general classrooms samples: in alignment with these studies, teacher and parent ratings were moderately associated for verbal and mathematical abilities (Miller & Davis, 1992) and weakly correlated for creative thinking (Geiser et al., 2016). The correlations between teacher and parent ratings of deductive reasoning and engagement were lower than expected (Chan, 2000; Miller & Davis, 1992). The low to at most modest associations between teacher and parent ratings might be connected to the relatively low observability of the item content (e.g., “understands abstract ideas”), high subjectivity (e.g., “good ideas” and “original solutions”), or differences in what situations are observed by parents and teachers.
The Association Between Teacher and Parent Ratings and School Grades
Our study offers important empirical support for the hypothesis that teachers and parents separately but also jointly predict students’ academic development (Glueck & Reschly, 2014). Specifically, it provides a detailed picture of different patterns of connections between teacher and parent ratings and students’ German and math grades.
In agreement with Benner and Mistry (2007), teachers’ and parents’ positive ratings of facets of giftedness were jointly related to good German grades, while negative ratings of these characteristics by both raters were linked to worse German grades—even though students had the same test or self-reported scores. Furthermore, our study indicated that the relationship between the two sources of ratings and German grades was additive in nature. Hence, both ratings sources are independently related to German grades, and their joint relationship with German grades is equal to the sum of their separate relationships. This result is of particular importance, as it emphasizes that teacher and parent ratings should be considered in conjunction with one another for such purposes as selecting students with high verbal performance in school or attempting to transform high verbal ability into high verbal achievement. However, there does not seem to be an interaction effect between teachers’ and parents’ ratings—at least not in a multiplicative manner that would have led to stronger or weaker associations between one rater’s rating and German grades depending on the other rater’s rating.
The results for the relation between ratings of facets of giftedness and math grades are more complex. In all cases, more positive teacher ratings were linked to better math grades after controlling for students’ test or self-reported score. However, various associations for ratings of different characteristics were apparent: only teachers’ but not parents’ ratings of students’ deductive reasoning and creative thinking were statistically significantly associated with math grades. In combination with the low to moderate accuracy and weak correlations of teacher and parent ratings for these two characteristics, this might indicate that teachers used more information about the students relevant for mathematics than parents when rating these characteristics. For ratings of mathematical abilities and engagement, the associations between teacher or parent ratings and students’ math grades were diminished when the other rater’s rating of that characteristic was positive. These results are similar to Benner and Mistry’s (2007) findings that positive parent ratings buffer negative teacher ratings and further showed that, vice versa, positive teacher ratings buffered negative parent ratings. Therefore, extending Helmke’s (2012) notion that it is educationally beneficial for teachers to slightly overestimate students’ abilities, our results indicate that teachers’ or parents’ ratings of students can be assets if they are higher than the other’s ratings—even among students who were nominated as gifted by teachers and had rather good grades already.
However, the explained variance in school grades was rather low, showing that teacher and parent ratings are a piece, but a rather small one, of the puzzle of what explains students’ school grades. Nevertheless, the mechanisms behind the joint associations between ratings and school grades should be investigated in addition to exploring the role of actual versus perceived agreement and disagreement. For example, ratings might influence adults’ expectations, feedback, and/or suggestions about students’ learning-related attitudes and behavior. When teachers and parents share positive ratings, they might create similarly highly stimulating learning environments at home and at school, and this congruence might in turn lead to better school performance (Christenson, 1999; Glueck & Reschly, 2014). Peet et al. (1997) reasoned that divergent ratings produce conflicting learning environments, resulting in mixed signals that might diminish a student’s school achievement. However, this might still be preferable to comparably discouraging learning environments based on similarly negative ratings by parents and teachers. Another (additional) possible mechanism behind our results may be that good and poor school achievement are easier to observe than average-level school achievement, leading to greater agreement between teachers and parents in their ratings of high- and low-performing students. Furthermore, combining teacher and parent ratings to create one diagnosis might be more reliable if they agree instead of disagree in their (high or low rather than mid-level) ratings of student characteristics.
Strengths and Limitations
The results presented here must be viewed in light of certain strengths and limitations of our study. First, we combined three important perspectives on facets of giftedness: teacher ratings, parent ratings, and student data (i.e., tests and self-reports). In addition, we measured a comprehensive range of characteristics. Therefore, we were able to investigate how ratings of different characteristics were connected.
Second, our focus on teacher-nominated gifted elementary school students expands the research on teacher and parent ratings, which has mostly been conducted on students from general classrooms (e.g., Geiser et al., 2016; Sommer et al., 2008). However, teacher-nominated gifted students are likely to differ somewhat from students identified as gifted on the basis of ability and achievement tests (Acar et al., 2016) or parent nominations. Furthermore, we focused on elementary school students, and our study was conducted in a German state. Students came from many different elementary schools in that state, so the results can be considered relatively representative of teacher and parent ratings of teacher-nominated gifted elementary school students in this region. However, the generalizability to ratings of, for example, teacher-nominated students of another age or nationality, parent-nominated gifted elementary school students, or students who are labeled as gifted on the basis of their test scores must be scrutinized.
Third, the study was conducted within an existing gifted education program, thereby increasing the study’s ecological validity and relevance for practice. However, several issues could not be avoided or adequately controlled for. Most prominent is the problem that parents were aware of the teacher nomination, as their children were already participating in the HCAP. Therefore, parents were probably influenced by teachers. We cannot determine how the ratings would have related to each other if parents had been unaware of their children’s nomination status or parents had been the ones to nominate their children as gifted. For example, parents might be more willing to accept teachers’ positive nomination decisions than teachers would have been of parents’ positive nomination decisions. Furthermore, students attended one or more courses from a broad spectrum of fields (e.g., MINT, arts, languages), but we could not control for course content. Hence, different patterns based on content areas might be hidden in our results.
Fourth, we used report card grades for German and mathematics as indicators of students’ academic achievement. School grades are frequently used as indicators of academic achievement but are also a form of teacher ratings (Alvidrez & Weinstein, 1999). They have extensive consequences for students’ academic careers, as they provide feedback for students and parents, are connected with students’ academic self-concept, and are listed on students’ school leaving certificates in most countries (Südkamp, Kaiser, & Möller, 2012). Nevertheless, future research should also include more objective measurements such as student achievement tests.
Finally, our results were based on a cross-sectional study. Hence, conclusions about the direction of causality between ratings and students’ school grades cannot be made. Longitudinal data are needed to examine the (bi-)directionality between ratings and school grades and between teacher and parent ratings.
Conclusion
It has been theorized that students perform best when teachers and parents agree on topics such as how to support and guide students (Christenson, 1999; Glueck & Reschly, 2014). To accomplish this, they have to interact on the basis of common ground. Their ratings of student characteristics relevant to giftedness, competence, and achievement might be one piece of this common ground. This study extends the literature on teacher and parent ratings by focusing on teacher-nominated gifted elementary school students and on the joint relations between teacher and parent ratings and students’ school grades.
Teacher–parent agreement on support can be understood as agreement on what type and areas of gifted education would be suitable for a given student. Similar views of students might enhance the probability of agreement. We reported low to moderate correlations between teacher and parent ratings of facets of giftedness and similar but low to moderate accuracy levels for both teacher and parent ratings. Hence, teachers and parents might profit from joint interventions to bring their views into greater alignment and work on their diagnostic competence in order to minimize the chance of jointly choosing a type and area of gifted education that does not fit the student’s profile.
With regard to how this common ground in ratings relates to supporting students’ school achievement, our study shows that better ratings by parents, and to an even greater extent by teachers, are associated with better school grades. Hence, agreement in ratings is only a favorable constellation when ratings are positive. Otherwise, disagreement might be a more favorable constellation, because more positive ratings on mathematical ability and engagement by either teachers or parents in comparison to the other was associated with less bad grades in German and even more so to less bad math grades.
Ultimately, administrators of gifted education programs set important framing conditions for their programs. Administrators should be aware that both teacher and parent ratings of facets of giftedness were at best moderately correlated with student data. Furthermore, halo effects affected both ratings, but parent ratings to a lesser degree. If teacher-nominated gifted students are in a certain course because teachers or parents or both overestimated them or because one salient and positively rated characteristic spilled over into other ratings, students might still have a chance to live up to these positive ratings. However, if instructors use parent or teacher ratings to adapt their courses to participants’ learning profiles, they might not be adequately prepared to support their students. Furthermore, if only ratings are available for programs to identify which of their participants are verbally and mathematically high achieving, they should use teacher ratings but supplement them with parent ratings.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The work reported herein was supported by grants from the Hector Foundation II.
Notes
Author Biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
