Abstract
Throughout the United States there is an increasing trend toward using value-added methods (VAMs) for high-stakes decisions. When policymakers use VAMs to identify, reward, and dismiss teachers, they may perpetuate the egg-crate model of schooling and undermine efforts to build instructional capacity schoolwide. At any time, in any school, some teachers are more knowledgeable, experienced, and skilled than others. Schools function best when they continuously leverage teachers’ expertise so that all students in all classrooms are well served. Drawing from research about the incentives and norms that influence teachers’ work within schools, this article illustrates what can happen when these methodologies are used to make job decisions and it identifies the hazards of using VAMs for this purpose. Contextualizing this within the larger discussion about performance evaluation systems, the article suggests how VAMs can be used productively as one source of information to promote improvement schoolwide.
In two sets of widely cited studies using value added methods (VAMS), researchers established clearly that the teacher is the most important school-level factor in students’ learning and that their effectiveness varies widely within schools (McCaffrey, Koretz, Lockwood, & Hamilton, 2004; Rivkin, Hanushek, & Kain, 2005; Rockoff, 2004). These studies have had far-reaching consequence for policy and practice. Evidence that teachers working in the same school have notably different effects on their students’ learning have led many reformers not only to focus on the individual teacher as the most promising unit of change but also to take advantage of VAMS as a mechanism for documenting differences in teachers’ performance.
Within a few years of these studies’ release, states and local districts introduced policies for assessing individual teachers’ performance and incorporating VAMS ratings in teacher evaluations. These actions were fueled, in part, by federal Race to the Top competition requirements, specifying that states and districts applying for grants incorporate measures of student achievement into teacher evaluations and “at a minimum” use those evaluations to inform decisions about “compensating, promoting, and retaining” teachers (U.S. Department of Education, 2010), all decisions that teachers would regard as carrying high stakes for them personally. Soon, 35 states and the District of Columbia required student achievement to be a “significant” or “the most significant” factor in teachers’ evaluations, and most used VAMS as a part of that determination (Layton, 2014). Current policies that incorporate measures of student achievement in teacher evaluation weigh that component at 30% to 50% of the teacher’s total score, with the rest dedicated to evaluators’ assessments of instruction based on observations and teachers’ professional behavior (Goldhaber, this issue, pp. 87–95; The New Teacher Project [TNTP], 2012).
Little explanation has been provided for why such wide variation in teachers’ effectiveness exists within schools, although the implicit assumption seems to be that teacher quality inheres in the individual rather than the school organization and that some teachers simply are more effective than others (for further analysis, see Kennedy, 2010). In today’s parlance, teachers who are better educated, skilled, and experienced are assumed to have more “human capital” to contribute than those with weak credentials, little expertise, or minimal experience (Auguste, Kihn, & Miller, 2010).
Based on this research about variation in teachers’ effectiveness within schools, some analysts argue that schools could be improved substantially by dismissing teachers who have low VAMS ratings and replacing them with teachers who have average or higher ratings. Hanushek (2009) proposed such an approach by urging policymakers and administrators to use student test scores to terminate, or “deselect,” the bottom 5% to 10% of teachers, replacing them with average or excellent teachers. Hannaway (2009) made a similar proposal to improve schools in Washington, DC. Subsequently, in a widely publicized study, Chetty, Friedman, and Rockoff (2013) used VAMS to estimate that the long-term effects of a single teacher on the lifetime earnings of a class of students was $266,000, thus giving further scholarly heft to claims from nonprofit organizations, such as The New Teacher Project (2010), and Newsweek (Thomas, Wingert, Conant, & Register, 2010) that the answer to failing schools is to “fire bad teachers” and increase a school’s human capital one teacher at a time. These strategies for improvement gained credibility among policymakers, who widely believe that VAMS accurately assesses the individual teacher’s contribution to students’ learning.
However, scholarly panels reviewing research about VAMS have raised serious concerns about the validity and reliability of the approach (Baker et al., 2010; National Research Council & National Academy of Education, 2010). These experts report that although VAMS might serve broad purposes, such as informing program evaluation, they are not sufficiently accurate and stable for making high-stakes decisions that affect individuals, such as annually reappointing probationary teachers, awarding or withholding tenure, dismissing tenured teachers, or selecting teachers for layoffs. In April 2014, the American Statistical Association (ASA) issued similar cautions about the use of VAMS in education, warning that it is “counterproductive” to “[attach] too much importance to a single item of quantitative information” and to do so “can be detrimental to the goal of improving quality” (p. 5).
Whether one regards a policy that relies on a VAMS score for 30% to 50% of a teacher’s evaluation as extreme—that is, “attaching too much importance to a single item of quantitative information” as the ASA (2014) report suggests—likely depends on one’s knowledge about, perspective on, and stake in any decision. Proponents suggest various ways by which public schools would benefit from this approach—raising professional standards and therefore the attractiveness of the teaching profession, providing current teachers with evidence about their effectiveness, and creating incentives for improved performance among current teachers (see Goldhaber, this issue). However, a very common rationale for policymakers is that using VAMS in employment decisions would improve the composition of the teaching force—weak probationary teachers would not be reemployed or granted tenure, ineffective tenured teachers would be dismissed quickly and easily, and newly hired teachers would perform more effectively than the teachers they replaced.
While proponents of using VAMS to assess teachers’ performance have gained substantial ground over the past decade, there has been no groundswell of support among teachers for its expanded use. Some argue that this is evidence of risk aversion among teachers, while others contend that teachers are as much concerned with the consequences these policies might have for their schools and students as they are about their own job security. In fact, teachers’ doubts and skepticism about the process have grown rather than abated with the implementation of VAMS. In conversations and debates about VAMS, proponents often contend that it is incorrect to give priority to teachers’ interests when students’ futures are at stake. However, opponents suggest that it is important to understand those responses because they may contribute to the kind of “unintended consequences” of using VAMS, which the ASA (2014) predicts might further compromise the quality of schooling.
Some districts had moved quickly to use VAMS to award pay bonuses for highly effective teachers (Mellon, 2008), but it was not until 2010, when The Los Angeles Times published the VAMS scores of individual teachers, that educators across the nation suddenly became alert to the potential personal impact that VAMS might have. As Strauss (2013) reports, other newspapers, including the New York Times (in 2012), and The Cleveland Plain Dealer (in 2013) followed suit. The high-profile court case in California (Vergara et al. vs. State of California et al., 2014) and similar suits in other states such as New York suggest that VAMS remain central to many influential reformers’ agendas (e.g., see LA School Report, 2014). These legal initiatives coupled with what teachers read in the education and popular press about the statistical challenges associated with using VAMS (Eger, 2014; Layton, 2014; Sawchuck, 2014a, 2014b) heighten teachers’ wariness about the role that VAMS might play in decisions affecting them.
Much scholarly attention to VAMS continues to focus on analyzing and improving alternative models for assessing teachers’ contributions to students’ learning. Far less research focuses on how VAMS are used in practice. It seems likely that statisticians and psychometricians will continue to address unresolved matters of measurement. However, it is important to step back and consider a prior set of questions about the use of VAMS: Will assessing and basing employment decisions on the individual teacher’s contribution to students’ learning—however refined and defensible VAMS may become—lead to better schooling? If so, by what process? Is this strategy of augmenting human capital one teacher at a time likely to pay off for students? Or will reliance on VAMS have unintended consequences that interfere with a school’s collective efforts to adopt or sustain promising changes in its instructional program?
In this article I bring an organizational perspective to the prospect of using VAMS to improve teacher quality. I suggest why, in addition to VAMS’ methodological limitations, reformers should be very cautious about relying on VAMS to make decisions that teachers view as important. Because the wide-scale use of VAMS is very recent, scant research exists with which to answer the organizational questions about the intended and unintended consequences of using VAMS in making consequential staffing decisions. However, relevant research about teachers and school improvement, coupled with my own experiences working with states and districts that are implementing new teacher evaluations, lead me to suggest that expanding the use of VAMS in teacher evaluations (even if it represents no more than 30% of the teacher’s total score) might compromise the school’s potential for improvement.
In his classic analysis, “Social Capital in the Creation of Human Capital,” James Coleman (1988) argues that human capital is transformed for the benefit of the organization by social capital, which “inheres in the structure of relations between actors and among actors” (p. S98). This would suggest that whatever human capital schools acquire through hiring can subsequently be developed by interactions among teachers, principals, and others within the organization through activities within subunits such as grade-level or subject-based teams of teachers, faculty committees, professional development, coaching, evaluation, and informal interactions. In the process, the school organization becomes greater than the sum of its parts, and in this way, the social capital that transforms human capital through collegial activities in schools increases the school’s overall instructional capacity and, arguably, its success.
Here, I explore whether expanding the use of VAMS to evaluate individuals is likely to compromise the school’s potential to transform its human capital through systematic reliance on its social capital. Drawing on the metaphor of the egg-crate school, I begin by suggesting that an organizational perspective suggests alternative ways to understand and respond to the variation in teachers’ performance within schools. Thus, I shift the focus from the individual back to the organization, from the teacher to the school. Next, I discuss some of the unintended consequences and organizational costs that a heavy reliance on VAMS may have for schools. Finally, I highlight several alternative strategies for school improvement that rely on social capital to augment the school’s human capital and expand its instructional capacity.
An Organizational Explanation of Variation in Teacher Quality Within Schools
Over 40 years ago, historian David Tyack (1974) and sociologist Dan Lortie (1975) depicted the school as an “egg crate.” Even today the image serves as a compelling metaphor for the familiar compartmentalized school structure, with its classrooms attached in rows along corridors. The egg crate symbolizes both a physical structure and an organizational one—teachers work in isolation, concentrating on their own students largely to the exclusion of others, interacting only intermittently with their colleagues. Tyack, who traced the history and development of the egg-crate school from the one-room schoolhouse to age-graded schools, explains that this segmented organization enabled administrators to respond quickly to enrollment growth by adding one classroom at a time without disturbing the rest of the organization. Similarly, with enrollment declines, adjustments can be made to some grades and classes that do not disrupt others.
However, scholars and practitioners have long argued that this atomized arrangement to school organization does not serve students well or support the needs of teachers as they move through their career (e.g., Darling-Hammond, 2001; Hargreaves & Fullan, 2012; Johnson, 1990; Sizer, 1984). According to these analysts, teachers’ experiences and opportunities for learning are limited because they fail to benefit from the varied models of instruction practiced by their colleagues or to adjust their teaching in response to what students learn or fail to learn in other grades and classes. When schools are organized like egg crates, important information about the challenges that teachers encounter, the problems that puzzle them, and the expertise they might offer their peers remains limited by the confines of the classroom (Hargreaves & Fullan, 2012; Johnson 1990; Kardos & Johnson, 2007; Little, 1990).
Given the prevalence of egg-crate organizational structures in public schools, it should come as no surprise that teachers within a school vary widely in their effectiveness. At any time, a school has a mix of teachers who differ in experience, knowledge, and skill. Compartmentalized school structures limit the potential development of individual teachers, who lack direct access to their colleagues’ expertise. However, social capital theory would suggest that if provided systematic opportunities to engage with their peers outside their classroom, the human capital of individuals—in this case, their instructional effectiveness—could be shared and augmented. Given this line of argument, the more robust the teachers’ instructional repertoire and the more opportunities they have to exchange and integrate promising ideas and techniques into their own teaching, the more likely it will be that all students—not only those assigned to the more effective teachers—will experience the benefits of expert teaching. This analysis suggests that teachers are not inherently effective or ineffective but that their development may be stunted when they work alone, without the benefit of ongoing collegial influence.
Recent studies have persuasively documented the benefits of systematic efforts to improve student learning through schoolwide improvement initiatives (Bryk, Sebring, Allensworth, Easton, & Luppescu, 2010; Little, 1982; McLaughlin & Talbert, 2001; Rosenholtz, 1989). Because students move through schools from class to class or grade to grade, they are better served when human resources are deliberately organized to draw on the strengths of all teachers on behalf of all students, rather than having students subjected to the luck of the draw in their classroom assignment. Successful schoolwide improvement increases norms of shared responsibility among teachers and creates structures and opportunities for learning that promote interdependence—rather than independence—among them. By contrast, a strategy for school improvement that focuses substantially on identifying, assigning, and rewarding or penalizing individual teachers for their effectiveness in raising students’ test scores depends primarily on the strengths of individual teachers.
Teachers’ Improvement Over Time and the Potential of Peer Learning
Recent research about teachers’ improvement suggests how social capital might augment human capital within schools. One of the most well-known, but puzzling, findings of research using VAMS is that on average, teachers’ value-added scores remain flat or decline after five or six years in practice (Rivkin et al., 2005). Evidence is clear that teachers improve during their first few years of work and that they are demonstrably more effective in their second year of teaching than in their first (Boyd, Lankford, Loeb, Rockoff, & Wyckoff, 2008; Rockoff, 2004). Generally, teachers improve for several years thereafter, possibly as a result of early mentoring, students’ responses to their teaching, or their own reflections on their pedagogy. However, studies suggest that on average, teachers’ development levels off and they reach a “plateau” of improvement relatively early in their career (Rivkin et al., 2005). This has led some school officials to consider replacing more experienced teachers whose improvement has stalled with new recruits who are committed to a school’s mission and more likely to work the long hours required to attain it.
Notably, new studies call into question the inevitability of the “plateau.” For example, Ladd and Sorensen (2014) analyzed longitudinal administrative data about middle school teachers in North Carolina and found that they improved “well beyond the first few years of teaching” (p. 24), not only in achieving higher test scores but also improving student behaviors, including the “amount of time spent reading for pleasure, amount of time spent completing homework, number of days absent, and number of reported disruptive classroom offenses” (p. 2). Kraft and Papay (2014) analyzed 10 years of teacher data from Charlotte-Mecklenberg, including both surveys about their school’s work environments and data tracking their students’ test scores over time. The authors found that teachers working in professional environments that the school’s staff had judged to be more supportive improved their effectiveness more over time than those who worked in environments judged to be less supportive. The authors found that over the 10 years studied, a teacher working in a school at the 75th percentile of professional environment ratings increased his or her effectiveness 38% more than a comparable teacher working in a school rated at the 25th percentile.
Therefore, changing the context in which teachers work could have important benefits for students throughout the school, whereas changing individual teachers without changing the context might not (Lohr, 2012). Given that possibility, it is worth learning more about the components of a teacher’s workplace that promote greater satisfaction and more interdependent work. Jackson and Bruegmann’s (2009) large-scale analysis of North Carolina teacher data revealed that student achievement within a grade level improved, both in the short term and over time, when a more effective teacher arrived to teach in the grade. They suggest that this improvement is the result of “peer learning,” although their study does not allow them to explain how peer learning develops. Other work suggests that dedicating time for teachers to meet in teams may be important. Ladd (2011), who analyzed statewide surveys of teachers in North Carolina, found that when elementary and middle school teachers had time allocated in their schedule for collaboration and planning, they were less likely to say that they intended to leave their school. Subsequently, Johnson, Kraft, and Papay (2012) found that the conditions in which teachers work matter a great deal to them and, ultimately to their students, whose student growth scores rose on average when their teacher worked in an environment judged positively by peers within the school in a state-sponsored survey. The factors that mattered most to teachers—the teacher’s colleagues, the school’s leadership, and the school’s organizational culture—were social in nature and therefore might contribute to developing its human capital. Therefore, both theory and empirical evidence suggest that students and their schools stand to benefit when teachers work closely and collaboratively with colleagues.
Possible Unintended Consequences of Increasing Reliance on VAMS
Collaborative professional norms and practices are difficult to develop, and once established, they remain vulnerable to disruptive changes in policy and leadership (Little, 1990). The egg crate continues to be the default organizational structure of most public schools. Therefore, if teachers become dissatisfied with efforts to coordinate and assess their work, they can decide to literally or figuratively “close their classroom door” and revert to working alone. Given that possibility, we next consider how relying on VAMS for a substantial portion of teacher evaluation might negatively affect current collaboration and shared responsibility for school improvement, thus reinforcing the walls of the egg-crate school. In what follows, I consider several unintended consequences that might result from increasing use of VAMS.
Making It More Difficult to Fill High-Need Teaching Assignments
Teachers’ confidence in relying on VAMS for evaluation ultimately depends on whether the methods adequately control for demographic differences among their students. A number of experts report that VAMS do not yet meet this standard (ASA, 2014; Darling-Hammond, Amrein-Beardsley, Haertel, & Rothstein, 2012; Goldhaber, this issue). Although teachers may not have read those methodological critiques, they are familiar with commentary about VAMS, which is seldom reassuring (Eger, 2014; Layton, 2014; Sawchuck, 2014a, 2014b). Moreover, their day-to-day experience often suggests that they may be at greater risk if they agree to teach students whose scores on standardized tests tend to be low, thus leading to a negative evaluation for the teacher. Stories are common among teachers in high-poverty, high-minority schools about respected colleagues whose evaluation ratings dipped because they were assigned to teach students with special needs, those living in poverty, or those still learning English. These teachers generally do not believe that VAMS are, in fact, evenhanded; recent analysis by educational statisticians support those views (McCaffrey, 2013; McCaffrtey & Buzick, 2014; Raudenbush & Jean, 2012). Therefore, heavy reliance on VAMS may lead effective teachers in high-need subjects and schools to seek safer assignments, where they can avoid the risk of low VAMS scores. Meanwhile, some of the most challenging teaching assignments would remain difficult to fill and likely be subject to repeated turnover, bringing steep costs for students.
Discouraging Shared Responsibility for Students
Many schools benefit greatly by encouraging collaboration among teachers, thus reducing their isolation and mitigating its effects on students (Bryk et al., 2010; Bryk & Schneider, 2002). For example, a grade-level team of elementary school teachers might regroup students so that each teacher takes responsibility for a subject she knows well. One teacher who is weak in math but has strength in social studies might teach several sections of social studies while her colleague teaches math. Arguably, all students—rather than just a few—can benefit from such arrangements, receiving the best instruction in each subject that a team of teachers has to offer.
However, using VAMS to determine a substantial part of the teacher’s evaluation or pay threatens to sidetrack the teachers’ collaboration and redirect the effective teacher’s attention to the students on his or her roster. For in order to avoid undeserved VAMS scores, teachers with the strongest overall knowledge and skills might well reclaim their students and conduct all instruction within a self-contained class, leaving less experienced and less effective teachers to cope on their own.
Some districts have tried to circumvent this problem by allocating formal responsibility to individual teachers who instruct each student in each subject, thus pinpointing where a student has learned (or failed to learn) the topics on which he or she is tested. Districts (e.g., Baltimore, Charleston, Chicago, Fort Worth, and Boston) as well as states (e.g., Ohio, Texas, and Washington, DC) have worked with external consultants to identify the portions of a student’s instructional day that each teacher is responsible for (Batelle for Kids, 2015). Ballou and Springer (this issue, pp. 77–86) describe New York State’s plan to use “fractional linkages” to allocate responsibility for students’ learning. However, these approaches can, at best, provide only a rough estimate of an individual teacher’s contribution to what a student knows and can do. A fifth-grade student who successfully solves word problems on his math test may actually owe that success to his fourth-grade ELA teacher who taught him to read carefully, not to his fifth-grade math teacher, whose lessons emphasized memorizing algorithms. Such diffuse effects of instruction, which are part and parcel of the educational process, would not be apparent, even with the most rigorous effort to attribute instructional responsibility to individual teachers. Given that reality, it seems possible that when districts rely on VAMS for a substantial part of teachers’ evaluations, their teachers may pull back from sharing collegial responsibility for the students in a school.
Undermining the Promise of Standards-Based Evaluation
Those who recommend using VAMS for personnel decisions often contend that this approach is superior to the counterfactual—evaluations conducted by administrators. In fact, critics have roundly criticized traditional teacher evaluation for being sporadic, based on simple checklists, and yielding uniformly satisfactory ratings (Donaldson, 2011; Toch & Rothman, 2008). Arguably, the past failures of evaluation are responsible for the continued employment of ineffective tenured teachers, who now are the target of reformers’ dismissal initiatives (Donaldson, 2011; Honawar, 2007; Tucker, 1997; TNTP, 2009; Vergara et al. vs. State of California, et al., 2014).
Over the past 10 years, however, districts have adopted much more sophisticated and informative standards-based assessments that cover a range of pedagogical elements and provide detailed rubrics describing practice at various levels of performance (Danielson, 2013). Since the Race to the Top set new standards for evaluations and the Measures of Effective Teaching study highlighted what can be learned from them if they are well implemented (Bill and Melinda Gates Foundation 2013a, 2013b), districts have spent considerable time and money adopting new instruments and training evaluators to use them validly and reliably. This investment seems warranted, given research findings that teachers’ instruction improves in response to standards-based observations and feedback from evaluators. Taylor and Tyler (2012) studied a comprehensive evaluation process in Cincinnati and found that when evaluators—in this case, peer evaluators—used a standard-based evaluation process to provide informed feedback to the teachers they observed, students’ test scores in those teachers’ classes improved, both initially and over time. Papay (2012, p. 134) argues that unlike VAMS, which provide no information about what a teacher does well or poorly, “standards-based evaluations can provide teachers with a clear ‘line of sight’ between their current practices and what they need to do to improve.”
As Goldring et al. (this issue, pp. 96–104) report, when principals supervise teachers, they tend to give priority to the behavior they observe, rather than to VAMS scores, suggesting that the influence of VAMS on teacher development will be modest. However, all principals do not have the knowledge and skills needed to coach teachers and model exemplary practice. Some have little or no teaching experience, and most others lack experience teaching the full range of grades or subjects they are responsible for supervising. As more states require that teacher evaluations include both classroom observations and data from standardized tests of student achievement, principals who are not instructional experts will be left to interpret discrepancies between what they see in the classroom and what they read on a VAMS score sheet. Will they doubt the validity of their observations or the accuracy of the VAMS score? If they are uncertain about judging instruction or believe VAMS to be more objective and precise than their own professional judgment, value-added scores may unduly influence their decisions.
Generating Dissatisfaction and Turnover Among Teachers
Using VAMS to make high-stakes decisions about teachers also may have the unintended effect of driving skillful and committed teachers away from the schools that need them most and, in the extreme, causing them to leave the profession. Since 2000, researchers have examined the prevalence, causes, and consequences of teacher turnover, especially in high-poverty schools. Ingersoll (2001) analyzed data from the national Schools and Staffing Survey and found that turnover rates were unexpectedly high. The problem, which he dubbed the “revolving door” (p. 11), is especially acute in urban schools serving large proportions of high-poverty, high-minority students. Teachers steadily leave such schools to take jobs in Whiter, higher income communities (Boyd, Grossman, Lankford, Loeb, & Wyckoff, 2006; Boyd, Lankford, Loeb, & Wyckoff, 2005; Hanushek, Kain, & Rivkin, 2004; Johnson and Birkeland, 2003; Leukens, Lyter, Fox, & Chandler, 2004; Marinell & Coca, 2013). Simon and Johnson (in press) recently reviewed the literature on turnover in high-poverty schools and concluded that teachers leave such schools because of poor working conditions, especially those that are social in nature.
Studies show that teachers’ relationships with colleagues matter as they decide whether to stay in their school or transfer (Allensworth, Ponisciak, & Mazzeo, 2009; Guarino, Santibanez, & Daley, 2006; Johnson & Birkeland, 2003; Johnson & The Project on the Next Generation of Teachers, 2004; Kardos & Johnson, 2007). Rosenholtz (1989) finds that teachers want to work in schools where colleagues share a common set of goals, are committed to innovation, share a “‘can-do’” attitude (p. 25), and take responsibility for bettering the larger school. Little (1982) explains that effective schools systematically engage teachers in “frequent, continuous and increasingly concrete and precise talk about teaching practice” (p. 331). Others report that teachers seek an inclusive work environment where the adults respect and trust one another and the students they serve (Achinstein & Ogawa, 2011; Bryk & Schneider, 2002).
When effective teachers become dissatisfied with their work environment, they often transfer or leave teaching, and as a result, their schools and students pay a price. Chronic turnover exacts instructional, financial, and organizational costs that can destabilize learning communities within schools and compromise student learning (Allensworth et al., 2009; Balu, Beteille, & Loeb, 2010; Guin, 2004). Recent research by Ronfeldt, Loeb, and Wyckoff (2013) documents the impact of turnover on students’ learning and finds that its negative effects are greater for low-performing and Black students than for their higher performing, non-Black peers. Thus, the very schools that most need effective teachers have the greatest difficulty attracting and retaining them.
Those who promote reliance on VAMS to make employment decisions about teachers often suggest that the best teachers will be more satisfied and remain at their school once ineffective teachers have been dismissed. However, if the dismissal process requires more testing or diverts teachers from collaborative improvement efforts, skilled teachers, who arguably have the most to offer the school, may lose confidence in administrators’ priorities and decide to go elsewhere, even if that takes them out of education. In line with Coleman’s (1988) theory about social capital, this suggests that a school would do better to invest in promoting collaboration, learning, and professional accountability among teachers and administrators than to rely on VAMS scores in an effort to reward or penalize a relatively small number of teachers.
Discussion
There is reason therefore for policymakers and administrators to carefully weigh the potential costs and benefits of relying on VAMS in evaluations of teachers’ performance—whether such measures constitute 30% or 50% of the final evaluation. It is possible that such policies will have their intended effects—raising professional standards and making teaching more attractive, reducing the variability in teachers’ effectiveness through dismissals and resignations, and motivating current teachers to use instructional strategies that are thought to improve students’ performance. However, it is also possible that such reliance on VAMS will make it more difficult to staff high-need classes, promote and sustain collaborative work among teachers, and develop shared responsibility among all teachers for school improvement and students’ learning. In response to these effects, turnover rates may increase, even among the very teachers whose expert skills and commitment could generate improvement among their colleagues.
This analysis is not meant to suggest that schools should continue to employ ineffective teachers. As I have argued elsewhere (Johnson, 2012), “neither individual teachers nor the schools in which they work can be ignored if students are to have the instruction they deserve” (p. 119). However, reformers should lead the way with efforts to improve the school throughout as an organization that supports effective teaching and rich learning. Therefore, it is important to ask what promising activities and programs might support and develop teachers, thus increasing the human capital within schools through deliberate reliance on social capital. What changes could ensure both accountability and development for teachers while continuing to ease, rather than reinforce, classroom boundaries?
One approach is to engage teachers, themselves, as leaders in a broad range of processes that ensure the selection, support, and assessment of skilled and committed colleagues. The first step of developing human capital is the process of recruitment and hiring. Research suggests that far too many teachers are hired solely on the basis of paper credentials or a single interview with the principal held late in the year (Cannata, 2010; Levin & Quinn, 2003; Liu & Johnson, 2006). Engaging teams of teachers and administrators to observe candidates’ instruction, either by video or in person, can ensure that newly appointed teachers have the potential to succeed. Having candidates debrief their demonstration lesson with prospective colleagues not only can lead to better hiring decisions but also ensure that the candidate has a good preview of the school’s expectations (Liu & Johnson, 2006) and that current teachers will accept responsibility to support their new colleague.
Once new teachers join a school, they can benefit from the support of an induction process that includes peer mentoring (Feiman-Nemser, 2012; Moir, Barlin, Gless, & Miles, 2009) or rapidly engages them in regular interactions with more experienced colleagues in an “integrated professional culture” (Kardos, Johnson, Peske, Kauffman, & Liu, 2001). In addition, schools can provide time for newly hired teachers to observe their colleagues as part of a schoolwide initiative to share feedback and best practices (Hamilton, 2013), including highly organized processes, such as lesson study (Fernandez, 2005) and instructional rounds (City, Elmore, Fiarman, & Teitel, 2009).
Ongoing supervision and standards-based evaluation by administrators and peer evaluators can encourage and support further development among teachers. Experienced teachers distinguished by their instructional expertise and interest in leadership may serve as both mentors and evaluators. For example, peer assistance and review (PAR) programs, developed jointly by labor and management in districts such as Montgomery County, Maryland; San Juan, California; St. Louis, Missouri; and Cleveland, Ohio, select expert consulting teachers to support and eventually assess all new teachers as well as tenured teachers whose performance falls below standards (Johnson, Papay, Fiarman, Munger, & Qazilbash, 2010). After several months of supervision and assessment using a standards-based evaluation instrument, the consulting teachers and the joint PAR panel to whom they report recommend whether teachers in the program should be reemployed or dismissed. Yusko and Feiman-Nemser (2008) studied PAR in Cincinnati and found that teachers there valued having a consulting teacher both to support their improvement and assess their performance. Because the PAR process is closely monitored to ensure that teachers receive due process, appeals and arbitrations are rare. The mentoring provided by consulting teachers leads to higher rates of retention among new teachers, while the formal evaluations lead to higher rates of dismissal and voluntary resignations among tenured teachers (Humphrey, Koppich, Bland, & Bosetti, 2011; Papay & Johnson, 2012).
New Haven’s Teacher Evaluation and Development System (TEVAL) incorporates multiple factors in teachers’ evaluation ratings based on standards-based observations, success in achieving goals for student achievement, and principals’ assessments of teachers’ professional practice (Donaldson, 2014; Donaldson & Papay, 2014; New Haven, CT Public Schools, 2014). Teachers consult with their evaluator in setting goals for improved student learning. By November, evaluators identify teachers who have low scores on observations and student achievement, putting them on notice that they risk dismissal. At the same time, evaluators identify teachers who demonstrate success on all components and are on track to receive exemplary ratings. An evaluator from outside the district then observes the instruction of teachers in both groups. If the external evaluator confirms a teacher’s ineffective rating, that individual is subject to dismissal; if the external evaluator confirms a teacher’s exemplary rating, that teacher is eligible for leadership roles, including becoming an evaluator. Over the first three years of TEVAL’s use, teachers increased their focus on student achievement (Donaldson, 2013), and the district reported more exits of teachers due to poor performance (Donaldson, 2014). These examples of hiring, induction, supervision, and evaluation practices illustrate how districts and schools can use organizational approaches and the social capital that exists within a school to develop its human capital. VAMS’ proponents might contend that the ratings determined by a statistical analysis of student achievement data would achieve the same outcomes more efficiently. However, the organizational benefits of programs such as school-based hiring or PAR extend well beyond a collection of discrete decisions to employ or dismiss individual teachers. They contribute to developing consensus among teachers and administrators about instructional goals, broad understanding of how the curriculum works in practice, and shared responsibility for the school’s students.
Thus far, research has focused almost exclusively on the technical side of VAMS, determining under what conditions these models can safely and sensibly be used. This work has been enormously important for cautioning and guiding policymakers and school officials about what can and cannot be learned from ratings based on VAMS. However, given some reformers’ and policymakers’ intentions to expand the use of VAMS along with teachers’ continuing apprehension about what VAMS might bode for them as individuals, it is time for other researchers to focus on how the use of VAMS in fact affects teachers’ attitudes and responses or what intended or unintended effects it may have for schools. We do not know how using VAMS as the measure of student achievement in the new generation of standards-based evaluations will influence administrators’ summative assessments, especially in states such as New York or Oklahoma, that require their use to determine a substantial part of teachers’ evaluations. It is important to know much more about how administrators use VAMS, not only to identify weak teachers but also to select and support those who seem to be effective. Will principals provide more or less attention to the many teachers whose performance is “good enough” but who could become much more effective with support and guidance?
Also, we need to understand how intensified use of VAMS advances or sets back progress on a broader school reform agenda. This calls for a set of comparative case studies conducted in a variety of policy contexts. Qualitative and quantitative methods can be used to study whether and how the use of VAMS interacts with initiatives to increase shared responsibility and collaboration among teachers in a school. Such studies might compare districts that rely on VAMS scores with those using alternative approaches to evaluation, including PAR or programs similar to TEVAL that engage teachers in setting goals for student performance in consultation with evaluators. Importantly, what effect, if any, do various approaches to assessment have on teachers’ readiness to share what they know and learn from their peers?
There is as yet no evidence that the intensified use of VAMS interferes with collaborative, reciprocal work among teachers and principals or sets back efforts to move beyond the traditional egg-crate structure. However, the fact that we lack evidence about the organizational consequences of using VAMS does not mean that such consequences do not exist. Convincing research is available about several relevant topics, including the importance of the work environment in retaining effective teachers (Johnson et al., 2012; Ladd, 2011), the key role of trust among teachers and principals in efforts to improve school (Bryk & Schneider, 2002), and the potential of peer learning both to improve students’ learning (Jackson & Bruegmann, 2009) and to accelerate teachers’ development over many years of their career (Kraft & Papay, 2014). Until more extensive studies are available about the use of VAMS in schools, policymakers and school officials would do well to rely on approaches that have been shown to improve schools as organizations. Ignoring the unintended consequences of using VAMs to make employment decisions may set back hard-earned, but still fragile, progress.
