Abstract
This article provides a brief history of K–12 education testing in the United States from colonial America to the present. In early America, students were examined orally. After the mid-nineteenth century, written tests replaced oral presentations. In the late nineteenth century, graded schools gradually replaced the single-teacher, one-room schools. In the beginning of the twentieth century, standardized intelligence tests were increasingly used to categorize and promote students. State departments of education have played a larger role in local school funding and policies in the past hundred years. Since the 1960s, the federal government has expanded its involvement in national education while also promoting the role of states. During the past three decades, the federal government and states increased the use of high-stakes national testing with initiatives such as America 2000, Goals 2000, No Child Left Behind, and Every Student Succeeds.
In the United States as of this writing, the public and policy-makers are concerned about the quality of our public schools. Facing increasing global education and economic competition, we want our schools to be among the best in the world. As a result, in the last three decades the federal government and the states have cooperated in developing national initiatives such as America 2000, Goals 2000, No Child Left Behind, and Every Student Succeeds. All these programs emphasize creating national or state education standards and high-stakes, test-based accountability. Indeed, many educators believe that our students face more high-stakes tests than in any other country (Koretz 2017).
How did this emphasis on high-stakes tests develop, and how has it affected U.S. society? There is an extensive literature on high-stakes testing, but most scholars lack an understanding of the history of education and of the emergence of educational testing over the past several hundred years. During debates today about who should decide how students are educated and evaluated, for example, appeals are frequently made to the past to justify or oppose current practices (Vinovskis 2015b). Scholars should understand school testing and the part that parents, local schools, states, and the federal government have played in its development. This article briefly discusses these issues in three time periods: colonial and nineteenth-century America, 1900 to 1960, and 1960 to 2016.
Colonial and Nineteenth-Century Education
Many Europeans settling colonial America in the early seventeenth century were Protestants who emphasized that everyone should be able to read the Bible. Parents were responsible for teaching their own children basic literacy, though sometimes they employed women who taught rudimentary literacy or hired itinerant school teachers. At times, local communities in New England admonished families whose children could not read. A few grammar schools were created in the mid-seventeenth century, which offered advanced subjects such as Latin and Greek for students preparing to enter Harvard University (Moran and Vinovskis 1992).
After the American Revolution, common (primary) schools rapidly expanded; and even some high schools appeared before the Civil War. The North and the Midwest rapidly expanded their common schools, while the South struggled to provide such facilities. Except for the larger communities, public and private schools usually were one-room “little red schoolhouses.” In the summer months, unmarried female teachers instructed younger students. During the winter terms, older children attended schools taught by male teachers (Kaestle 1983; Zimmerman 2009).
Before the Civil War, free African American children attended segregated Northern schools or learned to read elsewhere (such as in Sunday schools). In the South, the great majority of African Americans were slaves whose masters opposed teaching them reading or writing. Following the Civil War, however, there was a large increase in African Americans who attended segregated and underfunded Southern primary schools (Anderson 1988).
Outside the few larger cities, local boys and girls of all ages attended common schools that taught basic subjects such as reading, writing, and arithmetic. School committees frequently hired poorly prepared, but less expensive, teachers. Education reformers advocated employing better trained teachers. But education remained a parental and community responsibility; local school committees and parents, rather than outside educators, decided where children attended school, what teachers were hired, and what books were used (Kaestle and Vinovskis 1980).
Boston established the first public male high school in 1821. Six years later, Massachusetts required towns with at least five hundred families to create public high schools to teach American history, bookkeeping, geometry, surveying, and algebra. High schools grew more rapidly in the urbanized North than in the more rural South. Large cities and some smaller towns maintained at least one central high school (Reese 1999; Vinovskis 1995).
In the mid-nineteenth century, a major change occurred in Boston when rehearsed oral examinations were replaced by unannounced written questions (tests). This practice spread rapidly to other schools in larger cities and affected how students prepared for tests as well as what questions they were asked (U.S. Congress 1992, 107–10). Prior to the mid-1800s, the end of the common school year was celebrated with a public exhibition during which students orally demonstrated their individual achievements. Younger pupils recited the ABCs while older ones memorized and recited famous poems or orations. Similarly, high school students, like those in common schools, displayed orally their recently acquired skills. Parents and the public enjoyed these well-rehearsed oral displays and accepted them as evidence of student accomplishments (Reese 2013).
In the early 1840s, the Boston English grammar school masters and the secretary of the Massachusetts Board of Education, Horace Mann, became embroiled in a debate over the need to reform schools. The masters claimed that their students were doing excellent academic work and cited the Boston School Committee’s favorable comments based on oral examinations. On the other hand, Mann argued that Boston grammar schools were poorly run and students inadequately trained (Reese 2013).
By the summer of 1845, the school committee substituted written tests for the annual oral presentations. The written test results confirmed the poor performance of the school students. Boston’s experiences encouraged other large cities to use written examinations.
In 1800, exhibitions, recitations, and visual displays of learning were the central means to assess a school. A century later, these seemingly timeless practices had hardly disappeared, but they had lost much of their legitimacy. Except in the most backwoods districts, no one confused an exhibition with an actual examination. (Reese 2013, 158)
Some communities initially switched entirely to using only written examinations and publicly displayed the students’ scores. This upset many parents whose children did not do as well as the others. As a result, written examinations remained in use, but they were combined with what were seen to be more “positive” classroom teacher observations. Moreover, educators no longer publicized the rank ordering of the schools, teachers, or individual students by their written test scores (White 1891). Yet the shift from oral examinations to written ones was significant because it later made it possible to standardize testing not only in individual classrooms and schools, but also in states and the nation.
Another development in the mid-nineteenth century that played an important role in the subsequent K–12 schooling and testing was the movement toward graded schools. Most children of all ages attended local district schools. As school districts became more densely populated, it also became necessary to develop multiroom, multiteacher schools. At first, these new institutions were only roughly graded by student competency. But in 1847, John D. Philbrick introduced a system of “graded instruction.” Subjects were standardized and taught in order of increasing difficulty. Teachers became more specialized in what they taught. This also allowed teachers to work with more students on the same lesson and allowed schools to hire narrowly trained instructors (Angus, Mirel, and Vinovskis 1988; Tyack 1974).
By the 1870s, graded classrooms were used in many of the larger communities, but how students should be categorized and advanced in their classes remained a problem. Some schools left promotions up to teachers; others relied on principals or school superintendents. To determine a student’s knowledge for promotion to the next level, a few schools introduced special examinations (Shearer 1898). The use of graded schools gradually replaced the one-room public schools in the next century, as efforts increased to standardize the curriculum and testing.
Changes from 1900 to 1960: Scientific Approaches and Equity Concerns Influence State-Based Testing Regimes
Major changes in American society and public schooling during the first six decades of the twentieth century made it easier to create graded schools than it had been in the nineteenth century. First, the population of the United States grew from 76 million in 1900 to 179 million by 1960. Second, the proportion of those living in the communities with a population greater than 25,000 increased from 26 percent in 1900 to nearly 45 percent by 1960 (Carter et al. 2006a, 104). Third, the number of public school districts were consolidated from 127,531 in 1932 to 40,520 by 1960. As a result, the number of one-teacher public schools dropped dramatically from 200,100 in 1916 to 20,213 in 1960 (Carter et al. 2006b, 398).
When school districts were smaller, neighbors and parents usually had closer contact with school committee members and their local schools. School committee members in large cities were elected by wards, which often led to corruption and abuses in teacher hiring practices. Progressives succeeded in creating city-wide school committees, with members elected at large. Consequently, professionals or middle-class reformers were more likely to be elected, and they favored centralizing school operations as well as standardizing the curriculum and testing (Callahan 1962; Steffes 2012).
Prior to the twentieth century, the federal government occasionally provided limited assistance to elementary and secondary schools through special legislation, such as the Northwest Ordinances of 1785 and 1787 (Kaestle 1983). A major federal initiative for public schools was the Smith-Hughes Vocational Education Act of 1918. This legislation provided federal support for vocational education teachers and for hiring additional State Department of Education staff (Steffes 2012).
At the same time, states created stronger departments of education, and governors became more involved with school issues. In 1900, 82 percent of revenue for public elementary and secondary schools came from local sources, and 18 percent from the states. By 1961, 57 percent came from local communities, 39 percent from the state, and 4 percent from the federal government. After World War I, states increasingly became more directly involved in how their public schools were managed. Teacher professionalization, stronger teacher unions, and more powerful city school superintendents made it easier to standardize school practices at both the state and city levels (Steffes 2012; Vinovskis 2008).
The growing influence of the business community in the early twentieth century also impacted education. This growing influence coincided with the introduction of scientific management in business, including its reliance on efficiency experts and measuring worker productivity. As the overall costs of public schooling increased, the business community argued that educators could teach students better and with less money by introducing scientific management practices into teacher practices and improving teacher efficiency. Many public schools eventually implemented these strategies and this had considerable impact on how students were taught and tested (Callahan 1962).
Empirical and scientific approaches to education were introduced in the 1890s to improve teaching and measure student achievement. Joseph Mayer Rice standardized student tests in arithmetic, English composition, penmanship, and spelling. He also analyzed how different teaching methods affected student achievements (Ayers 1918). And Edward Thorndike developed rigorous education examinations, including one of the first standardized achievement tests in 1908 (Linden and Linden 1968). By 1917, there were at least eighty-four standardized tests for elementary schools and twenty-five for high schools (Monroe 1918).
Alfred Binet’s pioneering work in Europe measuring children’s mental ability stimulated intelligence testing in the United States. Before the introduction of Binet’s ideas here, there had been only limited interest in intelligence testing. As John Carson explained, Lewis Terman’s development of the Stanford-Binet intelligence scales for Americans “marked a fundamental divide in the American history of intelligence. After 1916 the equation of intelligence with IQ, understood as innate, quantifiable mental ability, gradually became accepted within parts of the psychological community and the broader culture as well” (Carson 2007, 183). During World War I, the U.S. military recruited psychologists to develop and administer intelligence tests to nearly two million soldiers (Resnick 1980). After the First World War, the United States increased its development of mental tests for schools. According to Paul Chapman (1988, 170), “The use of intelligence tests flourished in the 1920s. The use of tests began in urban areas and focused initially on elementary schools. Quickly testing expanded into junior high schools and high schools and reached into the rural parts of the country. … As the movement gained support, intelligence testing became common practice in American schools.”
Post–World War I, school enrollments rose rapidly and student diversity increased due to European immigration. Elementary students increasingly continued into secondary school. Education also became more complex as special schools multiplied, regular public schools offered more curricular options, and students prepared for a greater variety of careers. This encouraged educators to provide alternatives, such as more vocational education, and to track students within regular schools. Education tests were used to help educators and parents guide students into vocational or academic programs. As a result, it became more important how schools assessed and categorized students (Resnick 1980).
Many testing experts and educators claimed that intelligence tests provided a more rational system for grouping students than relying on the judgments of teachers (Brooks 1922; Whipple 1922). Psychologists and many school personnel welcomed the development and use of intelligence tests; but others questioned the specific characterizations of intelligence as well as using a child’s IQ by itself to determine grade and academic track placement (Chapman 1988). Over time, most educators and testing experts acknowledged that intelligence tests measured both an individual’s innate abilities as well as their environmental background (Linden and Linden 1968).
Students were increasingly placed into graded schools and promoted annually from one grade to the next. Most students were advanced, but a substantial number were held back due to inadequate academic preparation. Early testing experts, such as Leonard P. Ayres, complained that the curriculum was designed mainly for the brightest students and recommended compulsory school attendance, a more flexible grading system, and courses of study suited to the average student. In the 1930s and 1940s, frustrated by the difficulty of reducing the number of students repeating the same grade or dropping out, educators gradually accepted more lenient promotions to the next grade. Social promotion opponents, however, worried about possible grade inflation as well as disadvantaged students receiving a second-class education (Angus, Mirel, and Vinovskis 1988).
Although the Great Depression of the 1930s created severe economic problems for schools, it generally did not alter how education and testing was provided. States at this point played a larger role in financing schools, but the federal government did not intervene much through its New Deal programs (Tyack, Lowe, and Hansot 1984). During World War II, education was acknowledged as a national concern that needed a more equitable distribution of local resources. Some also called for more international understanding and cooperation in education to improve democracy at home and abroad (Kandel 1948; U.S. Congress 1992, 127). And while World War II disrupted high school attendance for some students, the GI Bill in 1944 provided federal support for veterans afterward (Frydl 2009).
After World War II, the courts increasingly played a larger role in education. For example, the 1954 Supreme Court decision in Brown v. Board of Education declared segregation in public schools unconstitutional and called on federal officials and governors to eliminate racial discrimination in schools. In the ongoing debates about desegregation, the results of student tests often were used to document the disadvantages that African American students faced (Kaestle 2016; Patterson 2001). And three years later the Soviet launch of Sputnik raised questions about the academic quality of American education and led to the National Defense Education Act of 1958, which provided funds for foreign languages, mathematics, and science instruction at all levels of schooling, as well as federal research funding and assistance for state statistics (Ravitch 1983; Urban 2010).
Changes from 1960 to 2016: Ascendancy of Federal Involvement with Standards and Testing
The U.S. population almost doubled from 1960 to 2016 and increasingly lived in urban areas. The foreign-born population grew to 14 percent by 2016, with most immigrants coming from Asia or Latin America rather than Europe. As of this writing, 61 percent of the population is non-Hispanic white, 18 percent Hispanic, and 13 percent black or African American (Vespa, Armstrong, and Medina 2018). Many families now are more affluent than in 1945, but about one out of five children under age 18 in 2015 lived in households below the poverty level (Snyder, de Brey, and Dillow 2018, 32; Vinovskis 2011).
Pre-kindergarten through grade eight students increased from 32 million in 1959 to 39 million in 2014. The number of youth in secondary school rose even faster from 9 million to 16 million. Although more high school students graduated in 2015 than in 1960, many still do not (Snyder, de Brey, and Dillow 2018, 60, 224).
Spending for K–12 education (in constant 2016 dollars) grew substantially from $138 billion in 1959 (3.2 percent of GDP) to $676 billion (4.1 percent of GDP) in 2013. During that same period, the federal government and the states increased their share of support for K–12 schooling. In 1959, 4 percent of K–12 revenues were from the federal government, 39 percent from the states, and 57 percent from local sources. By 2013, federal revenues made up 9 percent, the state share increased to 46 percent, and local contributions dropped to 46 percent (Snyder, de Brey, and Dillow 2018, 63, 82). The average cost per pupil in public elementary and secondary schools more than tripled from $3,760 in 1959 to $12,732 in 2013. And the average annual salaries for elementary and secondary public school teachers increased by 41 percent, from $41,371 in 1959 to $58,432 in 2013 (Snyder, de Brey, and Dillow 2018, 168, 390).
Policies in the 1960s were especially concerned with eliminating poverty and addressing civil rights issues. This period also witnessed increased federal involvement in preschools as well as in elementary and secondary education. This included additional federal school funding, more regulations, development of new tests, and increased research support. The federal government worked closely with states, which helped to lessen opposition to these changes. Proponents of more federal and state involvement turned to the courts to expand fundamental education rights. At the same time, conservative and local opposition to federal and state involvement in education continued (Reed 2014; Vinovskis 2011).
The federal government also became involved in developing student education measurements. The collaboration in the early 1960s between U.S. Education Commissioner Francis Keppel and Ralph W. Tyler, a leading test specialist, was particularly important. Funded by the Carnegie Corporation, Tyler chaired a committee that drafted a national achievement test based on student samples (Finder 2004). By 1972, the federal government assumed full funding of what became the National Assessment of Education Progress (NAEP). Under pressure from opponents of state-level assessment comparisons, the committee agreed that only regional results would be released. A prominent NAEP study group in 1987, however, recommended the release of state-level NAEP data. As a result, the Hawkins-Stafford Elementary and Secondary School Amendments of 1988 allowed the collection of state-level mathematics and reading assessments on a trial basis. Thereafter, more NAEP data were collected and released at the state level (Jones and Olkin 2004; Pellegrino, Jones, and Mitchell 1999; Vinovskis 2001).
In 1964 Congress passed the Economic Opportunity Act as part of President Lyndon B. Johnson’s War on Poverty. Under this legislation, the federal government launched an eight-week summer Head Start preschool program in 1965 to prepare half a million disadvantaged children to enter public school as kindergarteners. Gradually Head Start became a year-round program, mainly operated by local community action agencies. Head Start was one of the most popular federal education programs of its time, but critics questioned its long-term effectiveness and complained about the lack of federal or state oversight of the quality of its projects. After disappointing results on IQ or cognitive tests, some early childhood education advocates and proponents of the program instead emphasized that Head Start students were less likely to be held back in regular schools or arrested later in life. Unfortunately, these two measures were not good indicators of cognitive improvement, which was the original goal of Head Start policy-makers (Vinovskis 2005).
The other major federal program for disadvantaged students during this time period came in 1965 as the Title I of the Elementary and Secondary Education Act (ESEA). Modest federal education monies were provided to almost every congressional district without much guidance or oversight for how the funds were to be used. Senator Robert Kennedy (D-MA), however, insisted that ESEA programs be evaluated for their effectiveness. Unfortunately, those assessments were not routinely or rigorously carried out (Kaestle 2016; Vinovskis 2005).
Besides providing funding for disadvantaged students, ESEA also expanded the state role in education substantially. U.S. Education Commissioner Francis Keppel recommended that states distribute and monitor the Title I funds. As a result, the federal government provided money to state education departments to increase their staffs as well as legitimize state involvement in local education (Smith 1967; Vinovskis 2008). “Despite Washington’s greatly enlarged role, perhaps the most striking change in U.S. education in the last forty years has been the growth of centralized state control and the ascendance of governors over school policy in most states. Organizations of local administrators, teachers, and school board members dominated state policy agendas no longer” (Kirst 2004, 28).
During the 1970s, the White House, Congress, and the courts continued to shape federal involvement in K–12 education. President Richard M. Nixon discouraged busing students to desegregate schools. He also encouraged coordination of federal and state education programs. And Nixon unsuccessfully tried to reduce the role of community action agencies in Head Start projects (Vinovskis 2011). In the 1976 presidential election, both the American Federation of Teachers (AFT) and the National Education Association (NEA) for the first time endorsed a presidential candidate, Jimmy Carter. During the campaign Carter promised to establish a cabinet-level department of education. The creation of the Department of Education in 1979 and the growing involvement of the AFT and NEA in politics led to further partisan disagreements on education. Although President Ronald Reagan tried to abolish the Department of Education, he did not succeed (Radin and Chanin 2008).
During the 1970s and 1980s, concerns about low economic productivity and the growing belief that elementary and secondary public schools needed substantial improvements persuaded many southern governors to take the lead in education reforms. Before 1960, low-stakes testing was a normal part of elementary and secondary education. However, in the 1970s, thirty-three states instituted minimum competency testing, including eighteen states that required high school students to pass tests before graduating. Many policy-makers and educators, using minimum-competence test scores, now held students, teachers, schools, and states accountable for their education achievements. Evaluations also increasingly relied on criterion-referenced tests, which reported results on how well students had mastered subjects or skills (rather than just comparing students to each other). At first, some of these minimum competency tests were rigorous; but states quickly reduced their difficulty as a substantial number of students failed the initial examinations (Resnick 1980). As Daniel Koretz observed, The shift from using tests for information to holding students or educators directly accountable for scores is beyond a doubt the single most important change in testing in the past half century. Test-based accountability has taken varying forms … but the basic principle of shaping educational practice by means of accountability to test scores has grown only more central to educational policy in the United States (and in many other nations as well). It is not an exaggeration to say that it is now the cornerstone of American education policy. (Koretz 2008, 57–58)
During the Reagan administration, the U.S. Department of Education issued its widely publicized A Nation at Risk, which claimed that then-recent declines in student academic test scores revealed public school shortcomings (Vinovskis 2009a). The governors and the National Governors Association (NGA) had already been using state-wide tests to stimulate education reforms, as well as demonstrate statewide student academic achievements in the 1970s and 1980s. The NGA issued its widely circulated report, Time for Results: The Governors’ 1991 Report on Education, which called for state-level goals and better reporting of the results (Vinovskis 1999). Tennessee Governor Lamar Alexander called for “some old-fashioned horse-trading. We’ll regulate less, if schools and school districts will produce better results” (NGA 1986, 3).
In the 1988 presidential campaign, George H.W. Bush differentiated himself from Reagan by acknowledging a federal role in education and calling himself the “education president.” During the campaign, neither presidential candidate brought up the idea of national education goals. Nor did President Bush initially emphasize the need to create the national education goals; but many governors still called for developing state goals, and some also supported national education goals. Arkansas Governor Bill Clinton, with bipartisan NGA support, urged President Bush to meet with the governors at the 1989 Charlottesville Education Summit (Vinovskis 1999).
In preparation for the summit, Bush assigned the drafting of the national goals to two separate offices within the Department of Education. At the same time, Governor Clinton and his colleagues came up with a similar list. At the September 1989 summit, both the administration and the governors agreed on the need for national education goals but cautioned that the objective was to create “an ambitious, realistic set of performance goals. … National goals will allow us to plan effectively, to set priorities, and to establish clear lines of accountability and authority” (Vinovskis, 1999, 40). When the final goals were announced in the next year, the president and the governors reiterated that “as elected chief executives, we expect to be held accountable for progress in meeting the new national goals, and we expect to hold others accountable as well” (Vinovskis 1999, 40).
The six goals that were agreed upon in 1990 were much more ambitious than Bush’s initial version, and the president and the NGA promised to reach them by 2000 (Vinovskis 1999). Especially challenging were goals one, three, and four:
Goal 1: By the year 2000, all children in America will start school ready to learn.
Goal 3: By the year 2000, American students will leave grades four, eight, and twelve having demonstrated competency in challenging subject matter including English, mathematics, science, history, and geography.
Goal 4: By the year 2000, U.S. students will be first in the world in science and mathematics achievement. (Executive Office of the President 1990)
The national education goals were extraordinarily ambitious and unrealistic. Yet they continued to be used (and even further expanded) in America 2000, Goals 2000, and No Child Left Behind (NCLB). Policy-makers pledged themselves to be held responsible for reaching the goals; but when America failed to reach any of the goals by 2000, almost no one acknowledged their earlier promise.
These goals eventually forced policy-makers and teachers either to ignore the legislation or to find ways around it (such as encouraging grade inflation or lowering the state proficiency standards). But in the meantime, the goals led to immediate objectives that often created difficulties and anxiety for many students, parents, teachers, and policy-makers. While policy-makers and educators in the early 1990s mainly focused on improving student testing methods and outcomes, they had paid much less attention to considering the consequences of setting such unrealistic and unreachable objectives.
The Bush administration proposed the America 2000 program, which stressed the need for challenging national academic standards and voluntary national tests, but partisanship prevented its congressional passage. Bush, however, used his executive powers to implement portions of America 2000 and accept voluntary state and local involvement. Altogether, forty-four states and 2,300 communities adopted the six national goals (Vinovskis 2009a).
President Clinton incorporated many of America 2000 ideas when Goals 2000 was passed. Both America 2000 and Goals 2000 focused attention on education and provided additional federal involvement and resources. And partisan divisions over the proposed opportunity-to-learn standards (resource inputs needed to reach the goals) undermined some of the other more cooperative bipartisan achievements. By the mid-1990s, the Clinton administration and others quietly downplayed Goals 2000 and instead concentrated on specific subjects such as reading, math, and science, along with a few other small-scale federal education initiatives (Vinovskis 2009a).
The Clinton administration realized that the national goals by themselves were not enough to significantly improve American schools by 2000. As a result, they also embraced systemic reform, which meant working closely with states. Several different definitions of systemic reform have been used by policy-makers, but there was never consensus on what the term meant or the likelihood of implementing it nationwide. The Clinton administration’s approach was based largely on the work of the incoming education undersecretary Marshall “Mike” Smith and his colleague Jennifer O’Day (Fuhrman 1993; Ravitch 1995). Smith and O’Day called for statewide curriculum frameworks aligned with rigorous assessments and additional teacher training. While the idea of systemic reform was innovative and plausible, it was largely untested at that time (Vinovskis 1996).
With the election of President George W. Bush in 2000, much of the basic framework of the previous America 2000 and Goals 2000 continued, but now many Republicans and some Democrats wanted stronger standards and to hold states and teachers even more accountable in practice under the proposed NCLB. Democrats and Republicans compromised, and the revised NCLB passed with nearly 90 percent support on the final votes in the House and Senate (McGuinn 2006; Vinovskis 2009a). Later evaluations of NCLB showed only limited educational improvements. The National Research Council, for example, “focused on 15 test-based incentive programs, including the large scale policies of NCLB, its predecessors, and state high school exit exams … [and concluded that] test-based incentive programs … have not increased student achievement enough to bring the United States close to levels of the highest achieving countries” (National Research Council 2011, 4).
Following the limited success of both Goals 2000 and NCLB, many analysts thought that President Barak Obama would abandon the NCLB approach. Obama did turn to other reform strategies, such as Race to the Top, but he also endorsed NCLB. Again, there was only limited progress toward the national goals during the Obama years. And opposition continued from many Democrats and Republicans about excessive testing and the likelihood of penalties for the increasing number of schools defined as failing under NCLB assessment standards. Finally, in 2015 with bipartisan support, the Every Student Succeeds Act (ESSA) was passed. While ESSA maintained aspects of the earlier initiatives, it returned control of education policy to the states (Duncan 2018; Hess and Eden 2017; Maranto, McShane, and Rhinesmith 2016).
While some education experts had at first welcomed test-based accountability, others had reservations about how such tests would be created and used to improve American education (Hanushek and Jorgenson 1996; Heubert and Hauser 1999; Ryan and Shepard 2008). Over time, the methodology of test-based accountability and the disappointing results of America 2000, Goals 2000, and NCLB became evident. Consequently, there are growing complaints about the use of excessively narrow tests as well as the damages caused by employing such high-stakes assessments to schools, teachers, and students (Hout and Elliott 2011; Madaus, Russell, and Higgins 2009; Ravitch 2016).
Daniel Koretz, for example, has supported using appropriate, limited education testing; but he rejected the misuse of narrow, high-stakes tests and warned about the lack of attention to other aspects of schooling not addressed by these examinations. In his sobering summary of the costs and benefits of the recent reforms from 1992 to 2015, he stated, It is no exaggeration to say that the costs of test-based accountability have been huge. Instruction has been corrupted on a broad scale. Large amounts of instructional time are now siphoned off into test-prep activities that at best waste time and at worst defraud students and their parents. Cheating has become widespread. The public has been deceived into thinking that achievement has dramatically improved and that achievement gaps have narrowed. Many students are subjected to severe stress, not only during testing but also for long periods leading up to it. Educators have been evaluated in misleading and in some cases utterly absurd ways. Careers have been disrupted and in some ended. Educators have been indicted and even imprisoned. The primary benefit we received in return for all of this was substantial gains in elementary-school math that don’t persist until graduation. This is true despite the many variants of test-based accountability the reformers have tried, and there is nothing on the horizon now that suggests that the net effects will be better in the future. (Koretz 2017, 191)
Conclusion
The history of education testing in America that has been laid out in this article shows considerable changes and improvements. But when applied today in high-stakes K–12 testing situations, the tests do not always yield good results. In addition, there are some other test-related issues that can be improved, in areas such as education research, early childhood education, civics/history education, and bipartisan cooperation on education issues.
Both basic and applied education research should be expanded and improved. While there is some high-quality education research, more can be done to develop rigorous and relevant research on school improvement and student learning (Vinovskis 2015a). Particularly useful would be studies of how policy-makers and educators could use test-based incentives more judiciously and effectively to help disadvantaged children (Ryan and Shepard 2008; National Research Council 2011).
Early childhood education is vital if we are to help disadvantaged students. Head Start and similar state-sponsored programs are useful and should be continued. But some Head Start advocates overstate the long-term ability of the current programs to significantly improve the academic skills of preschool students. We need to objectively reassess Head Start and other state-supported early childhood education programs by using appropriate cognitive tests. This will help to ensure that disadvantaged students have access to the best quality preschools and teachers (Vinovskis 2005).
I came to America at the age of six as a refugee from Latvia at the end of World War II. My educational experiences in this country were welcoming and especially helpful. Several grade school teachers worked closely with me to help me learn English as well as gradually improve my initially low test scores. Many immigrants today face a more challenging situation. Yet the number of immigrants coming to the United States since 1965 has increased substantially and likely will continue to do so. We need to see them as a resource that will strengthen our country. To accomplish this, we must help them to integrate into our society.
Many of these immigrants are children coming without knowledge of English or American history. For these children to assimilate into our society, they will also need additional help in our schools. Unfortunately, American history is no longer considered part of the high-stakes testing requirements in K–12 education. Consequently, elementary and secondary schools and teachers are spending less time and effort on civics and history education today to prepare for the other required tests (Mirel 2010; Vinovskis 2009b).
Finally, partisan differences between Democrats and Republicans exist on many education issues. And strong disagreements over funding public and private schools as well as the extent and nature of testing in classrooms are making it more difficult for many Americans to work together to improve education. At the same time, the parties have been able to work together on programs such as Goals 2000, NCLB, and ESSA. We need to improve bipartisan political cooperation in education. States now are playing a larger role in education. As single-party control of individual states has increased, bipartisan cooperation is even more difficult. Given an aging population, rising health costs, increasing public and private debt, significant disillusionment with our schools, and wealthier parents providing better educational opportunities for their own children, we cannot be confident that the funding for public schools will continue to grow. Yet to have the necessary money to attract high-quality teachers and improve our public schools, we must again find ways of working together on education, including resolving more amicably our differences on issues such as testing (Vinovskis 2011).
Footnotes
Note:
I would like to thank Douglas Reed for his helpful comments and Amy Berman and Michael Feuer for their thoughtful suggestions and excellent editing skills.
Maris A. Vinovskis is Michigan University’s Bentley Professor of History and a professor at the Public Policy School. He has published Revitalizing Federal Education Research (University of Michigan Press 2001), The Birth of Head Start (University of Chicago Press 2005), and From a Nation at Risk to No Child Left Behind (Teachers College Press 2009). He worked in the Bush and Clinton administrations on educational research and policy issues.
