Abstract
The adversary evaluation model emerged in a context that favored democratic debate. It was successfully used in a variety of sectors following its inception, but it was abandoned once neo-liberal thinking and goal achievement approaches became dominant. It is time to give it a second chance. The judicial evaluation model (JEM) relies on human testimony, rules of evidence, cross examination and principled deliberation. These features contribute to evaluation independence, a characteristic that is sorely needed in today’s fractured social environment. JEM promotes civil interaction among groups committed to different ideologies. It encourages tolerance and respects pluralism by combining professional authority with direct citizen participation and neutral facilitation. Competently designed and managed, it resists capture by vested interests and it holds promise as an instrument of progressive evaluation focused on the public interest. Since it is demanding and time consuming, it is especially relevant for large, controversial and complex interventions.
Keywords
Adversary evaluation once had its place in the pantheon of influential evaluation models. Unfortunately, it no longer does. I argue that a revival of adversary evaluation would help in the restoration of civility in the public policy sphere. It would also enrich the evaluator’s tool kit and contribute to a renewed emphasis on the ethical and pluralistic values of a discipline that has always thrived on principled and respectful debate.
After recapitulating the origins and rationale of the adversary evaluation approach, I outline its defining characteristics and recognize the diversity of its applications and methods. Next, I examine the sudden rise and precipitous fall of this innovative evaluation model and acknowledge its limitations. I also take stock of its strengths and conclude with a plea to give adversary evaluation methods another chance as part of an urgently needed endeavour to strengthen evaluation independence and revitalize democratic evaluation.
The Advent of Adversary Evaluation
Adversary evaluation emerged in the early 1970s at a time of rising curiosity about fresh ways of assessing social policies and programs. Its advent was triggered by rising concerns about the ethical, technical, and practical limitations of the randomized control trial model that dominated the policy research field throughout the 1950s and early 1960s, and has paradoxically recently re-emerged as a favored instrument in the development evaluation domain despite the paucity of insightful evaluation information generated by experiments, quasi-experiments, and statistical inference tests. In the context of the participatory phase of evaluation history, the experimentialist approach lost its allure as practitioners came to the realization that alternatives to evaluation methods too narrowly focused on the attribution question and were insufficiently sensitive to the diverse needs and perspectives of intervention beneficiaries.
Accordingly, adversary evaluation pioneers sought to explore deliberative approaches to evaluation better attuned to the complexity of real-life decision-making and more reliant on qualitative assessment methods. Thus, Levine (1974) opined that “the experiment cannot deal with historical contexts and it requires the reduction of whole human events to be contrived into dimensions that can be quantified…. In the search for methodological purity, social scientists have often lost sight of the substantive problems the methods were meant to solve” (p. 674). Similarly, Wolf (1975) argued that “in seeking objectivity, the decision maker using these methodologies may exclude a factor that ought to be of fundamental concern: human judgment” (p. 185).
As the paradigm conflict between quantitative and qualitative methods heated up, the search for alternative evaluation model intensified and the adversary evaluation pioneers proceeded on several assumptions: that social policies and programs are multidimensional; that a mix of methods should be used to assess interventions; that the evaluator is not a purely rational, bias-free agent; that most decision makers see merit in alternative interpretations of evaluation evidence; and that, in a democratic society, open debate among rival political groups should be encouraged.
Thus, reacting against the parsimony of experimental techniques focused on goal achievement, the adversary evaluation proponents took on board a diversity of program objectives, acknowledged distinct value frameworks, and encouraged freewheeling debate in the market for ideas. Toward securing credibility in the newly competitive environment for evaluation theory, they borrowed from time-tested principles of public debate, administrative hearings, and jury trial procedures.
Adversary Evaluation Is a Big Tent
Eventually, adversary evaluation was applied out in different evaluation domains through a variety of processes. Its most visible embodiment, the judicial evaluation model (JEM), sought to simulate legal proceedings as closely as possible: It put the evaluated intervention on trial through a prosecutorial-defense process moderated by a “judge” and involving cross-examination of witnesses and a jury.
JEM was skillfully promoted by the standard-bearer of the adversary evaluation movement—Robert L. Wolf. In time, the wider public came to perceive JEM as the primary adversary evaluation model. Yet, in actual practice, adversary evaluation applications in the 1970s and early 1980s varied considerably with respect to the extent to which JEM tenets were observed.
Thus, Scriven’s (1991) definition is broad based. According to his authoritative thesaurus, adversary evaluation is a type of evaluation in which, during the process and/or the final report, presentations are made by two individuals or teams whose goal is to provide the strongest case for (and against) a particular view or evaluation of the program (for example). There may or may not be an attempt at providing a synthesis, perhaps by means of a jury or a judge or both. (pp. 50–51)
Adversary evaluation alternatives to the strictly constructed JEM model strike a balance between confrontational, competitive stances and collaborative, participatory approaches. For example, controversial social programs tend to elicit sharply different evaluative judgments shaped by diverse stakeholders’ interests and evaluator’s dispositions. In such cases, it is not uncommon for two or more evaluations to be commissioned in order to accommodate different value frameworks and evaluation methods, a triangulation process that illuminates diverse facets of the evaluand and facilitates mediation of competing interests.
Similarly, adversary evaluation approaches could include any genuinely independent evaluations that yield evaluative narratives that differ starkly from the rosy assessments typically embedded in program managers’ self-evaluations. In such circumstances, drawing lessons out of the evaluation endeavor requires resolution of a robust contest that pits self-evaluation against independent evaluation, an inherently adversarial process.
Finally, for most evaluations, the final deliberative phase that precedes evaluation use displays adversary characteristics where diverse stakeholders’ views and interests need to be taken on board and reconciled. Specifically, in formative evaluations of complex programs, the pursuit of fairness calls for an honest brokering role before competing preference functions can be aggregated.
It follows that the essence of adversary evaluation lies in the recognition that stakeholder groups’ interests and perspectives vary. They are frequently incommensurable and tend to differ from those of the intervention sponsor. Thus, adversary and committee hearings are frequently put to work in democratic settings to help resolve divergent views about the performance and value of a social intervention and what might be done to improve its impact so as to serve the public interest.
As stressed by Datta (2005): “among the defining elements of a hearing is the assumption that there are two (or more) sides to almost any question, sides that are adversaries to each other” (p. 214). In sum, while not always the most desirable, there are circumstances (large and complex programs, sharply divergent interests, and polarized public opinion) where adversary hearings are highly appropriate (Smith, 1985).
For Levine (1982), adversary hearings clarify issues through human testimony on the assumption that truth emerges from a hard but fair fight in which opposing sides, after agreeing upon the issues in contention, present evidence in support of each side. The fight is refereed by a neutral figure, and all the relevant evidence is weighed by a neutral person or body to arrive at a fair result. (p. 270)
Not all adversary evaluations comply with this definition. There is room for wide diversity of adversary evaluation applications that may contain the following: multiple, complementary roles for evaluators; provision of relevant and reliable information to the contestants; amplification of weak and neglected social groups’ voices; independent verification of self-evaluation claims; facilitation of principled debate; impartial advice to policy makers; and meta-evaluation. Owens and Hiscox (1977, abstract) summarize the hallmarks of adversary evaluation as better communication between evaluators and decision makers, greater attention to the formulation of key evaluation issues, and increased concern for meta-evaluation.
The Rise and Fall of Adversary Evaluation
Evaluation models are apt to thrive when aligned with the prevailing political consensus. Conversely, their fortunes tend to decline when they diverge from the spirit and mood of the times. In a seminal article, the eminent Swedish evaluation thinker, Evert Vedung (2010), famously depicted the history of evaluation as a succession of waves propelled by larger tides of political ideologies.
According to this vivid metaphor, each wave of evaluation diffusion rises and falls over time leaving behind layers of intellectual sediment that nourish evaluation theory and practice. In the 1950s and 1960s, the experimenting society held sway and the first evaluation wave was rationalist, positivist, and meritocratic. In reaction to its purported elitism, a dialogic wave swelled in the 1970s: It was inclusive, participatory, and supportive of social learning.
During this era, adversary evaluation came into being. In his unpublished doctoral dissertation, Wolf (1973) challenged the quantitative methodologies that failed to involve citizens in a meaningful way. Wolf advocated instead for an evaluation model informed by judicial processes and endowed with constructivist, pluralistic, and deliberative features thus placing a premium on educative inquiry, human testimony, interactive dialogue, and value-driven judgments.
Despite a promising start, adversary evaluation did not survive the neoliberal, positivist wave that engulfed the evaluation discipline in the 1980s. In the new policy context that fed its rise, support for adversary evaluation evaporated. Under the aegis of the “new public management” movement, market thinking invaded the evaluation domain and power holders sought to occupy the commanding heights of policy research. “Let managers manage” was the new mantra and utilization focused evaluation, responsive to managers’ needs and closely adapted to the fee-dependent evaluation market, assumed enormous influence.
The triumph of the neoliberal single narrative also reignited the paradigm wars. Proponents of quantitative methods, disdainful of the subjectivity of qualitative methods, carried on with their increasingly sophisticated approaches despite being assailed with considerable effect by three groups of critics: naturalists who consider social interventions as inextricably involved with human values, interpretivists who reject the very notion of quantitative objectivity, and critical theory advocates who take umbrage at the inherent elitist bias of goal achievement models favored by experimentalists (see, e.g., Gage, 2009, pp. 4–10).
Through attrition and mutual exhaustion, the methodological confrontation eventually abated. An uneasy truce took hold. Mixed methods came to the fore as a pragmatic and convenient compromise. However, the multidisciplinary approach it implied proved elusive given the doctrinal incompatibility of the protagonists. On the one hand, the proponents of quantitative approaches continued to seek an even sharper edge in their pursuit of incontrovertible evidence of verifiable “results.” On the other hand, a wide range of qualitative evaluations and action research initiatives, ranging from interpretative–qualitative studies to critical theory analyses, sprouted and competed for public funding and legitimacy within the academy.
As the evidence-based wave that swelled in the 1990s and continued to wash over the evaluation discipline in the early 21st century, a genuine rapprochement between qualitative and quantitative researchers remained out of reach and a wide range of distinct evaluation models shared the limelight. Rather than signaling a genuine reconciliation among the contending factions, the fragile truce among them heralded an evaluation scene characterized by competitive turmoil under the big umbrella of a rapidly expanding and increasingly cosmopolitan evaluation movement (Chelimsky & Shadish, 1997).
Despite the new and remarkable tolerance of the evaluation community toward a wide diversity of approaches, not all evaluation models survived. Adversary evaluation was one of the losers: It has yet to return to the fold. Thus, according to Datta (2005), adversary evaluation as embodied in the JEM, “appears to be no longer practiced…[and] seems to be a dead end among the approaches developed in the 1970s” (p. 216). To help understand the causes of JEM’s persistent eclipse, it is necessary to draw the implications of its courtroom antecedents, examine its performance record over its short life span, and recognize the heavy odds against it.
Why Did the Judicial Model Stumble and Fail?
Early adversary evaluation initiatives were a mixed lot. Some consisted of hearings used to inform a political debate. Others provided summative inputs for decision-making about program continuation or cancelation. Still others had formative aims regarding program design (Owens & Hiscox, 1977). But these efforts were not deemed sufficient to provide a unified framework to practitioners, or to comply with the coherent standards ideally expected of evaluation models.
Paradoxically, the fluid and generic conception captured by the Scriven definition contributed to a perception that the approach was trivial: a structured mode of oral presentations that articulate multiple stances within a group setting or in which two sides argue the pros and cons of a specific question. As such, adversary evaluation was deemed insufficiently detailed by influential critics who insisted on the design of detailed protocols and standards allowing systematic scrutiny.
Ironically, not all other evaluation models meet the exacting standards put forward by adversary evaluation critics. Surviving models have the advantage of ambiguity and bear the imprint of appreciative inquiry and cooperation. They do not evince the discomfort associated with the confrontational stance commonly attributed to adversary evaluation. For example, House and Howe’s (2000) admirable democratic evaluation model may have survived (although rarely practiced today) in part because it avoided the trap of detailed protocols and procedures, relying instead on shared values and basic operating principles that ensure inclusion of all relevant stakeholders, promote dialogue with and among stakeholders, and involve stakeholders in extended deliberation processes. It seems that the bar for adversary evaluation was set much higher than for other models. In response, adversary evaluation proponents fatefully embarked on systematic efforts to strengthen and document the protocols of the JEM.
In doing so, they adopted the legal system as JEM’s guiding metaphor, and over and above the prosecution–defense rivalry feature of judicial processes, they sought to replicate as closely as possible the evidence testing, cross-examination, procedural rules, and jury deliberations of customary jurisprudential practice. This proved suicidal within an occupational group that prizes its social research antecedents (Stern, 2005) and that tends to neglect the role of governance rules in its work (Picciotto, 2016). Furthermore, most evaluators are temperamentally suspicious of models extrapolated from other professional domains such as art, literature…or justice (Datta, 2015).
Justice and Evaluation
Ideally, a democratic judicial system seeks justice. How then does evaluation relate to justice? The pursuit of justice differs from the pursuit of truth—the main object of the evaluation enterprise. But justice cannot be delivered without truth. This means that the administration of justice, just as the conduct of evaluations, relies on methods intended to get as close to the truth as possible by securing relevant evidence, assessing its reliability, drawing its implications, and exercising sound value judgments.
Similarly, adversary evaluation can be conceived as an application of the deliberative democratic evaluation model which aims at social justice; respects conflicting values, reaches out to diverse stakeholders; provides representation opportunities to poor and neglected communities; helps decision makers resolve conflicting claims; provides reliable information; and complies with legitimate ethical guidelines (House & Howe, 1998). Adversary evaluation writ large may even accommodate transformative social justice evaluation models by giving affirmative status to marginalized groups, interrogating systemic power structures, and using evaluations to protect human rights (Mertens & Wilson, 2012).
On the other hand, the JEM is severely undermined by perceptions of rampant legal machinations abundantly displayed on popular television series: presumptions are put forward as facts; testimonies are subjective and often contradictory; admissible evidence is not credible; solemn oaths do not guarantee truthful or objective statements; witnesses are often biased, swayed by clever arguments, or cowed through intimidation in cross-examinations; and so on. This jaundiced view of the legal process has led to widespread scepticism regarding claims of blind, even-handed justice. The bottom line is that public perceptions of unfair treatment by the court system are widespread, especially among minorities (Anderson, 2014). Hence, and unsurprisingly, the interface between evaluation and the legal profession has been heavily contested.
The judicial model ultimately relies on a choice between rival advocacies. Each side uses highly selective evidence to argue its case. The outcome is shaped by judges who are not always free from bias. In the last analysis, lawyers—unlike evaluators—do not seek to secure truth from facts. This is what juries are supposed to do. Instead, lawyers are tasked to wield all the advocacy tools at their disposal to prosecute or defend individuals or groups accused of misdeeds. Similarly, it is, in principle, for evaluation stakeholders to assess the validity and implications of evaluation findings. It is at the interface between evaluation commissioning, delivery, decision-making, and citizens that evaluation tools and concepts are expected to protect the public interest.
In other words, blatant advocacy is the bread and butter of the legal occupation. By contrast, unless based on evidence and restrained by professional impartiality standards, advocacy of any kind is prone to undermine the credibility of evaluators suspected of having an axe to grind (Chelimsky, 1998, p. 40). Just as scientists, evaluators are expected to fight all human bias, including their own.
Validity tests are therefore central to the legitimacy of the evaluator’s craft and independence of mind is a key competency for evaluation practitioners, as it is for auditors. They are duty bound to aim at strict objectivity, stay at arm’s length from decision-making, and recognize the limitations of available evidence. They have an ethical duty to combine fulsome engagement with stakeholders with proper distancing and recognition of the inevitable fallibility of formative evaluation judgments (Scriven, 1997, pp. 480–481).
This said, it is plausible that evaluation, a meta-discipline, could benefit the legal profession and enhance its credibility. For example, rules governing admissibility of evidence and interpretation of expert witness testimonies would benefit from the validity tests routinely used by evaluators. Conversely, it is not far-fetched to posit that judicial processes developed and fine-tuned over centuries of legal practice might have something to offer to evaluation. Thus, skilled cross-examination could improve the quality of human testimonies in evaluation; the composition of focus groups might usefully emulate the strict standards of jury composition procedures.
Indicating its potential benefits, House (1978) gave pride of place to adversary evaluation when probing the philosophical assumptions underlying various evaluation theories. While theories of justice and theories of evaluation are closely linked, they are distinct, and adversary evaluation began to develop assumptions that linked evaluation and justice. Unfortunately, the premature write-off of adversary evaluation as a legitimate model grounded in subjectivist ethics, intuitionist/participatory methods, and pluralistic politics did not allow the approach enough time to emerge and flourish.
Thus, adversary evaluation rapidly lost its luster when it overreached in its efforts to comply with a courtroom model and to mimic rather than adapt its practices to the prevailing consensus about the role of evaluation in society. Before concluding whether the revival of adversary evaluation would be timely and useful, I now turn to recognizing its limitations and strengths in the connected realms of theory and practice.
The JEM Is Demanding
Prominent critics of JEM (e.g., Popham & Carlson, 1977, pp. 3–6) have pointed to such risks as faulty definitions of evaluation issues, overreliance on hearings, inevitable disparity in the debating abilities of contestants, and frequent fallibility of judges and juries. These are indeed serious obstacles to high-quality adversary evaluation. Still similar risks affect other evaluation models where evaluation designs are weak.
Other judicious arguments levied against the JEM embodiment of adversary evaluation center on the complexity and demandingness of the process and the pitfalls associated with the courtroom model: (i) adoption of trivial or inappropriate features of courtroom procedures, (ii) their frequently polemical and divisive features, (iii) the risk that articulate lawyers enjoy an unfair advantage, (iv) a tendency to give equal weight to rival cases irrespective of merit, (v) unnecessary polarization as advocates do all they can to “win” the argument, and (vi) high costs and long elapsed times.
Another related consequence of emulating judicial processes is the adoption of some of its most theatrical features. These tend to amplify complaints, sow division, and aim at proving guilt rather than at securing a sober and balanced judgment about the merit, worth and value of an intervention. Specifically, the disputatious and manipulative persuasion techniques used by prosecutors and defense lawyers are often effective in swaying jury decisions through selective use of evidence, intimidation of witnesses, veiled appeals to prejudice.
Emphasizing extreme positions regarding the effectiveness of an intervention may neglect the nuanced middle ground that is often the sweet spot of legitimate evaluation findings. Conversely, emphasis on achieving evenhandedness when one of the two rival positions can be shown to be demonstrably false can be highly detrimental. The risk of a hung jury may deny a just evaluation by delaying it. By emphasizing polemics, manipulation or misrepresentations may be unwittingly encouraged.
Excessive zeal in argumentation may lead to specious judgments that are not solidly grounded in evidence. Imbalance in the debating skills of the protagonists may not be fully compensated by an impartial judge or the common sense of a jury. When rigorously transposed to the evaluation process, legal practices may amplify the all too frequent and sadly mistaken public disposition that equates evaluation with faultfinding, retribution, and punishment.
Nor do adversary evaluation methods necessarily eliminate evaluator bias: They merely seek to balance contradictory biases, while other biases may remain unaddressed. Finally, by stressing differences in the interpretation of available evidence, judicial processes may generate polarization and hinder the sensible compromises that would help achieve positive evaluation results where stakeholders’ interests conflict.
Adversary Evaluation Is Democratic
Despite its drawbacks, adversary evaluation has merit in democratic societies. Using the law as a conceptual framework for the design of evaluation processes recognizes that peaceful conflict resolution and the exercise of voice lie at the core of democratic decision-making (Hirschman, 1994). Equally, for Habermas (1987), communicative dialogue among principled individuals is the only way to generate standards “based on attitudes that require critical consideration by means of arguments, because they cannot be either logically deduced or empirically demonstrated”(see Appendix in Habermas, 1987).
Public hearings are standard deliberative tools in a democratic society. At its best, adversary evaluation helps to explore alternative solutions to social problems, to select pertinent issues for debate, to expose hidden assumptions, to give voice to persons affected by the policy intervention, to produce alternative inferences from the evidence gathered, to provide structure to deliberations among stakeholders, and to enable a balanced set of recommendations to emerge (Owens & Wolf, 1985).
Arguably therefore, adversary evaluation writ large can be made to comply with democratic pluralism tenets according to which the public interest is best served by deliberative processes responsive to diverse citizens’ views channeled through representative groups (Wolf, 1979). It involves systematic exploration of the policy context, participatory selection of evaluation issues, rigorous field investigations by opposing evaluation teams, fulsome involvement of representative witnesses drawn from persons affected by the program or policy being scrutinized, careful preparation of arguments by the contending evaluation teams, open display and rigorous critique of admissible evidence, and moderated hearings that allow cross-examination of witnesses.
Adversary Evaluation Is Adaptable
Most of the limitations highlighted by critics can be addressed through proper design and adaptation of the approach to individual contexts. It is not beyond the realm of feasibility for adversary evaluation processes to be guided by an experienced moderator and to culminate in a public presentation of evaluative methods and findings by opposing parties (including opening and closing arguments) followed by principled deliberations conducted by a panel of informed citizens simulating a jury.
Well conducted and in the right circumstances, adversary evaluation would provide decision makers and the wider public with a wide array of information. It would encourage advocates to produce high quality, pertinent evidence. It would subject opposing arguments to rules of evidence. It would provide for cross-examination of witnesses, and as a result, it would reduce the scope for unsupported claims and minimizes evaluator bias. All legitimate stakeholder groups would be able to participate. Neutral facilitation would ensure that decision makers are prevented from exercising undue influence.
In the realm of actual practice, adversary evaluation methods appear to have proved their worth in the education sector (Owens, 1971). Later, applications were also successfully put to work in ethnographic inquiry (Schensul, 1985), employment research (Braithwaite & Thompson, 1981), and the military sector (Miller & Butler, 2008) while Barker and Pistrang (2005) have outlined quality criteria for the use of adversary hearings in relatively small, community-based research programs.
The adversary evaluation architecture is flexible. The judge may be a senior manager, an executive director, an elected representative, a respected personality, and so on. The jury may include beneficiaries’ representatives, other stakeholders, regular citizens, and subject matter specialists. Costs can be contained, and the use of a full-blown courtroom procedure is not imperative. Identifying clear-cut, relevant issues through a participatory process, gathering a common evaluation database prior to the adversary hearings, and ensuring highly skilled and impartial chairmanship of hearings would greatly enhance the cost effectiveness of adversary evaluation in a wide variety of operational domains.
Skeptical observers have convincingly argued that adversary evaluation is less effective than other models when nuanced judgments are in order when the issues to be addressed are not clearly identified and when formative evaluation is called for. They have nevertheless acknowledged that the adversary evaluation approach has considerable merit when used to assess large, controversial interventions that affect diverse stakeholders; to elicit polarized stances among affected groups; and when clear-cut decisions are required as to whether interventions should be funded—or, if they are ongoing, whether they should be continued or terminated (Worthen & Todd Rogers, 1980).
Unfortunately, the short shelf life of adversary evaluation approach has prevented establishing, with certainty, the utility of the approach. Nevertheless, a systematic review of adversary hearings conducted over a 15-year period suggests that the verdict of evaluation history has been too harsh. Smith (1985) examined nine formal applications ranging from small-scale projects undertaken by volunteers to large and complex professional exercises requiring substantial time and resources.
Smith found that recurring difficulties included imbalances in the qualifications and persuasiveness of adversaries, neglect of relevant issues, and the difficulties experienced by citizen panels in producing cogent and helpful recommendations. On the other hand, his review concluded that: the participants and especially the organizers of these studies report general satisfaction with the adversarial hearings. Although they acknowledge a number of problems…. they remain convinced that the procedure is an effective way of to provide for public participation and scrutiny of human testimony in evaluating complex programs and policies. (p. 745)
Adversary Evaluation Supports Independent and Ethical Evaluation Practice
Why then was the adversary evaluation model prematurely consigned to the dustbin? Could it be that vested interests are not normally eager to subject the programs they manage or sponsor to genuinely independent scrutiny and fulsome debate? Is it because, wary of public reactions, program funders or leaders are unlikely to embark on an evaluation process that leaves no stones unturned unless they are supremely self-confident or providentially uninformed? Might therefore self-interest and reputational concerns help explain why managerial-oriented evaluation approaches have flourished while adversary evaluation has fallen on hard times?
The burden of objective, no-holds-barred evaluation falls on the evaluator—often one who needs to make a living and is ultimately dependent on the goodwill of the evaluation commissioner who may be the very manager, sponsor, or beneficiary whose program is being evaluated. This potential dilemma underlies the distinction between bona fide evaluators who can call a spade a spade and evaluation consultants who are prone to avoid making direct evaluative claims that may upset their clients. Scriven (1997) wrote that we have not yet fully developed “evaluation as a major social service” (p. 492). He suggested that “evaluators and clients are more inclined toward the evaluation consultant role” (Scriven, 1997, p. 492). Hence, the need for institutional arrangements that protect the evaluator from retribution by vested interests so that evaluations can help make politicality responsible. Adversary evaluation could serve that purpose.
Rarely highlighted in the literature is the use of adversary evaluation methods to strengthen transparency and accountability in legislative processes, independent commissions, environmental assessments, and administrative reviews. Nor is the proven potential of adversary evaluation to overcome the toxic effects of fee dependence and commissioner-driven contractual obligations on the integrity of the evaluation process widely recognized.
From this perspective, Mark Twain’s reported quip about the “greatly exaggerated” report of his death comes to mind. While adversary evaluation is now excluded from the indexes of most recent evaluation textbooks, the adversary evaluation approach broadly defined (i.e., not limited to the JEM) is alive and kicking whenever self-evaluation, often carried out with the help of evaluation consultants, is subjected to systematic scrutiny or full-scale reevaluation by bona fide, independent evaluators tasked with attesting to the validity of self-evaluation claims.
For example, independent evaluation is a standard operating practice in the multilateral development banks where independent evaluators report to the supreme authority of the institution. They are tasked with objective verification of all operational self-assessments and have oversight over self-evaluation methods and processes. The independent evaluation units also generate independent policy evaluations and meta-evaluations of their own.
Similarly, evaluations overseen by the legislative branches of Western democracies are often used to challenge or complement self-evaluation by the executive branch. This is also how, without fear or favor, the General Accountability Office is expected to operate. If both the legislative and executive branches have been captured by vested interests, evaluation can still exert influence and serve the public interest when it reports to civil society organizations that operate at arm’s length from government through principled adversary review processes.
Adversary evaluation broadly conceived is therefore an appropriate model for the implementation of democratic evaluation since it can be equipped with protocols that ensure evaluator independence, protect the integrity of evaluation processes, and facilitate the design and conduct of evaluations so that they serve the public interest. It is less vulnerable to manipulation than the prevailing market-driven, managerial-oriented system and can be designed to create a level playing field for the exercise of voice and the interpretation of results. Expertly designed and handled by competent evaluators, it would help to ensure that power holders’ promises are compared with actual delivery of results through impartial, fair, and transparent assessment processes.
Is It Time for a Second Look?
In the wake of the 2008 financial crisis, the policy world has become sharply divided and the ghost of inequality is again haunting the world. In the new and now global Gilded Age, the world’s three richest people own wealth equivalent to the combined GDP of the world’s poorest 48 countries. The top 1% have more wealth than the remaining 99% of people (Oxfam International, 2015).
While differences in income and wealth attributable to effort, skill, or entrepreneurship are not widely resented, social cohesion and trust in government are severely undermined when distorted rules of the game, predatory economic behavior, or unethical business practices prevail. With social mobility under threat, wages stagnant, and environmental crises looming, hope for the future has been overtaken by nostalgia for an imagined past, and angry resentment has come to dominate politics in many countries.
Instead of inducing free and open debate, the new, unregulated information economy has created obstacles to mutual understanding by atomizing social media users into like-minded groups that fail to interact and reject compromise. Growing intolerance of minorities, rising contempt for authority, and disdain for fact-based policy have divided public opinion, degraded the political discourse, increased the influence of radical groups, and contributed to the spread of populist myths.
These trends could presage a worrying slide toward civil strife and authoritarian rule facilitated by public disdain for expertise and characterized by fact-free policy making. On the other hand, they could be reversed through social activism fueled by a rededication to enlightenment values and liberal democracy tenets. Irrespective of the outcome, and since public opinion remains sharply divided as to the policy remedies required to achieve sustainable economic growth, restore social harmony, and reduce inequality, principled debate across the ideological divide is now at a premium.
A return to civility in public affairs would help usher in a systematic, no-holds-barred reconsideration of the flawed policies that underlie the current social predicament. In turn, such an evolution would favor a greater role for adversary evaluation. This would reflect a sea change in the political sphere and trigger a new wave of evaluation diffusion geared to social impact (Picciotto, 2015). In this context, it seems to be an appropriate time to rediscover adversary evaluation and continue to explore ways in which the significant problems can best be tackled.
Values would be put back at the very center of the evaluation discipline, induce a shift in its policy directions and redirect evaluation toward social impact and participatory democratic tenets. In such a context, adversary evaluation might achieve a successful comeback. In the meantime, to keep the flame alive and to enhance the independence and integrity of the evaluation process, some bold and principled evaluators might opt to explore discrete opportunities for the revitalization of democratic evaluation through adversary evaluation applications adapted to the local context.
Conclusion
Esoteric experimental techniques have not facilitated principled public debate. Manager-driven evaluations often lack credibility. By contrast, fulsome citizens’ involvement, a privileged reliance on human testimony, well-crafted rules of evidence, rigorous cross-examination, and careful deliberation can add up to a deeply meaningful evaluation process that does not yield to vested interests. Here lies the niche of adversary evaluation. It is demanding but so are other evaluation approaches.
Whether modeled according to judicial type norms or simply as public hearings in support of independent evaluation, adversary evaluation appears well adapted to the contemporary social environment. It can promote civil, disciplined interaction among groups committed to different ideologies. Its contestability dimension signals that pluralism in society calls for evaluative practices that accommodate individual as well as group diversity.
The benefits of a carefully designed, well-conducted adversary evaluation process that takes adequate account of lessons drawn from the admittedly limited and curtailed experience of the 1970s and 1980s may well outweigh the significant costs especially where stakeholders hold sharply divergent views, the program is large and controversial and/or dichotomous decision-making is at stake.
By combining professional authority with direct citizen participation and neutral facilitation, adversary evaluation designed to meet the unique needs of individual interventions should have a place in the sun. It would encourage tolerance and mutual understanding. Given its promising record and its unique features, it holds significant promise as an instrument of democratic evaluation mobilized to facilitate a transition phase toward a more equitable economic and social order. When all is said and done, a revival of adversary evaluation would strengthen evaluation independence and help ensure that public value is the ultimate arbiter of the evaluation process.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
