Abstract
The contribution discusses the risks and benefits of the push for evidence-based decision making. More particularly, it focuses on the risks of a results-focused approach; the special risks of the preferential treatment of RCTs; and the benefits of an evidence-focused approach.
Keywords
Although most participants at the UKES conference were from the UK, there were also visitors from Africa, Asia and North America. This ‘Counterpoint’ was the reaction of one practicing US evaluator to the advocacy by many of the UK policy-based participants in the conference of experimental approaches often grouped together under a ‘what works’ flag.
Introduction
In 2014, I attended my first meeting of the UK Evaluation Society. As an evaluator with nearly two decades of experience with clients in the United States, I was struck by some of the similarities in the evaluation discourse between the USA and the UK. At the UKES meeting, ‘What Works’ was introduced as a repository for evaluation studies meeting a high standard of rigor. What Works promises to transform public services by providing a repository for rigorous impact studies which can be used to guide programme funding decisions. Randomized controlled trials (RCTs) were clearly favoured as inherently meeting the standard of sufficient rigor for inclusion in this repository.
Results-based accountability has also been a buzzphrase for the past decade in the US. The push for RCTs as the ‘gold standard’ for evidence-based evaluation is a theme often heard in the US. As an evaluator, I welcome any call to increase the rigor of evaluation studies. My concern is that this current focus on impact evaluation detracts from the equally important role of other types of evaluation. This effect is further exacerbated by the mistaken belief that RCTs represent the ideal evaluation design in the real world setting. I am sceptical of the value of preferential treatment for one specific evaluation design above others. I am also concerned that a focus on demonstrable results may overshadow the importance of evaluating implementation. In my experience as a practicing evaluator, I believe there are potential risks inherent in policies that overemphasize the value of RCTs and evidence-based funding. If unchecked, this disproportionate emphasis on results will invariably lead to less impactful and less innovative programmes. In this contribution, I discuss the risks and benefits of the push for evidence-based decision making.
Focus on results may draw attention away from the sometimes greater importance of process evaluation
We worked with one client who was very concerned about rigorous evaluation design. Sites were carefully randomized to receive a more intensive intervention, a less intensive intervention, or no intervention. Randomization was done in such a way as to ensure comparability of population across sites. A several month washout period was diligently observed throughout the entire state to eliminate possible exposure to the intervention. The rollout of each phase of the intervention was carefully coordinated across all sites and intercept survey data collection was planned on all possible days of the week and at several different times during the day to maximize representation of ages, ethnicities, work sectors, and parents or childless adults. Multiple survey teams were launched to ensure that data were collected at each site at the same time each day. Sample sizes were generous to ensure sufficient statistical power for multivariate analysis. The study design was as rigorous as possible for a population-based intervention. There was one problem that was overlooked by the client: oversight to ensure consistent implementation of the intervention at all sites. No process evaluation was included in the scope of work, and no supervision was provided by the client to ensure that implementation was faithful to the model. Our team fortunately incorporated some process measures into the evaluation. Our main findings were gross inconsistencies in implementation across sites – to the extent that many of the intervention sites did not provide sufficient exposure to the intervention to draw any conclusions. Intervention materials were completely missing at many sites and the quality of service delivery varied widely across sites. This evaluation was very resource-intensive. Rather than focusing on rigorous measurement of outcomes, this organization would have better used their funds had they invested in a detailed process evaluation and evaluability assessment to determine whether the intervention was even ready to yield measurable outcomes. Indeed, recent research in the US has found that fewer than 30% of reviewed programmes are ready for outcomes evaluation, suggesting a strong role for formative evaluation to strengthen programme quality (Leviton et al., 2010). It appears that an emphasis on evidence-based research may not itself be based on evidence.
Results-focus may discourage innovative thinking
In the US, the focus on standardized testing in schools can lead thinking away from fundamental policy questions. We were involved in analysing the school readiness scores of kindergarten students who had attended academic-focused state-sponsored nursery schools in a large state in the US. The goal of the evaluation was to determine whether children who attended nursery school were more ‘school-ready’ than those who did not. The results indicated that yes, those children had an edge in kindergarten. However, there was another finding with potentially greater policy implications: The number one predictor of school readiness was not nursery school, it was gender. Girls had an overwhelming edge over boys in primary school. This should have led to deep questions about why school environments and educational practices are currently designed and practiced in a way that advantages one gender over another. However, this result was quietly ignored in favour of the ‘evidence’ that the preschool programs ‘worked’.
Results-focus prejudices interventions toward superficial approaches that are easy to evaluate
This is an example of the evaluation tail wagging the intervention dog. Recently a large urban school district in the US embarked on a very public, very resource-intensive transformation to address abysmal student scores. Over $1 Billion (US) was poured into the intervention, which focused on improved teaching methods. The improvements in teaching were expected to lead directly to improved test results. This project was a dismal and embarrassing failure for the city involved. Why? The intervention failed to address any underlying causes of poor student performance – extreme poverty, parents working three jobs, parents without jobs, parents who were drug abusers or mentally ill, and schools that had been already abandoned by higher performing students in favour of private education. Poverty, substance abuse, unemployment – these are complex problems that require multifaceted responses engaging all community stakeholders to create solutions. Addressing only one cause in a multi-causal model may sometimes yield some tangible results in the short term, but will never result in long-term success. Had this intervention instead been focused on unpeeling the layers of challenges that led to poor school performance, and had striven to measure progress on indicators such as parental education, availability of jobs, accessibility of substance abuse treatment – the investment may have yielded slower, but more lasting outcomes.
Results-focus encourages cherry-picking
Programmes working with vulnerable populations with complex problems face particular risks when pushed to produce verifiable outcomes. Examples would include programmes to reduce homelessness, a target population that invariably includes individuals with substance abuse and/or mental illness. Dual-diagnosed individuals face great difficulties in maintaining employment and housing and have complex challenges that are difficult for programme staff to address. Requiring programmes such as these to demonstrate tangible outcomes for all clients as a condition of funding can result in the programme cherry-picking their clients to ensure demonstrable outcomes. Programmes establish exclusionary criteria that restricts clients to those without mental illness or substance abuse problems – in short, leaving the programme to ‘treat’ the population that would be most likely to escape homelessness without assistance. When faced with the requirement that the continuation of funding be tied to tangible results, many programmes will be hesitant to work with populations with complicated and difficult problems. This further reduces the likelihood that any breakthroughs can be made in working with these populations.
The special risk of preferential treatment of RCTs
There are potential risks inherent in limiting the definition of a rigorous study to narrowly defined menu of designs and methods. Misunderstandings of what constitutes an RCT abound, and randomized quasi-experimental designs are often mislabelled RCTs. This results in evaluators overlooking study designs that would be more appropriate to the intervention being evaluated.
Calling it an RCT does not make it an RCT
There is a difference between randomization and a randomized controlled trial. It is rare to be able to properly execute a randomized controlled trial in the real world. It is often impossible to control access by the control group to a similar intervention, or even to the specific intervention being tested. One of our clients was interested in studying the impact of an ethics curriculum in secondary schools. They were interested in a randomized design. It was not possible to coordinate with the school district to actually randomize student assignment to the classes in which the curriculum would be given, but the schools were willing to randomize the specific classrooms that would receive the curriculum. Randomizing the schools themselves was not feasible because the school populations were too different from one another. Randomizing classrooms helped to ensure that the students exposed to the curriculum were demographically similar to the control students. However, there was no way to prevent the flow of information from students who viewed the intervention to their friends who had not participated. In all likelihood, some students would discuss the ethics examples over lunch with their friends who had not been in class. The resultant bias in such a study is not random. Failure to eliminate exposure of the control population to some aspect of the intervention will always result in dilution of the effect size and will increase the likelihood that the evaluation supports the null hypothesis. Promoting the RCT as a gold standard, while simultaneously being unable to truly control exposure to the intervention in the vast majority of programmes, will lead to an increase in programmes that are erroneously determined to be ineffective.
RCTs are not a ‘gold standard’
It may be incorrect to assert that RCTs are appropriate to hold up as a gold standard for evaluation. What works in vitro does not always work in vivo. The clearest evidence for this stems from pharmaceutical trials, the birthplace of RCTs. Drug trials often have so many exclusionary criteria and are restricted to such a closely defined population that the results do not always apply to the population for whom the drug is designed. This can result in drug interactions and side effects that were not evident during the RCT process. In the evaluation world, interventions that can actually be evaluated using an RCT may be so specific to a particular sample group that they may never be scalable to a larger population. There may be very few real-world opportunities to implement a true RCT.
Benefits of evidence-based design
As an evaluator I welcome any call to increase the rigor of evaluation studies. Our team has seen numerous examples of ways in which programmes are strengthened when administrators are required to demonstrate statistically significant evidence of achieving outcomes. This increased pressure for programme accountability encourages organizations to test their assumptions and increase the rigor by which their work is evaluated. New NGOs or government programmes can particularly benefit because they are then required to think more strategically about their work. Our firm’s work with a number of clients has illustrated these benefits.
Evidence-focus encourages examination and testing of assumptions
One client we worked with was convinced that their pilot intervention was changing client lives and reducing care costs, but was basing this belief on case reviews and anecdotes. The pilot funding was coming to an end, and the client wanted to know whether and how to absorb the programme into their general operating budget. We systematically reviewed their data and discovered that the intervention was indeed highly effective, but only for 25 percent of their clients who fit a particular high-risk profile. Lower risk patients did not appear to benefit much from the intervention. The client learned that their assumptions were only partially correct when tested by examination of the data. They also learned that they could cost-effectively integrate the pilot into operations by restricting the programme only to clients who met the appropriate profile.
Evidence-focus encourages increased rigor in evaluation design
We have worked with several organizations that were interested in rigorous, defensible evaluations, but were concerned that the qualitative data that they were most interested in would not hold up under examination. When these questions arise, it is an opportunity for us to discuss what rigor means in qualitative research. We are able to introduce choices in sampling methods, achieving thematic saturation, and coding of results using grounded theory as approaches to achieve replicability and rigor in qualitative design.
Evidence-focus helps fledgling NGOs start thinking more strategically about their work
One of our clients was a charitable organization that provided support to families in medical crisis. They sought an evaluation to generate the outcomes data that was increasingly being required by donors. In helping them design a results tracking system, we also helped them to refine their goals and objectives to become more strategic, targeted and realistic. Moreover, systematic data collection revealed the unexpected finding that Spanish-speaking families in medical crisis were not connected to the extensive support networks of English-speaking families and were therefore highly dependent upon this small NGO. The external pressure to produce evaluation results resulted in this client sharpening the focus of their goals and objectives as well as gaining tangible evidence of their unique contributions to Spanish-speaking families.
Implications
This would only be a matter of academic debate were crucial policy decisions not hinging on interpretation of the evidence presented in ‘What Works’. The most serious risk is the way the call for increased rigor could be interpreted by policy makers. Many NGOs and programme staff understand the complexity of their work and the long-term nature of their efforts. Foundations are sometimes willing to fund experimental or long-term programmes in the interest of increasing the reach of their dollars. Policy makers, however, do not have such an intimate understanding of the problems faced by the populations served, and are bound by their positions to make prudent use of public funds. Policy makers depend upon sources such as What Works to distill research and guide policy decisions. If the studies presented in What Works are limited to simple problem-simple solution, policy makers could be led to believe that those interventions are the only ones that should be funded, and that future policy should be based upon those studies. Bearing the stamp of rigor, What Works has the potential to lead us away from difficult strategies that address root causes and instead divert resources to programmes with little lasting impact and minimal reach.
The purpose of this essay is not to suggest abandoning ‘What Works’, but rather to encourage consideration of the inherent risks and to prompt discussion of ways to mitigate those risks. Rather than single out a particular method as meeting the bar of rigor, the policy and evaluation community should strive for rigor in all evaluation designs. Evaluators, implementers and decision-makers should all guard against the impulse to leapfrog from program design to impact. Evaluability assessment and process evaluation can be used to carefully assess success in implementation in a variety of populations and settings. For programmes of sufficient maturity to warrant impact evaluation, a wide range of methods exist (Stern et al., 2012; Wilson-Grau and Britt, 2012) that provide ample opportunity to increase rigor. In cases where an unsullied control group is unattainable, the lowly pre-post design may yield the most reliable results. Tailoring the data collection methods and evaluation design to the intervention ensures that interventions will not be biased to become evaluation-friendly.
In closing, what evaluations should be included in What Works to ensure that policy makers have access to the best data to make decisions? There are no short-cuts in assessing the quality of evaluation studies. Each must be reviewed individually with attention to sampling design and methods, suitability of the design to the objectives of the study and to the population being studies, appropriateness of analytic methods, replicability, and the extent to which conclusions are supported by the data. Both qualitative and quantitative research have standards for rigor and replicability which can be used to screen potential evaluations. Vigilance to unintended consequences of evidence-based decision-making and dialogue on solutions will ensure that What Works delivers on its promise.
