Abstract
In this commentary I discuss how Getting to Outcomes and other empowerment strategies have potential to contribute to a theory of evaluation practice, address the epistemology of quality improvement, and inform the external validity of outcome evaluations.
Abe Wandersman’s article with Alia, Cook, Hsu, and Ramaswamy (2016) in this issue of American Journal of Evaluation derived from his keynote address to the 2015 Eleanor Chelimsky Forum at the Eastern Evaluation Research Society. In my commentary, I will focus on three issues: (1) Why his article exemplifies important connections between evaluation practice and theory, (2) how his article addresses evaluation in the context of continuous quality improvement (CQI), which presents thorny and fascinating epistemology and methods challenges, and (3) how the approach he outlines helps to inform the external validity of outcome evaluations.
Connecting Evaluation Practice to Theory
My impression is that past speakers at the Chelimsky Forum tended to focus on evaluation theory first and then derive implications for practice. In contrast, Abe’s work exemplifies how a theory of evaluation derives seamlessly from practice. Many things are refreshing about the empowerment evaluation approach, but chief among them in my opinion is the immediacy with which abstractions derive from practice guidance. In particular, both Getting to Outcomes (GTO) and the Interactive Systems Framework for Dissemination and Implementation (ISF; Wandersman et al., 2008) are central underpinnings of the Chelimsky Forum article, abstracted from and challenged by evaluation practice and observations about programs.
The clarity and simplicity of GTO and ISF are deceiving. They reflect a theory about formative evaluation in the sense that Cronbach and his colleagues (1980) used the term—that evaluations are often used in a formative manner regardless of the age or established nature of the program. The cycle of steps in GTO uses what we know about programs and implementation to codify formative uses of evaluation as theory. In general, people sometimes think empowerment evaluation is just formative evaluation by a new name, but that is not correct. GTO, in particular, is not formative evaluation as conventionally viewed. Rather it is a cycle of evaluation inquiry, ideally based on the life cycle of programs or practices (Scheirer, 2012). GTO certainly employs a stage of conventional formative evaluation, but also outcome evaluation, of course. It also incorporates strategic planning, an area that is often neglected in conventional program planning and evaluation, and CQI.
Abe’s work is deeply informed by research on implementation science (Lobb & Colditz, 2013) but with additional abstracted principles. For example, underpinning the ISF is an extensive literature review on what makes organizations ready to implement an evidence-based intervention with quality (Assistant Secretary for Planning and Evaluation, Department of Health and Human Services, 2014). The Robert Wood Johnson Foundation is now supporting the development and psychometric testing of an instrument to assess readiness based on this review. The project will test the instrument’s predictive validity and in the process will build theory about implementation.
For implementation science, GTO also starts to formalize a theory of how factors at various levels of influence produce variation in implementation and outcomes. These range from international factors to provider and recipient factors. In this way, GTO connects to the emerging focus on systems thinking that is beginning to guide more evaluation planning (e.g., Parsons, Jessup, & Moore, 2013). The examples Abe and his colleagues present are fairly conventional in that they focused on the levels of community, hospital, provider, and patient characteristics and actions as sources of influence. Evaluators often presume that these are the most powerful levels to produce or control variation, but the reality may be quite otherwise depending on the topic. For example, we are discovering that a surprising variety of nonhealth policies at local, state, and national levels have profound effects not only on health outcomes (influence on the social determinants of health) but also on the implementation of health programs (Robert Wood Johnson Foundation, 2015). The value of Abe’s table 1 is the presentation of a full range of factors and levels. Knowing where the power over implementation resides helps to identify leverage, which in turn helps to prioritize evaluation questions and overall strategic planning for a program.
Consider the case of tobacco control. The powerful sources of influence that we understood and manipulated have varied over time and place. For several decades after the original surgeon general’s report on smoking and health, the evaluation focus was strictly on smoking cessation methods implemented by practitioners for individuals (Department of Health, Education and Welfare, 1964). This is still an important leverage point (Cochrane Tobacco Addiction Group, 2016). In the late 1980s, however, medical care providers’ advice and assistance became a topic of study. It was often prompted by a reminder in the medical chart, an organizational leverage point (Burns, Cohen, Gritz, & Kotke, 1993). Affordability of nicotine replacement therapy, a pricing issue, became important as well. In the present day, power over implementation of counseling and therapy has shifted somewhat to states and private insurance plans only some of which pay for cessation therapy and counseling (Kofman, Dunton, & Senkewicz, 2012). Societal norms about smoking have changed radically since the 1960s (Center for Public Program Evaluation, 2011). Along the way, other policy and environmental factors were discovered that discourage tobacco use, and the level of government that imposed these policies varied over time. In the 1990s, it emerged that cigarette taxes deter youth from initiating smoking, and both state and federal taxes on cigarettes increased. The federal government banned cigarette advertisement aimed at children because of its association with a rise in youth smoking. By the 2000s, it was established that smoke-free indoor air laws help people to quit, increasing the passage of such laws at the local and state level (Center for Public Program Evaluation, 2011). These policies varied greatly by locality, but state and federal laws have largely preempted local policies, making for more uniform taxation and smoke-free environments. With new questions about e-cigarettes, policy control is once again shifting to state and local levels (Public Health Law Center, 2015).
Certainly, many evaluators in the present day focus on multiple levels of influence—but perhaps they do not prompt stakeholders to consider carefully all the levels of influence that might affect a problem and a program. What differentiates GTO is that the stakeholders are engaged in strategic planning and use of findings. If the source of powerful variation is Level X, and if stakeholders can modify or work around Level X, then they pose evaluation questions about strategy for Level X.
Epistemology: How Do You Evaluate a Process of CQI?
CQI is probably needed to achieve population impact through dissemination and adoption of evidence-based programs (Chambers, Glasgow, & Stange, 2013). I would argue that empowerment evaluation strategies, and GTO specifically, assist us with the thorny problem of CQI as an object for evaluation. One reason is that it helps to define the evaluand when both the situation and the interventions can be complex or at least complicated (Hawe, 2015). By definition, implementation changes are ongoing in CQI, a response to local context. Yet without adequate description, the result can be enormous, idiosyncratic variation, as seen in research on quality improvement in medical care. In response, new publication guidelines in medicine require better description both of the improvements themselves and the context that required the CQI (Davies, Batalden, Davidoff, et al., 2015; Hoffmann et al., 2014).
Abe and his colleagues describe failures to replicate the effectiveness of surgical checklists. My colleagues in medical care quality improvement believe that these were likely cases of “cargo cult science” named after the cargo cults of the South Pacific after World War II. This occurs when “poor understanding of what an intervention really consists of, what it does, and how it works thwarts the meaningful replication of interventions that were successful in their original context” (Davidoff, Dixon-Woods, Leviton, & Michie, 2015, p. 5). By contrast, the GTO process aims to introduce discipline into the description and phased study of improvements. Planning, documentation of implementation, and consideration of the fit between an intervention and the context then culminate in outcome evaluation and subsequent CQI; potentially, they create the conditions for more articulate and systematic evaluation of CQI.
Causal inference about the effect of CQI is challenging because implementation is not stable over time and place. For example, the Plan, Do, Study, Act rapid-cycle strategy was introduced to medical care from industry. It uses multiple rapid cycles over time that modify various aspects of implementation and then incorporate only those that seem to promote the aim (Berwick, 1996). The aim may involve an overall change or reduction in variation or both. The statistical techniques to assess improvements are far advanced, including the presentation of time series on regularly generated indicators, such as medication errors, waiting time in clinics, or falls in elderly patients (Moen, Nolan, & Provost, 2012). There is often quite a wealth of time-series data to assess whether quality improved; the problem is causal attribution. Because there is no single, defined point in time when improvement begins, interrupted time series as conventionally used in evaluation is not appropriate for CQI evaluations. Dennis Ross-Degnan (2015) is currently examining some conditions under which interrupted time series might usefully be employed for quality improvement interventions.
GTO handles the problem of causal attribution by specifying a stage for outcome evaluation, a moment in time when the effects of intervention are stabilized for examination. However, that does not address the value of the entire GTO cycle of steps or whether improvements might be attributed to CQI specifically. In a cluster randomized experiment, Chinman, Acosta, Ebener, Malone, and Slaughter (2016) evaluated the effects of adding GTO to an evidence-based teen pregnancy prevention program implemented at 32 boys and girls clubs. Clubs assigned to GTO had better fidelity and youth outcomes.
For rapid-cycle improvement in medicine, Schouten et al. (2008) conducted a systematic review using a fairly limited set of controlled studies, indicating moderate effects. Another strategy is to compare effect sizes of interventions that do and do not employ rapid-cycle improvement methods. In our cluster randomized trial of 114 hospitals, one arm committed to a specific change in neonatal care, supplemented by rapid-cycle collaborative improvement, while the control arm did not commit to the change. The trial achieved one of the largest effect sizes in the medical care quality improvement literature, which is otherwise noteworthy for small effects at best (Horbar et al., 2004).
These positive developments notwithstanding, CQI is likely to remain slippery and elusive when subjected to conventional evaluation tools and methods. Indeed, a formal process like GTO is probably necessary to overcome slipshod descriptions of improvements and to establish a moment in time when outcomes can become stable enough for assessment.
Assessing External Validity
Because GTO offers a systematic way to involve practitioners in adopting and implementing evidence-based interventions, I believe that it can be helpful to assess external validity. External validity in evaluation is closely allied with the spread and translation of evidence-based interventions. Both GTO and ISF aim at precisely this. GTO assesses the fit of a program to context—a method of assessing the prior probability that the program will be successful in context. Once adopted, the formative process examines adaptations that might be required, while maintaining fidelity to core components.
For a comprehensive description of a revised approach to external validity, see Leviton (in press) and Leviton and Trujillo (2016). Here are the highlights as they might pertain to GTO and other empowerment approaches. External validity requires inferences about whether a “causal relationship holds over variation in persons, settings, treatments, and measurement variables” (Shadish, Cook, & Campbell, 2002, p. 20). As Leviton and Trujillo (2016) attempt to demonstrate, practitioners have important things to say about that variation. Their input is best used in a structured, thoughtful way, and we offer suggestions for that. It requires a fresh take on program theory, the logic of induction, and sampling.
Program theory guides us to some of the important features of context that are likely to facilitate or impair implementation and effectiveness. Practitioner input assists theory by helps us to identify those context features to begin with. In addition, practitioners can challenge and sharpen program theory by applying it in diverse contexts. Practitioners also adapt interventions in response to context features—treatment variation. Through consistent interaction in a framework such as GTO, practitioners can help to determine which adaptations are helpful, harmful, or irrelevant.
Induction helps us to rule out the causal relationship between program and outcome in contexts that impair implementation. It also identifies moderators of effects that reduce the strength and integrity of intervention. This type of inductive logic resembles very closely GTO’s step of assessing the fit between program and setting. Induction also helps to rule in contexts where a program is likely to fit a setting—induction as a statement of probability. By using the program theory, the logic of induction can help to establish whether adaptations are likely to be helpful or harmful. Ideally, some adaptations should then receive independent tests of effectiveness.
Sampling practitioner input informs us about the frequent context features and adaptations because those are the ones that will have consequences at a population level. Context features and adaptations that are important (based on program theory) and frequent (based on an expanded sample, using practitioner experiences) are those with the greatest information value and might receive priority for further tests of effectiveness.
It may seem odd to suggest that empowerment evaluation could assist with the puzzle of external validity given its popularity for evaluation of single-site interventions. Yet taken together in a systematic effort, practitioners’ experiences of implementing evidence-based interventions will offer a broader, more diverse sample of contexts than most researcher–developers are likely to see. Why not utilize this experience as a source of information about variation across contexts? There are very few formalized ways to do this, and empowerment evaluation offers several.
Summary
Empowerment evaluation and GTO specifically have substantial power to assist with the spread of evidence-based interventions with quality. In the process, they aim toward a theory of evaluation practice, assist with the thorny problem of evaluating CQI, and have potential to improve assessment of external validity. Empowerment evaluation is not the panacea for evaluation because there can be no such thing. Evaluation is inherently a flawed, unsatisfactory, but necessary activity, and subscribing to the one perfect approach is like believing in the tooth fairy. Yet, the ability of empowerment evaluation to engage the practitioner in a structured, sensible process represents an important contribution to both practice and theory of evaluation.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
