Abstract
This article considers a new method, called Cross-Site Attributional Model Improved by Calibration to Within-Site Individual Randomization Findings (abbreviated as “CAMIC”), which seeks to reduce bias in analyses of the contributions of a program’s various design, implementation, and contextual characteristics to its overall impacts. It requires a multi-site experiment where some (or all) sites randomize individuals to one of three arms: a standard treatment group, an enhanced treatment group (that receives the standard treatment plus a program “enhancement”), or a control group (that has no access to the program). A recent evaluation—that of the Health Profession Opportunity Grants (HPOG) program—provides a motivating example of this design and its potential for both methodological and substantive lessons; and the article considers other evaluations that would be fitting for deploying and learning from CAMIC. We conclude that the promise for CAMIC lies in situations where the correlations between the selected program enhancement and alternative program characteristics of interest are relatively high, implying that producing an experimental estimate of the enhancement can reduce bias in the estimation of other non-randomized program characteristics.
Introduction
Experimental evaluations—in which study participants are randomly assigned to treatment and control groups within sites—provide researchers with a method for understanding the overall effectiveness of programs without selection bias. However, once we understand the magnitude and sign of program impacts, the next question is why the program did or did not have its intended effect. Programs that operate in many locations—and the multi-site evaluations that accompany them—offer an opportunity to “get inside the black box” to explore how program characteristics determine its impact on outcomes of interest. They can do so by leveraging both planned and natural variation in the program world. Beyond multi-site experiments, multi-armed experiments offer greater potential still.
For example, randomly assigning individuals within each site either to a standard treatment group or to an “enhanced” treatment group that is offered the standard treatment plus a program enhancement is a multi-armed design that has both planned and natural variation. The planned variation comes from the three-armed experiments where the two treatment arms represent variants of the program; and the natural variation comes from the fact that the evaluation takes place in multiple locations. Researchers can produce an experimental estimate of the influence of the program enhancement on impact magnitude via the multi-armed experiment. But, the program components themselves are not randomized to the locations within which evaluations take place. Although they could be, the program’s design (choices about the components that are part of the program) is not often randomized. Indeed, there are relatively few circumstances in which randomizing at the program level would be feasible or desirable. Further, some characteristics of a program—such as the nature of its implementation (as generated by the dynamism of a program manager or office culture, e.g.,) cannot be randomized. Instead, it is routinely the case that each site chooses its own configuration of program components to adopt, each possesses its own set of implementation features, and each operates within its own context with its own target populations. In a multi-site evaluation, design and implementation characteristics tend to vary naturally across sites.
In this case, researchers will rely on that naturally occurring cross-site variation in program characteristics to produce non-experimental estimates of the influence of program characteristics on impact magnitude. Early examples of this sort of work include Greenberg et al. (1993; 1994) and Bloom et al. (2003) who pooled multi-site job training and welfare reform experiments to assess how certain program design and implementation characteristics associated with experimentally based impact estimates. As those researchers pointed out, impact estimates may be subject to omitted variable bias if there are unmeasured or omitted site-level variables that are correlated with the program characteristics included in the model and that influence treatment impact magnitude (e.g., Bloom et al., 2001; Moulton et al., 2014). For example, a local program manager’s enthusiasm and leadership might be associated with both the choice to offer access to peer support groups as part of their job training; and those same traits may also lead the program to have larger or smaller impacts. In this situation, the manager’s traits are associated both with what is offered and with impact magnitude. If the researcher fails to control for these traits (which are often difficult to measure), then cross-site non-experimental estimates of the relationship between program characteristics and impact magnitude will be biased.
This article considers whether having an experimental estimate of the contribution that one program enhancement makes to impact magnitude can help to reduce the bias in cross-site non-experimental estimates of the relationship between other program characteristics and impact. More specifically, the approach we use is called the Cross-Site Attributional Model Improved by Calibration to Within-Site Individual Randomization Findings Method (abbreviated as “CAMIC”). The CAMIC method uses the difference between experimental and non-experimental estimates of the program enhancement to choose the statistical model with the least-biased non-experimental estimate of the program enhancement. The method then uses this model to estimate the influence of program characteristics for which random assignment did not take place.
The article proceeds as follows: first, we discuss two common approaches that prior researchers have used to estimate the influence of program characteristics on impact magnitude when those characteristics are not randomized. Next, we explain how multi-site, multi-armed experiments can eliminate bias in estimates of the effect of randomized program components and, in turn, potentially reduce bias in estimates of the effect of non-randomized program characteristics, via the CAMIC Method. Finally, we offer some evidence on how the CAMIC method performs in simulations, and discuss implications for evaluation research and program practice.
Methods for Analyzing the Contribution of Program Characteristics to Impacts
Two methods have commonly been used in recent years as analytic means for assessing the contribution of program characteristics to impacts in the context of experimentally designed evaluations: (1) a multi-level modelling approach, and (2) an instrumental variables approach. Both of these use an experimental evaluation design, essentially estimating the relationship between the site-level program characteristics and experimental site-level program impacts. This section describes briefly each of these two approaches and then describes how multi-site, multi-armed experiments provide a way to eliminate bias in estimates of randomized program enhancements. Capitalizing on the multi-site, multi-armed design, the next section will discuss the CAMIC approach, a distinct kind of “within-study-comparison” (WSC) opportunity to improve the non-experimental analysis by leveraging experimental results.
Multi-Level Modeling of the Contribution of Program Characteristics to Impacts
To inform how program characteristics contribute to impact magnitude, a number of studies have taken advantage of the naturally occurring cross-site variation in the specific services offered by programs (program components) and in how these services are delivered (implementation features), while controlling for other contextual influences (e.g., Bloom et al., 2001, 2003, 2005; Dorsett & Robins, 2014; Godfrey & Yoshikawa, 2012; Greenberg & Robins, 2011). These analyses commonly use multi-level modeling based on Bryk and Raudenbush (1992) to compute non-experimental estimates of the relationship between site-specific program characteristics and experimentally estimated impact magnitude. This method of relating program characteristics to impact requires the assumption that the model is not subject to omitted variable bias in that there are no site-level variables that are both correlated with the program characteristics included in the model and influence treatment impact magnitude (e.g., Moulton et al., 2014).
This kind of analysis originated in the work of Greenberg et al. (1993), who may have been the first to categorize the questions of interest as being about “prying the lid from the black box.” That is, impact analyses treat “the program” as a “black box,” an unknown space within which varied activities would take place. The black box nature of the characteristics surrounding program implementation begs the questions: which of those activities, and which implementation and contextual factors, are responsible for any observed program impacts? Placing a multi-site evaluation in a multi-level framework provides a means for answering those “black box” questions, at least correlationally (Greenberg et al., 1994).
Using Site-by-Treatment Interactions in an Instrumental Variables Analysis
More recently, similarly motivated analyses have used each of the site-by-treatment interactions that exist in a multi-site experiment as instruments for estimating the contribution of site-level characteristics (e.g., Bos & Granger, 2000; Kling et al., 2007; Magnusson & McGroder, 2002). One key instrumental variables assumption is the exclusion restriction. In this context, the exclusion restriction means that the effect of the program on the outcome must be mediated—and mediated alone—by the program characteristic of interest. However, programs are commonly multifaceted—especially those subject to these “prying the lid from the black box” sorts of questions—which makes it difficult to separate the effects of one program characteristic from the effects of another characteristic using instrumental variables (Gennetian et al., 2002). An additional challenge of using instrumental variables is that, if the impacts of the treatment on the program characteristic do not vary significantly across sites, then the use of multiple instruments may lead to substantially decreased precision and increased finite sample bias (Reardon et al., 2013).
A Multi-Site, Multi-Armed Experiment’s Distinctive Method
Going a step further, the impact evaluation of the first round of Health Profession Opportunity Grants (HPOG 1.0) included design elements that could eliminate bias in estimates of the contribution of specific program characteristics (Peck et al., 2018). The impact evaluation assesses whether providing access to career pathways training for healthcare occupations improves participant outcomes overall while also examining what about the multi-faceted intervention associates with impacts. To do this, eligible individuals were randomized to the HPOG treatment group or to a control group that did not have access to HPOG-funded services. Importantly, in some sites, there were two treatment groups, where one treatment group had access to HPOG while the second treatment group had access to HPOG enhanced with one of three selected program components: (1) emergency assistance for specific needs, (2) noncash incentives designed to encourage desirable program outputs and outcomes, and (3) facilitated peer support groups. This design enabled the researchers to produce experimental estimates of the contribution of each of these enhancements—peer support, emergency assistance, and noncash incentives—to the HPOG Program’s impact. 1
As described further below, the HPOG evaluation design allowed for both experimental and non-experimental evidence on the relative effectiveness of each of these program components. As a result, HPOG involves a special kind of WSC that provides an opportunity for additional methodological learning. The goal of other WSC studies (beginning with LaLonde, 1986) has been to learn which approach to measuring program impact, subject to selection and other sources of bias using observational data, best replicates an experimental finding for the same impact quantity. A series of such studies has begun to point evaluators toward the conditions that yield more-reliable non-experimental findings than other options do (see Cook et al., 2008; Glazerman et al., 2003). 2
The CAMIC Method
The CAMIC method was conceived in the HPOG 1.0 Impact Study as a possible way of improving on the Multi-level Modeling of the Contribution of Program Characteristics to Impacts Method by employing a WSC-style approach, thereby reducing bias in non-experimental estimates of the contribution of selected program characteristics to overall impacts (Bell et al., 2017). It does so by identifying the most useful measures to include as covariates in the statistical model.
As depicted by Figure 1, consider the experimental evaluation design, where two sets of sites exist, with respective designs as follows: • Design in Sites that Randomize Individuals to an Enhanced Treatment Group. Individuals in a subset of sites are randomly assigned to one of three arms: a standard treatment group, an enhanced treatment group (that receives the standard treatment plus an enhancement component), or a control group (that has no access to the program). For example, HPOG’s emergency assistance enhancement added access to emergency funds to meet needs stemming from imminent eviction from housing, utility shutoff, vehicle repair needs, and childcare needs to the standard intervention. • Design in Sites that Do Not Randomize Individuals to an Enhanced Treatment Group. Individuals in another subset of sites are randomly assigned into one of (at least) two arms: a standard treatment group (where the treatment does not include the program component that was—in other sites—randomized to as an enhancement) and a control group.
3
CAMIC design requirements.

Given this design, the CAMIC method has the potential to reduce bias in cross-site non-experimental estimates of program components and implementation features by completing the following steps.
Experimentally estimate the impact of the program enhancement (e.g., the emergency assistance program enhancement) using the sample limited to sites that randomize individuals to an enhanced treatment group using three-armed random assignment (as depicted in the top panel of Figure 1). An experimental estimate of the impact of the program enhancement can be computed in these sites by comparing mean outcomes for individuals randomly assigned to the enhanced treatment group with mean outcomes for individuals randomly assigned to the standard treatment group. More specifically, we can produce an experimental estimate of the enhancement in these sites by estimating the following model In this equation, Y is the outcome of interest for individual i in site j; Estimating equation (1) through linear regression, we obtain—among other things—an estimate
Definition of Model Terms.
In this step, we calculate several non-experimental estimates of the contribution of the program enhancement to impact magnitude using cross-site variation in whether the enhancement is included in a given site’s program (this follows the Multi-level Modeling of the Contribution of Program Characteristics to Impacts Method from above). To produce this estimate, the sample is limited to (1) the control and enhanced treatment arms from sites that conducted three-armed random assignment using the enhancement and (2) the control and standard treatment arms from sites that did not randomize to the enhancement.
4
We use variations of the following model to produce non-experimental estimates of the contribution of the enhancement component to impact magnitude In this equation, Y is the outcome of interest for individual i in site j and Estimating equation (2) through linear regression, we obtain a non-experimental estimate of the influence of the enhancement component on impact magnitude Bias in the estimates of the influence of the enhancement component Importantly, in this step, we produce several different estimates of the non-experimental estimate of the influence of the enhancement component on impact magnitude
The next step in implementing the CAMIC method is to identify the version of the equation (2) model from Step 2 that minimizes the difference between the experimental estimate of the program enhancement’s effect computed in Step 1 and non-experimental estimate of the program enhancement’s effect computed in Step 2 (using the best “similarity” metric available from WSC scholarship; e.g., Steiner & Wong, 2018), where bias is measured subject to sampling variability as
Apply the “best” model from Step 3 to estimate the contribution of other program characteristics of interest
CAMIC Method Performance
Bell et al. (2017) report on simulations that investigate whether the CAMIC method is expected to accomplish its goal of reducing bias in non-experimental estimates of the influence of program characteristics on a program’s total impact. The simplified framework used for the simulations focuses on the site-level relationship between the true impact of a program ( • • • • • • • •
As is the setting for the CAMIC method’s potential application, suppose we observe an unbiased measure of the true value of the enhancement component
As discussed above and shown in Appendix A, bias arises in non-experimental cross-site estimates of the influence of program characteristics on impact magnitude when the program characteristics are correlated with site-level factors that also influence impacts but that are omitted from the analysis. Similarly, in this simplified framework the key parameters that determine the bias are the correlations among the three observed program characteristics (the enhancement component, focal component, and the bias-reduction component), the correlation between each observed program characteristic with the unobserved factor, and the true influence of the bias-reduction component on impact. To calculate bias, one must select a value for each of these seven parameter values.8,9,10
To understand the CAMIC method’s potential to reduce bias in our estimate of the focal component
Table 2 presents an excerpt from Bell et al.’s (2017) Exhibit 9, and it reveals the following: when the correlation between the enhancement component and the focal component is high (ρ = 0.70) and the correlation between each of these two components and the bias-reduction component is low (ρ = 0.25) to moderate (ρ = 0.50), the CAMIC method performs favorably in more than 80% of the simulations (Scenarios A and B in Table 2). In contrast, when the correlations are high (ρ = 0.70) across the board (as is the case for Scenario C) CAMIC performs favorably in only 29% of the simulations. When the correlation between the enhancement component and the focal component is low (ρ = 0.25) to moderate (ρ = 0.50), the findings are inconclusive in that the CAMIC method sometimes performs favorably and sometimes performs unfavorably (Scenarios D through M).
Results for Selected Scenarios, Focused on the Relative Magnitude of the Correlations.
Source: Excerpt from Bell et al. (2017), Exhibit 9.
In advance of implementing the CAMIC method, these observations suggest that researchers compute the correlation between the enhancement component and the focal component of interest to ensure that these two program characteristics are highly correlated. 11 If this criterion is satisfied, then the researcher could deploy the CAMIC method, calibrating the model specification using only those additional bias-reduction components that have a low (ρ = 0.25) to moderate (ρ = 0.50) correlation with the enhancement component and the focal component of interest.
These simulation findings ignore the role that variance might play in the CAMIC method. Although the CAMIC method seeks to select the model that yields the least biased estimate of the non-experimental estimate of the enhancement component, noise in the estimates may result in selecting the wrong model. This limitation of the CAMIC method is not reflected in the simulations, where perhaps further analysis or application is needed.
Discussion and Conclusion
This article has discussed the problem of bias that arises when estimating the influence of program characteristics that have not been randomized, but instead are selected and implemented in ways and in contexts that correlate with their relative contributions to overall impacts. Both multi-level analytic methods and instrumental variable estimation are approaches that analysts use to pry the lid from the black box—to use Greenberg et al.’s (1993) phrase—and attempt to ascertain what it is about a program that is responsible for its impacts.
Beyond these analytic approaches, recent years have seen an increase in the numbers of experimental evaluations that include both multiple sites and multiple treatment arms. Through design, if individuals are randomly assigned to either a standard treatment group or an enhanced treatment group that receives the standard treatment plus a program enhancement, then the bias in the estimate of the randomized program enhancement can be eliminated. Advancing further still, this kind of design offers additional analytic opportunities for bias reduction in the non-experimental estimation of the contribution of other non-randomized program characteristics to overall impacts.
Past examples of evaluations that share the multi-site plus multi-arm experimental design include the following. The U.S. Department of Health and Human Services-funded National Evaluation of Welfare-to-Work Strategies (NEWWS) covered many sites, where some sites included three-armed experiments testing competing models of welfare reforms and other sites included an enhanced treatment arm designed to estimate the contribution of case management to the treatment’s overall impact (Hamilton et al., 2001). Similarly, the U.S. Social Security Administration’s Benefit Offset National Demonstration (BOND) also included—in its national, multi-site design—a third treatment arm that randomized eligible individuals within a subset of the study sample to have access to support services (Gubits et al., 2014). The U.S. Department of Housing and Urban Development’s First-Time Homebuyer Education and Counseling Demonstration also randomized to a control group and two treatment groups, with about 5800 individuals randomized across 28 metropolitan areas, with service access to 63 local and two national service providers (DeMarco et al., 2017). These studies and others with similar designs—randomized treatments, many sites, data on varying design and implementation—have the potential to use the CAMIC method to explore bias reduction in estimates of non-randomized program features, thereby improving understanding of why or how an intervention had its observed effect. Going forward, we encourage researchers to do so, where fitting.
This article contributes to the literature in two main ways. First, it explores the use of a multi-site, multi-armed experimental design as a way to produce unbiased estimates of an enhancement component. In addition, it suggests a method that could reduce bias in the cross-site non-experimental estimates of non-randomized program characteristics, in the case where the study design includes three-armed randomization that allows for experimental analysis of the contribution of at least one specific program characteristic. Over time, CAMIC method-based tests of cross-site impact attribution specifications may yield consistent results with enough replications. Given the investment needed to carry out this kind of evaluation design (with both multiple arms and multiple sites), we recognize that it will take time to amass replicate evidence of the influence of program components on impacts. Certainly as the state of the science of within-study comparisons improves as well, we should gain knowledge regarding the calibration of non-experimental to experimental results.
The evolution of CAMIC as a method among those in the evaluation toolkit has implications both for evaluation research and for program practice. Future evaluation research should consider opportunities to test CAMIC in settings where the design is fitting: multi-site, multi-armed experiments, where at least a subset of the sites randomize individuals to either a standard treatment group, an enhanced treatment group (that receives the standard treatment plus a program “enhancement” or modification), or a control group. This design allows the researcher to compute both experimental and non-experimental estimates of the influence of the program enhancement, and CAMIC may be able to leverage that evidence in order to improve the estimation of the contribution of other non-randomized program characteristics of interest. The CAMIC method is likely to achieve its promise of reducing bias in estimates of the relationship between a focal/non-randomized program component and program impact only when the correlation between the randomized program enhancement and the focal/non-randomized program component of interest is high and the correlation between each of these two components and alternative bias-reducing covariates are low to moderate.
One important consideration for researchers interested in using data from multi-site randomized experiments to estimate the influence of site-level program characteristics on program impacts is the sample size at the site-level (Moulton et al., 2014). Because it can be expensive to conduct evaluations in many varied locations, most past experiments have been confined to a small number of sites. Having a small number of sites severely limits researchers’ ability to conduct cross-site comparisons due in part to lack of statistical power. Additionally, having a small number of sites increases the potential for omitted variable bias of non-experimental estimates due to the need to leave out key site-level variables in order to conserve degrees of freedom. Acknowledging that the minimum sample size required to achieve a given level of statistical power depends on a number of factors, including the expected effect size and the amount of measurement error, Green (1991) suggests that the sample size for regression analysis should be at least 50 + 8 m, where m is the number of independent variables (VanVoorhis & Morgan, 2007). This limitation on the number of site-level predictors included implies that analysts need to be selective in which site-level measures enter into the analytic model; and the CAMIC method provides one way to be selective.
Another consideration when designing a multi-site evaluation is the cost and related tradeoffs between competing design and analytic options. We do not recommend that the potential application of the CAMIC method be the reason that an evaluator decides to implement a multi-site, multi-armed experimental design. Instead, the policy questions are what should determine the design. In the situation where an evaluation is already designed to meet the multi-site, multi-armed criteria that CAMIC requires, the added cost of applying CAMIC is likely to be marginal. CAMIC can be thought of as another tool for “within-study-comparison” and could be a promising low-cost method for choosing an improved model specification for non-experimental analyses that inform why or how an intervention works.
Indeed, to the extent that the CAMIC method helps reduce bias in estimates of the influence of non-randomized program characteristics on impact magnitude in real-world applications, CAMIC also has implications for program practice: program managers can expect more reliable estimates of the relative contributions of program characteristics to overall impacts. For example, the proliferation of this line of analysis in practice means that the field should be able to learn how programs operate differently in strong versus weak economic conditions, which aspects of multi-faceted programs are important drivers (or suppressors) of impact, which implementation features are essential to a program’s success, and how participant compositional factors influence (or not) programs’ effectiveness. These are practical lessons that inform administrative decisions on the ground as well as higher-level policy decisions about what to fund.
Footnotes
Acknowledgments
The authors are grateful input on this draft from Larry Buron and Eleanor Harvill at Abt Associates, and from Nicole Constance, Hilary Bruck, and Amelia Popham at ACF/OPRE. Eleanor Harvill and Stephen Bell are co-authors on a prior iteration of this work, and we are grateful to their input on developing and testing the CAMIC method. The views expressed in this article do not necessarily reflect the views or policies of the HHS/ACF/OPRE.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the U.S. Department of Health and Human Services (HHS), Administration for Children and Families (ACF) (No. HHSP23320095624WC), Office of Planning, Research and Evaluation (OPRE).
Notes
Source of Omitted-Variable Bias in Non-Experimental Cross-Site Estimates of the Influence of Program Characteristics
Estimating the equation (2) model will produce estimates of the
To see how bias arises, consider the following version of equation (2) that includes a term for the omitted site-level factor F
j
and the coefficient
When this equation is estimated with maximum likelihood methods (Bryk & Raudenbush, 1992) with
