Abstract
The commenter’s proposal may be a reasonable method for addressing uncertainty in predictive modeling, where the goal is to predict y. In a treatment effects framework, where the goal is causal inference by conditioning-on-observables, the commenter’s proposal is deeply flawed. The proposal (1) ignores the definition of omitted-variable bias, thus systematically omitting critical kinds of controls; (2) assumes for convenience there are no bad controls in the model space, thus waving off the premise of model uncertainty; and (3) deletes virtually all alternative models to select a single model with the highest R 2. Rather than showing what model assumptions are necessary to support one’s preferred results, this proposal favors biased parameter estimates and deletes alternative results before anyone has a chance to see them. In a treatment effects framework, this is not model robustness analysis but simply biased model selection.
The scientific challenge that Young and Holsteen (2017; hereafter, YH) seek to address is that authors typically run many models in the course of their research but may publish only two or three preferred estimates. Often, there are many other plausible models, which may support very different empirical conclusions. Choice of model specification can drive the results of empirical research in ways not visible to readers. This reflects a fundamental a problem of model uncertainty and is closely connected to the growing “crisis in science”—the sense that much published research is tainted by false positives and not broadly credible. The YH method aspires to show more transparently what credible estimates can be found in the data, especially if one of more of the modeling assumptions is wrong.
Slez (2018; hereafter, “the commenter”) criticizes the YH robustness analysis and proposes an alternative using Bayesian model averaging (BMA) that weights the results by a transformation of model fit. The commenter’s approach seems reasonable on the surface but has fundamental flaws that undermine the purpose of robustness analysis.
First, the commenter draws on tools from predictive modeling that are inappropriate for a treatment effects framework. The goals of predictive modeling are satisfied by the model fit measures such as the BIC or R2. Treatment effects analysis, in contrast, is focused on controls, covariate balance, and the problem of omitted-variable bias: None of these concerns are addressed in the BIC or R 2. In treatment effects analysis, the goal is not to predict y but rather control for factors correlated with both x and y. An assumption that “goodness of fit” is equal to “goodness of analysis” will only hold in unusual circumstances and systematically omits critical kinds of control variables. In short, model fit is the wrong metric for evaluating different empirical models in a causal analysis framework. I discuss alternative methods of weighting models that recognize the problem of omitted-variable bias.
Second, using the posterior probability to weight models, as the commenter advocates, leads to nonsense weighting of models with remarkably similar fit. In practice, the method tends to simply select one model with the most favorable BIC and then weights all other models at zero. Rather than showing the results from alternative models, this method deletes them. This is not model robustness testing but rather knife-edge model selection.
The commenter’s proposal lies at the intersection of these two flaws: It uses the wrong criterion to evaluate models, and then based on that metric, formulates extreme judgments (weights) about which model(s) to consider. This method dissolves the problem of model uncertainty by selecting a single model that will often contain biased parameter estimates.
YH, in contrast, is not a model selection procedure because model selection imposes the assumptions that our method seeks to relax and treat as open to question. The YH method does not simplify life for researchers: It asks them to defend their preferred model in light of what other plausible models show. This is how the YH method addresses the problems of model uncertainty, transparency, and asymmetric information in social science research.
In the following response, I elaborate and explain the central points of disagreement between YH and the commenter. Addressing the commenter’s concerns speaks to a number of fundamental issues in social science research today.
Predictive Modeling Versus Treatment Effects Analysis
Regression models are “designed and estimated for a purpose,” and that purpose should be central to how models are evaluated (Hansen 2005:63). We start from a “treatment effects” framework where the goal is to understand the true causal effect of a variable of interest, using a conditioning-on-observables analysis. In treatment effects analysis, the main purpose of the model is to achieve covariate balance between treatment and control conditions. This is different from the goal of maximizing the predictive accuracy of the outcome equation. No one objects to predictive accuracy with respect to the outcome, but it is not a valid or sufficient criterion for including a control variable in a treatment effects analysis.
The distinction between causal “treatment effects” research and predictive modeling has been a long-standing source of confusion. Indeed, there is a common intuition that efforts at “prediction” and “explanation” should converge on the same answers (Hofman, Sharma, and Watts 2017; Shmueli 2010). In practice, prediction and treatment effect explanation are different kinds of questions, and they do not converge on the same model nor answer each other’s questions except under rare conditions. Treatment effects analysis uses substantively different criteria for selecting variables and thus for evaluating models than does predictive analysis—the focus is on conditioning rather than on predicting.
Consider an analysis of an outcome of interest (y) focusing on a key treatment variable (x) and a set of possible controls (
In predictive modeling, the goal is to find a prediction of the outcome
Predictive analysis calls
Treatment effects analysis is aspiring to achieve conditional independence between treatment and control conditions: The treatment variable x takes on different values “as if” by random assignment, after controlling for confounding factors. Success in meeting this goal is difficult to evaluate and is the subject of a vast literature in econometrics. In this literature, the model fit of equation (1), by itself, is not how models are evaluated.
Model Fit Is the Wrong Metric
In prediction problems, it makes sense to weight models by measures of fit, such as the BIC or the
In treatment effects analysis, the correlation
The model fit from this treatment equation can be denoted
Morgan and Winship (2007) have noted that in contemporary practice, “a focus on equations for outcomes” has led many to lose sight of what is important in treatment effects analysis (p. 13). Those familiar with matching methods for conditioning on observables will be more comfortable thinking in terms of both a treatment (or propensity) equation and an outcome equation. Ordinary least squares regression can also be written as involving an outcome (equation 1) and treatment (equation 2) equation in a process that strips out the effect of z from both equations, which better reflects the intuition of adjustment for omitted-variable bias (see Morgan and Winship 2007:136-37 for review).
By ignoring the relationships in the treatment equation (equation 2), predictive models (such as the commenter’s method) systematically fail to control for omitted-variable bias (Athey, Imbens, and Wager 2017; Belloni, Chernozhukov, and Hansen 2014; Belloni et al. 2017). What are dropped from a predictive analysis are confounding controls that have a low correlation with y but a high correlation with x. This is often a serious error; variables with a low correlation with y can still cause substantial bias in the parameter estimate if they are strongly correlated with x. Belloni et al. (2014) implement a model selection algorithm designed for treatment effects analysis when there are many possible controls. It is a three-step procedure: (1) select controls that predict the treatment variable
Intuitively, rather than focusing on the
As YH note, the most influential z controls in applied analyses are often those with the smallest correlation with y (see pp. 17, 21). It is also common that variables with the highest correlation with y have no influence at all on the estimated treatment effect (p. 27). This easily understood when thinking about the omitted-variables bias formula or the Belloni et al. (2014) selection algorithm but makes no sense when focusing on
It should be emphasized that YH do not advocate automatically including all influential controls, and I do not recommend adopting the
Bad Controls
In a treatment effects framework, there are important reasons to exclude statistically significant control variables. The problem is variously referred to as posttreatment bias, the problem of collider variables, endogenous controls, or conditioning on intermediate outcomes (Acharya, Blackwell, and Sen 2016; Cole et al. 2009; Elwert and Winship 2014; Nyhan, Montgomery, and Torres. 2017; Rosenbaum 1984). The issue is best illustrated with a vivid example.
Consider the question of gender discrimination in high-tech industries. The basic model for this is
This initially tempting conclusion of gender parity has serious problems. Their central control variables, promotions and performance evaluations, are classic posttreatment variables, under the discretion of the company and central to how men and woman can be sorted into different career tracks. An equally valid conclusion from their analysis is that the gender wage gap occurs because Google gives women bad performance evaluations and does not promote them.
Google—and the tech sector in general—seems to have a problem with nurturing female talent and appropriately recognizing the contributions of female workers. Valid controls would be measures of skill and ability prior to entering the company, and how long they have been with the company. But there are good reasons to think that negative evaluations and nonpromotion are the result of being female in the tech sector not an exogenous explanation of their low pay. By controlling for subjective performance and promotions, the Google report controlled away the pathways by which gender discrimination occurs.
In a causal analysis of gender discrimination, the goal is not to simply predict the wages of Google employees with the best possible model fit. Indeed, the variables that strongly predict wages (evaluations and promotions) confound the analysis of discrimination. If we give greater weight to the models with higher
This is not a secondary concern in social science nor a “misguided fear” as the commenter suggests. A recent review of articles in top political science journals over five years concluded that 40 percent of articles included “obvious” posttreatment controls (Acharya et al. 2016). This suggests the problem is widespread and too few researchers consider the problem of posttreatment bias.
There is no simple solution to this problem of endogeneity—where the controls factor out the mechanisms we are trying to understand. But we cannot resolve the conflict by only showing models with a high
Excluding Significant Controls Often Make Sense
Bad controls are part of a broader reason to routinely consider excluding controls that are statistically significant. Consider a general form of model uncertainty in which the “true” set of control variables is neither known nor observed so that none of models under consideration entirely captures the causal forces at work. In this general case, the omitted-variables bias formula becomes much more complex. It is very hard to say which imperfect model gives the “best approximation” of the true model. If there is only one omitted control variable, z, then adding it will reduce (to zero) the omitted-variable bias
Learning From the Data
The commenter argues that “failure to account for fit is tantamount to a refusal to learn from the data” (p. 18). This is a naive statement that speaks to how much misplaced weight the commenter puts on the model fit as well as to his narrow and mechanical concept of “learning.” The
The influence analysis in YH uses and reports full information from all coefficients, for all explanatory variables, from all possible model specifications. It shows how the estimated treatment effect (
In robustness analysis, the influence analysis is more informative than the overall model robustness statistic and much more informative than model fit statistics. Through the influence analysis, analysts and readers learn from the data by seeing what model features drive the estimated treatment effect. As we will see, the commenter’s method sharply reduces the information available to analysts and readers.
Averaging With Zero Weights: Model Robustness Versus Model Selection
The model robustness framework ultimately calls on authors to justify their model selection criteria—to defend their preferred model—in full light of what other reasonable estimates can be found in the data. In contrast, the commenter seeks to impose the selection criteria of the
The BMA technique does not answer questions about robustness. The technique generates weights
Few readers will have immediate intuition of what this formula does. When the commenter describes the method, it sounds like a modest reweighting of models to take into account the
This pattern of essentially selecting a single model is seen clearly in the commenter’s Table 2, which shows a small-scale illustration of the model weighting procedure. Table 2 shows eight models with estimates ranging from −0.18 to +4.79. One model shows no evidence at all that gender is a factor in mortgage lending. Another model finds a five-percentage point higher mortgage approval rate for women over men (model 8). These models offer very different substantive conclusions, so the results clearly depend on how the model is specified. However, the commenter concludes that there is virtually no model variance, meaning that the results do not depend on decisions about which variables to include.
This robustness conclusion emerges because six of the eight regression estimates are given a weight of .000 and thus excluded entirely from consideration. Model 8, with the estimate showing the largest degree of gender disparity (in favor of women), is given 93.6 percent of the weight, and model 7 is given the remaining 6.4 percent of the weight, although the estimate is notably smaller. The difference in the
It is not credible to say one is considering eight plausible models, when six of them are given literally zero weight. The commenter claims that this approach is “easily justified in terms of basic probability theory” (p. 2). It is better described as a nonsense weighting of closely comparable models with no substantive justification at all. This is not model averaging or model robustness analysis—it is model selection. Moreover, it is not testing to see what happens if one of more of our modeling assumptions is wrong. It is just discarding models that do not maximize
The commenter is aware that his method frequently deletes the results from every plausible model but one. He writes that the difference between YH and his method is mostly driven by “the extent to which the weights used for model averaging tend to concentrate on a single model” (p. 12). I agree this is a key problem: The commenter is doing model selection, while YH is doing model robustness.
Any approach which assigns 93.6 percent of the weight to the single model with the highest
Learning in Two Stages
In a sense, both YH and the commenter’s BMA are two-stage procedures, using the same first stage but diverging dramatically on the second stage. Both begin by generating a plausible model space and estimating all models to obtain a distribution of estimates.
At the second stage, the commenter’s BMA approach “learns from the data” by deleting the vast majority of the modeling distribution, leaving the estimate from the model with the highest
The second stage of YH, in contrast, proceeds to the influence analysis, which is a meta-regression showing which elements of the model space (such as which control variables) explain why different estimates are possible. 5 Here, analysts learn from the data by unpacking the model assumptions and seeing which aspects of the model specification are critical to the results. This does not automate model selection, and it remains the responsibility of analysts to explicitly invoke and justify any model assumptions that are necessary to sustain their preferred result.
If one desired a model selection routine from YH, it would be based on the influence analysis. The selection criteria would be: All variables that significantly influence the treatment effect should be included, while others can be dropped (or not) without biasing the parameter estimate. This requires assuming that all the controls are strictly exogenous and that there are no other omitted variables or other specification errors. Sometimes authors and readers will be comfortable with that assumption, other times not. One can imagine an option that compares the full modeling distribution to the restricted distribution of results (including all influential controls) under the complete exogeneity assumption. Users can easily implement this approach after examining the influence results, and an example similar to this is offered in YH (p. 22, figure 3).
Conclusion
As Heckman (2005) has noted, there is no “assumption free” way to conduct an empirical analysis. It is not “embarrassing” to show that an analysis depends on certain assumptions. On the contrary, social scientists should be open and transparent about what assumptions are necessary for their conclusions. Today, too many empirical papers bury the evidence about which assumptions matter. Often, they bury model uncertainty in an effort to imply that their conclusions are true under all possible scenarios. Model uncertainty represents a limit to our knowledge that applied researchers are reluctant to admit.
Embracing transparency and showing what estimates are available—as well as which aspects of model specification affect the results—is central to thinking through the challenges of causal inference in a conditioning-on-observables analysis.
If a result is not robust to model specification, this does not mean we have failed to learn anything. Instead, it means we must focus on the influence analysis, which shows which model assumptions are needed for a strong conclusion. If those necessary assumptions are credible, authors should make an explicit case to readers justifying the model assumptions. We should not simply hide this information from readers, so they do not know that the results are dependent on key assumptions. And we should not hide this from readers by using opaque, scientific-sounding methods that delete the alternative models while claiming to average across them.
We need to guard against naive enthusiasm for more complex methods (Glaeser 2008). The YH method asks authors and readers to consider all the estimates that can be credibly supported by the data. This may strike some as too simple. But complex weighting schemes based on incorrect criteria are not an improvement. Methods which simply show the available estimates should not be rejected because they are easy to understand. As one philosopher put it, sometimes we can see a lot just by looking.
I wish to emphasize that the commenter arbitrarily limits his analysis and critique to the choice of alternative control variables, while the YH framework extends to alternative variable definitions (Brady, Beckfield, and Seeleib-Kaiser 2005), alternative cleaning and coding strategies (Leahey, Entwisle, and Einaudi 2003), alternative standard errors, data sources, and different estimation commands and functional forms (Kane et al. 2013). This, I believe, is the research frontier of robustness testing.
I am passionate about the potential of model robustness analysis to produce research that is more transparent and more convincing than current practices in sociology. The tremendous growth of computational power has given rise to serious problems of asymmetric information between analysts and readers (Young 2009). The result is a growing crisis of confidence in science today, owing to the fact that analysts, but not readers, can see and select a preferred estimate after running large numbers of similar models. We need better methods of showing what results are possible under a reasonable set of model assumptions. YH worked hard to develop a robustness command that would be easy to use and can perform transparent robustness analysis across many aspects of model specification in a wide range of empirical settings. The project is ongoing (Stewart and Young 2018; Young 2018; Muñoz and Young 2018), and I welcome input and ideas on how to make it better and more useful. However, a proposal that casts aside the objectives of robustness analysis and conducts single model selection based on the
Footnotes
Acknowledgment
The author would like to thank Michelle Jackson, Sheridan Stewart, John Muñoz, and Katherine Holsteen for helpful feedback and suggestions.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
