Abstract
Measuring the effect strategic choices have on electoral outcomes is problematic, because this requires an assessment of the outcome under a counterfactual that is not observed. To overcome this problem, we extend the synthetic control approach for causal inference to circumstances with multiple treated cases and use it to estimate the effect of vice-presidential candidates on their home states’ vote. Existing research has concluded that vice-presidential candidates have little effect on the outcome of elections in their home states. However, our results from elections spanning 1884-2012 suggest that vice-presidential candidates increase their tickets’ performance in their home states by 2.67 percentage points on average—considerably higher than previous studies have found. In addition, our results suggest that the vice-presidential home state advantage (HSA) could have swung four presidential elections since 1960, if presidential candidates had chosen running mates from strategically optimal states.
Introduction
Political actors involved in presidential election campaigns make a variety of strategic choices with the hope that these will help them win the White House. 1 Although political scientists have shown that the “fundamentals” (e.g., the state of the economy, whether the country is at war, etc.) are the strongest predictor of the outcome of an election (see, for example, Sides & Vavreck, 2013), there is evidence that campaign strategy can be important, particularly in close elections (Erikson & Wlezien, 2012; Shaw, 2006). Among the many different strategic choices made in presidential election campaigns is the selection of a vice-presidential candidate. One key question in this regard that has puzzled political strategists is whether there is any electoral benefit to selecting a vice-presidential candidate from a particular state. 2
A series of studies devoted to this question have concluded that running mates generally do not improve the ticket’s chance of winning in their home states. Attempts to measure the vice-presidential home state advantage (HSA)—the number of points of the two-party vote a vice-presidential candidate adds to her party’s performance in her home state—have consistently shown that vice-presidential candidates do not have much impact (see, for example, Devine & Kopko, 2011, 2013; Dudley & Rapoport, 1989). The most recent estimate of the average vice-presidential HSA is a mere 0.69 percentage points, a number statistically indistinguishable from zero (Devine & Kopko, 2013). 3 Importantly, this conclusion has not gone unnoticed outside the realm of political science: Republican strategist Karl Rove cited Devine and Kopko’s (2011) article as evidence that Mitt Romney should base his 2012 running mate selection on the candidates’ experience and potential strengths as vice-president post inauguration, rather than seek out a viable running mate with roots in a battleground state (Rove, 2012).
In this article, we question the validity of the formula most commonly used to measure the vice-presidential HSA and propose an alternative measure relying on the synthetic control method. Our suspicions of the traditional formula are founded on a series of individual vice-presidential HSA measurements that appear to be implausible estimates of the actual effect these vice-presidential candidates had on their home state’s vote based on historical assessments of the relevant elections and candidates. For example, although Lyndon Johnson’s selection by John F. Kennedy in 1960 is frequently described as a clever tactical move to gain support in an increasingly less Democratic Southern state, the HSA measured by Devine and Kopko puts Johnson’s contribution at −14.72. If correct, this effect would indicate that Johnson cost Kennedy votes in Texas and that JFK would have increased his vote share by nearly 15 percentage points in Texas by not selecting Johnson. This case, and other individual measurements, which seem unrealistic based on historical assessments, suggests that there is a serious flaw in the current measurement of the vice-presidential HSA. 4 We argue that this flaw stems from the formula’s failure to account for the fact that vice-presidential candidates are often selected from states in which their party has recently seen a decline in support (see Figure 1). In other words, a number of vice-presidential candidates violate the parallel trends assumption on which inference in the difference-in-differences framework the current literature relies on is based. In these cases, the HSA measure compares realized vote totals with an artificially high “expected” or counterfactual result and therefore underestimates the effect vice-presidential candidates have on the vote in their home states.

Treated states’ average two-party vote share in comparison with the U.S. average over nine pre-treatment and treated elections.
We propose a new approach to estimating the vice-presidential HSA, which applies the synthetic control method (Abadie, Diamond, & Hainmueller, 2010, 2015). Using this method allows us to estimate the vice-presidential HSA without the bias that would otherwise be caused by time-varying heterogeneity across states. Our approach produces an average vice-presidential home state effect of 2.67 percentage points—a number that is statistically distinguishable from zero and close to the average 3.61 percentage point increase presidential candidates are purported to produce in their home states (Devine & Kopko, 2013). Substantively, a 2.67 percentage point increase in a ticket’s share of the two-party vote—ceteris paribus—would be enough to swing nearly 20% of state-years in our data from one party to another. These results indicate that vice-presidential candidates, although on average less influential than presidential candidates in their home states, can in fact affect election outcomes and that the commonly accepted narrative that running mates lack the capacity to add to their ticket’s performance in their home state is incorrect. Indeed, we show that a carefully chosen vice-presidential candidate from an important battleground state could plausibly have altered the outcome of four presidential elections in our sample period.
Localism and the Vice-Presidential Home State Advantage
The hypothesis that candidates for elective office should perform better in their home state is derived from V.O. Key’s concept of “localism.” Key (1949) noted that in the South, “candidates for state office tend to poll overwhelming majorities in their home counties and to draw heavy support in adjacent counties” (pp. 37-38). Early quantitative tests of Key’s theory of localism include Black and Black (1973), who found that George Wallace performed better in his home county than should be expected based on levels of Black and working class voters, and Tatalovich (1975), who found that candidates in local races in Mississippi received a higher percentage of the vote in counties closer to the county in which they reside.
On a national level, Brogan (1954) argued that “state pride and the interests of the state machine” produce a presidential HSA (Brogan, 1954, p. 197). Lewis-Beck and Rice (1983) were the first to measure this presidential HSA, relying on the following formula:
in which
Dudley and Rapoport (1989) and Devine and Kopko (2011, 2013) use the same formula to measure the vice-presidential HSA. Using the 2008 election as an example, the vice-presidential HSA for Joe Biden can be calculated as
suggesting that the Democratic ticket received a 6-point bump in Delaware, relative to a counterfactual Democratic ticket in which Joe Biden retained all of his experience and characteristics but was not from Delaware.
By Lewis-Beck and Rice’s measure, however, Biden appears to be an anomaly. Both Dudley and Rapoport and Devine and Kopko find a very small average vice-presidential HSA (0.3 and 0.69 points, respectively). Almost without exception, other attempts to measure the effect of vice-presidential candidates on election outcomes have come to similar conclusions. 5 However, although the vice-presidential HSA appears inconsequential, this average hides a wide spread of results, ranging from Ed Muskie’s (D-ME) 13.32-point increase in 1968 to Lyndon Johnson’s (D-TX) 14.72-point decrease. Although variation across candidates is to be expected, some of these individual measurements are unlikely to be correct representations of the actual HSA produced by vice-presidential candidates. Indeed, politicians who were widely regarded as popular in their home states are, in the Lewis-Beck and Rice HSA measurement, argued to have produced negative (Johnson, John Nance Garner [D-TX], John Sparkman [D-AL]) or surprisingly low (Calvin Coolidge [R-MA]) effects. Each of these vice-presidential candidates had a substantial voting base and a proven record of winning statewide elections by landslide margins in their home states. 6 As such, it is unlikely that presidential candidates would have performed better without these highly popular running mates.
Home-State Advantage as Difference-in-Differences
We argue that the Lewis-Beck and Rice measurement, when used to estimate the vice-presidential HSA, requires onerous assumptions that are unlikely to be met in practice. To start, consider Equation 1 above: A slight rearrangement of terms reveals that Lewis-Beck and Rice’s measure is precisely the difference in two differences: the difference between
First, previous researchers assume that an appropriate control group for states “treated” with a vice-presidential candidate selection consists of an average of all other U.S. states. Second, they assume that any difference between treated and control units is time-invariant. 8 These two assumptions are related: Comparing treated states with national averages is sensible only insofar as treated states’ pre-treatment trends closely track the pre-treatment trends in the national average. Figure 1 provides evidence that this is not true, by plotting the share of the two-party vote in treated states versus the U.S. average in the pre-treatment and treated periods. Figure 1 demonstrates that vice-presidential candidates often come from states whose support for their party was on a downward trajectory. Thus, the pre-treatment trends in treated states are not well-approximated by all other U.S. states over the same period. 9 If treatment and control groups do not share parallel trends in the outcome variable prior to treatment, there is no reason to believe that they would share parallel trends post treatment, in a counterfactual world where the treated group went untreated. Without satisfying this assumption, inferences based on difference-in-differences—even the informal versions used by researchers in the HSA literature—are biased.
Previous researchers amplify both problems by aggregating results from five prior elections into a single pre-treatment period. This not only assumes that the difference between, for instance, Texas and the remaining U.S. states will stay constant from one election to the next; it also assumes that the average difference between Texas and the other U.S. states over the past five elections will remain in the treated election, obscuring shifts in vote totals that differentially affect treated units. The logic of combining several elections is that it reduces the effect of idiosyncratic events on estimates of HSA. In practice, though, combining election results going back 20 years into a single pre-treatment level ignores seismic shifts in American politics that affect states in different ways. 10
To illustrate the problem, Figure 2 graphs the Democrats’ Texas and national vote share, respectively, from 1928 to 1960. When calculating Lyndon Johnson’s HSA in 1960, Dudley and Rapoport and Devine and Kopko combine elections from 1940 to 1956 into a single pre-treatment level. By including 1940, 1944, and 1948 in the term

Democratic performance in Texas and the United States in the run-up to the 1960 election (1928-1960).
A Synthetic Control Approach
Our alternative to the current HSA measurement is to apply the synthetic control method developed by Abadie et al. (2010) to a situation with many treated units. The synthetic control method was designed for the purposes of estimating causal effects in comparative case studies where only a single unit was treated. 12 By weighting a set of control units—in this case, states—to closely follow the pre-treatment outcomes of a single treated unit, we are able to approximate the key counterfactual for causal inference: what the treated state would have looked like in the absence of treatment. In the case of Texas in 1960, we seek a weighted combination of control states whose Democratic vote shares most closely match that of Texas prior to 1960. Assuming that a set of control states and weights exists that approximates pre-treatment Texas well, we can use the actual Democratic vote share in that “synthetic Texas” in 1960 as an estimate of what the Democratic vote share in Texas in 1960 would have been in the absence of treatment.
Consider the simplest case with 50 states, indexed by
where
The primary improvement Abadie et al. (2010) offer over the standard difference-in-differences model is to allow for the effect of unobserved confounders to vary over time.
13
A set of control units and weights, which satisfies two conditions, is an unbiased estimate
where
In practice, small amounts of noise remain between the treated and synthetic control units even in very well-matched cases. Figure 3 shows this problem in the case of Maine: Prior to the treated election in 1968, Maine’s Democratic vote share was slightly higher than that of “synthetic Maine.” This pre-treatment difference would bias our result upward if it remained constant into the treatment period. Correcting for such noise is possible by differencing out the pre-treatment difference from time

Democratic performance in Maine relative to a “synthetic Maine,” 1936-1968.
Note that, by differencing out pre-treatment noise, we reintroduce the assumption that there is no time-varying heterogeneity between treated and control groups. Therefore, if the treatment group is subject to idiosyncratic shocks between times
We apply Abadie et al.’s (2010) method by constructing synthetic control units for each of our 53 treated cases individually, restricting the set of possible control units to states, which were not treated during the pre-treatment period under consideration. For instance, in the case of Delaware in 2008, we seek to construct a synthetic control unit that closely matches Delaware over the period 1976-2004. The donor pool excludes any states treated by a major-party presidential or vice-presidential candidate in 2008 (Illinois, Alaska, Arizona). We eliminate a further set of states (14 in total) because they had major-party candidates during the pre-treatment period 1976-2004. 18 Each match in our sample can be judged for quality visually or by using the metric Mean Squared Prediction Error (MSPE) during the pre-treatment period. 19 To illustrate the high quality of matches that we generate, we ranked our 53 cases on MSPE and plotted the pre-treatment trends for the case that fell at the 25th percentile. Figure 3 shows that even a synthetic control on the lower end of the range (25th percentile) is well-matched.
This matching strategy—which combines fractions of different control states to compose a single synthetic control “state” for each treated state—may appear non-intuitive or even objectionable. Our comparison, in the case of a single treated state, is no longer between two actual states; rather, it is between a treated state that exists and a conglomeration of one or more existing states. We note that other empirical methods are equally susceptible to such critiques, including workhorse methods such as cross-sectional matching and linear regression. The synthetic control method makes weighting explicit and transparent—indeed, by reviewing the supplemental appendix, one can see the precise weights that our method generated for each case—in contrast to many common alternative methodologies.
Overall, our empirical strategy deals with each of the major criticisms of existing studies we leveled in the previous section. We minimize assumptions about temporal changes, inherent in prior studies, by using a single pre-treatment election in calculating the difference-in-differences estimate, rather than five. More importantly, we are able to offer close pre-treatment matches between treated units and their respective synthetic control units, which prior studies that use all untreated states as a control group cannot. This increases the probability of satisfying the key identifying assumption required for difference-in-differences.
Politicians Playing at Home
Our outcome variable is the two-party vote share for Republicans and Democrats in each state, over Presidential elections from 1872 to 2012. 20 Our quantity of interest is the SATT, which we calculate by averaging the unit-specific effects of 53 treated cases. We also discuss alternative estimators, including the simple difference between treated and synthetic control units, and the MSPE ratio proposed by Abadie et al. (2010). We prefer the post-processing difference-in-differences on both theoretical and substantive grounds, but report inferences from alternatives in the supplemental appendix—in short, our conclusions are unchanged.
As discussed in Abadie et al. (2010), we choose a set of predictor variables based on their power to predict vote share. Importantly, predictor variables are distinct from control variables in regression models or even matching variables in standard matching models. Unlike propensity score or other forms of matching, which generally maximize balance between treated and control groups on a set of covariates thought to influence selection into treatment, synthetic control methods maximize balance on pre-treatment values of the outcome variable itself. Predictor variables are balanced only to the extent that they are good predictors of the outcome variable and are also used to check fit between the treated and synthetic control units. 21
We emphasize this point because the conditions for credible identification are distinct from those that apply to cross-sectional matching models. Our approach does not require us to assume away selection on unobservables, and the predictor variables are not included to adjust for confounding. Rather, predictor variables provide the synthetic control algorithm more information to use in its minimization function, and provides a second check on the quality of matches produced. 22
Data covering the long time series that we study are sparse. Even attempts to forecast modern Presidential elections use Spartan models (see, for example, Lewis-Beck, 2005). Our inclusion of the late 19th century and early 20th century means that our choice of predictor variables is necessarily constrained. We use a set of demographic controls constructed from U.S. Census records, as well as Census variables that capture the economic structure of individual states. We include the percentage of African Americans and the percentage of the population that was foreign-born, because these groups have historically voted differently from Caucasians and the native-born. We also include the percentage of rural inhabitants, the total acreage used in agriculture and the average acreage of farms, as rural states often have interests distinct from urban, industrialized states. The number of manufacturing establishments in each state further captures economic structure. For political variables, we include the two-party vote share in gubernatorial elections, because swings in pre-treatment Presidential vote share may not reflect the actual partisanship of a state (Congressional Quarterly Press, 2010). Finally, we also include indicator variables for census region and division, as regional differences are particularly pronounced throughout our sample period. 23 As mentioned previously, the particular set of predictor variables is less important in synthetic control models, relative to matching models. Nonetheless, in the supplemental appendix, we report results from several alternative specifications; our results are robust—indeed, virtually unchanged even at the level of individual cases—to alterations in the set of predictor variables used. Balance plots for our predictor variables are similarly available in the appendix.
We show our primary results visually in Figure 4, which plots the gap in two-party vote share between the treatment group and their synthetic control units over time. We estimate that vice-presidential candidates increase their tickets’ two-party vote share by 2.67 percentage points in their home states measured across this entire period, demonstrated by the substantial jump in the gap between treated and synthetic control units at time

The gap in two-party vote share between the treatment group and their synthetic control units (solid line).
In Figure 5, we report these primary results and a series of subgroup analyses as point estimates, separating our treated cases by the party of the candidate, by the size of the state, and providing estimates for cases where the treated state is electorally competitive—the effect size does not vary appreciably when we limit the analysis to competitive elections. 24

Aggregate and subgroup estimates of the vice-presidential home state advantage using synthetic controls.
When aggregating cases matched individually, it is possible that our results could be driven primarily by those cases that are poorly matched. To ensure that our results are not an artifact of merely poorly matched cases, we divided the sample and recalculated the sample average treatment effect on the treated among only well-matched cases. These results also appear in Figure 5, which shows that our estimated effect size actually increases when we limit the analysis to the half of the sample with the best pre-treatment matches. Our estimate is also not driven by large positive outliers: The median effect size is 2.43%.
As Wawro and Katznelson (2014) have pointed out, averaging effects over long historical stretches runs the risk of obscuring crucial temporal changes that influence political outcomes over time. To study the effect of the vice-presidential HSA over time, we present the average HSA over each of the “party systems” covered in our sample in the bottom portion of Figure 5. 25 As the results illustrate, HSA varies across time periods, with the lowest and highest average HSA in the sectional party and New Deal party systems, respectively. However, much of this variation may be an artifact of small sample sizes, because we estimate effects for just three sectional period elections.
Another possible cause for concern is strategic allocation of campaign resources by opposing campaigns. In cases in which vice-presidential candidates were selected with the specific intent of helping their ticket win their home state, the opposing ticket might counter by increasing their campaign activities in the run-up to the election, eliminating or reducing the VP HSA in such cases. If true, this would mean that the average HSA we find—although substantial—would only apply in cases where this advantage is not particularly relevant—that is, states that are either safe for one or the other party or that do not carry many electoral votes. 26 To provide a rough check of this argument, we include HSA estimates for states that could reasonably be considered safe or competitive at the time, and states that had a higher-than-average number of electoral votes. 27 As can be seen in Figure 5, candidates from states that both carried many electoral votes and were competitive—the combination most likely to include those candidates selected with the HSA in mind—do, indeed, have a slightly lower average effect, but the effect is still substantial.
As with traditional matching (Abadie & Imbens, 2008), the proper method for conducting statistical inference in synthetic control models is unknown. Abadie, Diamond, and Hainmueller suggest using a series of placebo tests to calculate non-standard p values for inference; this same idea has been used in recent working papers to calculate non-standard confidence intervals in the case of multiple treated units (see, for example, Acemoglu, Johnson, Kermani, Kwak, & Mitton, 2013). Although using methods such as those of Acemoglu et al. (2013) produces very similar conclusions regarding uncertainty, we prefer to situate inference within the randomization inference framework pioneered by Ronald Fisher, as applied to observational studies by Rosenbaum (2002). 28
Fisher’s framework has several advantages. First, its properties under a variety of distributional assumptions and in varying sample sizes are well-understood. 29 More importantly, Fisher’s framework utilizes the control group that is best justified by our research design. Abadie et al.’s (2010) suggestion differs in that it relies on placebo tests among the entire sample of possible control units (i.e., the entire donor pool), implicitly assuming that all units in the donor pool are equally suitable counterfactuals for the treated unit. This assumption is violated in practice—as we showed in Figure 1. Not all states are good counterfactuals for our treated states, illustrated by the fact that pre-treatment trends in our treated units are poorly approximated by the pre-treatment trends in an average of all available control units.
In contrast, we apply randomization inference among the treated and synthetic control units only, because the latter are the best counterfactuals for their respective treated units. Treatment assignment is plausibly ignorable between these two groups, but not between the treated and all possible control units. As a result, our approach to inference should be more conservative and more accurate than that advocated by Abadie et al. (2010). 30
Fisher’s approach to inference is straightforward. 31 Consider inference as a missing data problem under the potential outcomes framework. We know the realized potential outcomes for treated and control units, because they are observed. We do not know the potential outcomes for those same units under the counterfactual condition. Under a sharp null hypothesis of no treatment effect for any unit, we can fill in the missing potential outcomes under the counterfactual, because they are assumed to be precisely the same as the realized, observed potential outcomes.
What is the distribution of possible treatment effects under such a null hypothesis? We can calculate the randomization distribution (or empirical null distribution) by randomly assigning units to placebo treatment and control conditions and calculating the treatment effect obtained. By repeating this process for many possible realizations of randomization, we obtain the empirical null distribution. Then, a comparison between our treatment effect estimate and the distribution allows us to perform statistical inference by asking how likely our result is under the null—how far into the right or left tail of the empirical null distribution does our result fall?
Our research design mirrors a matched-pair design discussed by Rosenbaum (2002). We have 53 matched pairs, each consisting of a treated and synthetic control unit, with treatment assigned within each pair. To calculate the randomization distribution, we randomly assign placebo treatment within each pair, assigning each unit probability of treatment equal to .5. For each iteration, we calculate the difference-in-differences test statistic. A paired design with 106 cases produces
One advantage that we derive by using the synthetic control method is that the estimated effect from each case is individually meaningful; in expectation, each is an unbiased estimate of the unit-specific treatment effect. As one would expect, these estimates vary from −12 points to as high as 16 points. We report each candidate’s estimated treatment effect in Table 1. 33
Unit-Specific Estimates of the Vice-Presidential Home State Advantage.
Note. State-years with no treatment effect were excluded from the analysis based on our sample inclusion rules.
Despite our earlier objections to the method for statistical inference proposed by Abadie et al. (2010), it has the significant advantage of being estimable for singular cases. Because any approach to inference at an aggregate level—including our own—assumes that treated and non-treated cases across different years are drawn from the same “treated distribution” and “non-treated distribution,” respectively, we use Abadie et al.’s (2010) method to provide estimates of “extremeness” for each individual case, relative to untreated cases in the same election year. Again, we note objections to this approach, particularly that unit-specific p values are underestimated, because treatment is not ignorable across units in the same election year. This objection motivated our decision to use randomization inference, but we report p values calculated using Abadie et al.’s (2010) method in the interest of estimating case-level uncertainty, despite its drawbacks.
We rank each treated case compared with all non-treated cases on the basis of effect size in the same election year and calculate the p value as the rank over the total number of cases in that year, applying Abadie et al.’s (2010) method to each treated case in our sample. Under the null hypothesis of no effect, the distribution of p values calculated in this way should be uniform on the unit interval. Figure 6 plots our 53 effect estimates and their p values and shows that the distribution is heavily skewed toward positive effect estimates and small p values—highly suggestive of a positive average effect. 34

Distribution of Effect Estimates and p Values for Each Treated Case.
Falsification and Robustness Checks
Our investigation shows a positive and substantial effect of vice-presidential candidates on the two-party vote share in their home states. Our conclusions differ markedly from past findings. To increase confidence in our results and guard against specific threats to inference, we subject our findings to a series of falsification tests and robustness checks.
In our first falsification test, we exploit the fact that a state’s treatment in our study is time-limited. Vice-presidential candidates are on the ticket for one or two terms and then either exit the ticket or run as presidential candidates themselves. If a state is represented on the ticket in just one election, any increase we observe from the HSA should dissipate in the next election. Among our sample, there are 25 such cases, in which a state was represented on one of the major-party tickets in precisely one consecutive election. In Figure 7, we restrict the sample to these “singletons” and plot the gap in two-party vote share in times

The gap in vote share between treated and synthetic control units, among those treatment cases which lasted a single election cycle.
An additional threat to valid inference comes from interpolation biases. In practice, a synthetic control unit could be composed of highly dissimilar states that, when combined, mimic the pre-treatment outcomes in the treated unit. A simple example of such a phenomenon would be to combine highly Democratic Washington, D.C. and highly Republican Utah to create a synthetic control unit for Wisconsin in 2012. Abadie et al.’s (2010) suggestion in this regard is to trim the pool of donor units, limiting the pool to those states that would provide sensible matches to the treated unit. We trim on the basis of vote share in
A further source of potential bias in our results arises from the regional appeal a running mate might have. For instance, those who argue LBJ was selected as a running mate who could deliver Texas often suggest this was part of a larger strategy by the Kennedy campaign to improve the ticket’s chances across the South, violating the no-spillover assumption needed for valid inference. However, to the extent that the states that make up our synthetic control units are often geographically proximate to the treated unit, this possibility would only serve to reduce the estimated advantage vice-presidential candidates provide in their home states. If, for example, Johnson improved the Democrats’ performance in Mississippi and Florida—the two states that comprise our synthetic control unit for Texas—then our results actually underestimate the vice-presidential HSA.
The results we have presented are both substantively meaningful and statistically significant according to our permutation tests. We have also shown that the treatment effect is temporally bound in cases in which treatment lasts for a single electoral period, and that our results are robust to—or even underestimated because of—the most likely sources of bias. As a final check on our results, in the next section, we investigate several cases in which our results diverge most dramatically from the findings of past studies.
Sensitivity Analysis of Divergent Cases
How do our individual estimates compare with those of previous studies? As would be expected, in several cases, there are sizable differences between our estimates and those of Devine and Kopko (2011). To further check the quality of our estimates, we subjected our results to two additional tests concerning “divergent cases”—those cases in which our estimates differ most dramatically from those of past studies. For eight divergent cases, we use a simple sensitivity analysis suggested by Abadie et al. (2015). For each case, we take the list of states that were included in the synthetic control, iteratively drop two at a time, and re-estimate the unit treatment effect. 38
Consider Massachusetts in 1920, whose synthetic control unit consists of New Hampshire, Rhode Island, Vermont, and North Carolina, in order of decreasing weight. In a first iteration, we remove New Hampshire and Rhode Island from the pool of possible donor states to the synthetic control and re-estimate; our effect estimate drops from 8.2 points to 5.6 points, and the synthetic control is constructed from Connecticut, Vermont, Texas, and Georgia, respectively. In a second iteration, we remove New Hampshire and Vermont from the analysis, and allow Rhode Island to return to the donor pool. We repeat this process, removing two states at a time from the pool, until we have exhausted all possible combinations.
In general, the primary synthetic control units that we report in our main results are prima facie logical. Arbitrary limitations on the states included in the donor pool, by design, never improve the pre-treatment fit between treated and control units. 39 However, our confidence in this formal approach to donor selection is buttressed by the fact that we generally prefer our primary synthetic control unit to the available alternatives on qualitative, case-specific grounds as well.
The results from our sensitivity analysis are summarized in Figure 8, which—for each of the eight cases—plots our primary estimate in a gray circle and each alternative estimate in hollow circles versus estimates generated by Devine and Kopko (2011) in gray squares. We find very few substantively important differences in estimates across available alternatives. In most cases, the alternatives return effect sizes similar to the primary results we report; more importantly, there does not appear to be any systematic bias in our primary estimates toward a larger-than-normal effect size.

Primary models, alternative models, and a comparison with Devine and Kopko’s model for “divergent cases.”
Conclusion and Implications
In this article, we have attempted to solve a problem common to the study of strategic choices in electoral campaigns: estimating the counterfactual outcome in a world where the strategic choice was not made. Our assessment of the vice-presidential HSA shows the limitations of existing approaches in this regard. Our solution is to apply the synthetic control method for comparative case studies to an instance with multiple treated units. We argue that this approach allows us to estimate the vice-presidential HSA without biases that plagued earlier studies.
In academic literature, as well as in assessments of modern presidential campaigns by politicians and members of the media, it has become conventional wisdom that vice-presidential candidates are, electorally speaking, nearly irrelevant—generally unable even to affect the vote in their own home states. We argue that this viewpoint is partly based on an HSA measurement with considerable inherent flaws. In contrast, our results suggest that the average vice-presidential candidate improves his or her tickets’ performance in his or her home state by 2.67 points of the two-party vote.
The effect size we estimate is close to that attributed to presidential candidates, and can have substantial consequences for electoral outcomes within particular states. From our sample of 1,524 total state-year elections, a 2.67 percentage point improvement in vote share would have swung the outcome to the opposing party in 294 elections, or 19.3% of state-years. Although most presidential elections are not decided by a single state, a 2.67-point swing in a single state can still affect the outcome of the presidential election as a whole. For example, had a strategically optimal running mate been available and selected, and had she produced at least a 2.67-point effect, such a choice would have swung the 1960 and 1976 elections to the Republican side, and the 2000 and 2004 elections to the Democrats. Specifically, a 2.67-point bump in New York in either 1960 or 1976 would have resulted in a Republican electoral college win. An increase of 2.67 points in any of six states (Florida, Missouri, Nevada, New Hampshire, Ohio, or Tennessee) in 2000 or two states (Florida or Ohio) in 2004 would have resulted in Democratic victories. 40
Although there is broad consensus among political scientists that the “fundamentals” are mostly responsible for shaping the outcomes of presidential elections, strategic choices are still ubiquitous in election campaigns. These include the type of choices campaigns make on a day-to-day basis, from having candidates hold campaign appearances in specific locations to buying advertisements in specific media markets. Studying the effect these choices have on electoral outcomes presents a minefield of inferential problems. We believe our use of the synthetic control method to estimate the vice-presidential HSA can serve as a blueprint for investigating the impact of many other strategic choices in election campaigns. Strategic choices are unlikely to single handedly change the winner of most elections. However, strategic decisions aggregate over the life of a campaign. Indeed, even taken in isolation, we have shown that the choice of a vice-presidential running mate could have altered the outcome in several key elections over the last 60 years.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
