Abstract
The “rally 'round the flag” effect—a short-term boost in a political leader’s popularity during an interstate political dispute—was first proposed by Mueller (1970) more than half a century ago. However, there is no scholarly consensus on its empirical validity and the circumstances under which the effect becomes most prominent. In this paper, based on a natural experimental design, we analyze large-scale worldwide surveys of 34,118 responses and causally identify the effects of 46 militarized interstate disputes (MIDs) on the approval ratings of political leaders in 27 countries. We find that MIDs, on average, decrease public support for national leaders. However, the public backlash could be attenuated depending on theoretically relevant contexts. Our finding implies that political leaders cannot rely on MIDs for public support increases, as they are generally penalized for such decisions.
Keywords
Introduction
With growing concerns that interstate conflicts are re-surging across multiple regions of the world, there has been renewed scholarly interest in understanding the association between public opinion and security policy-making. Mueller’s (1970) theory on the “rally 'round the flag” effect, proposed more than half a century ago, is one such foundational theory in international relations, positing that the onset of a militarized interstate dispute (MID) improves the incumbent political leader’s popularity, at least for a short time. For decades, numerous scholars have examined its empirical validity. Some studies find the overall effects of interstate disputes on U.S. presidential approval ratings to be close to null, while others show the conditional effects of several relevant factors.
It is important to resolve this long-standing debate because this well-known theory continues to exert influence on broader interpretations, despite the lack of a strong empirical foundation. In particular, proponents of the existence of the rally effect consider a logical consequence of Mueller’s (1970) theory and propose an extended theory: the “diversionary theory of war” (James 1987; Levy 1989; Morgan and Bickers 1992). They argue that political leaders can mobilize public support by initiating wars, incentivizing them to engage in MIDs. Such arguments could define political-societal conversations in mainstream media, with potentially significant consequences for the conduct of foreign policy by influencing voters’ policy attitudes.
In this article, we revisit the literature and re-test hypotheses previous scholars have proposed and tested. Specifically, we examine the fundamental hypothesis that an MID will result in higher approval ratings for national leaders. We also investigate three conditional effects. First, we examine whether rally effects will be smaller in aggressor states that initiate international conflicts than in states on the receiving end. Next, we test whether the hostility levels of conflicts affect the magnitude of the rally effects. Finally, we test whether rally effects will be greater if a dispute occurs in the context of a broader (and continuing) international security crisis.
In the existing literature, the results of testing these hypotheses are mixed at best. We believe that this inconclusiveness is partly due to two methodological problems. To produce more credible estimates, we address them using the following approaches.
First, we use novel cross-national data that have never been used in the existing literature. Specifically, we merge large-scale multinational and individual-level survey data from the Gallup World Poll with incident-level MID data from the Correlates of War (CoW) project. 1 The cross-national data allow us to generalize our findings regarding the rally effect beyond specific countries, especially the United States, and certain regions, such as Europe. Most existing literature focuses on the United States, which has unique historical and political characteristics, as Most and Starr (1989) and James and Rioux (1998) point out. Furthermore, the applicability of the theory has been constrained by limited cross-national studies that tend to focus on Europe, and usually in the context of a different type of incident, i.e., terrorism (Chowanietz 2011; Turkoglu and Chadefaux 2022; Falcó-Gimeno, Muñoz, and Pannico 2022; Godefroidt 2022)
Second, we compare the differences in the approval ratings of political leaders worldwide between respondents interviewed just before the onset of a militarized dispute and respondents interviewed just after. Under a set of reasonable assumptions, we can interpret the differences in the approval ratings between the two groups of respondents as causal. Focusing on a short period is also suitable for examining MIDs’ short-term effects, which we would not be able to estimate using monthly or quarterly data. This improves upon most existing studies, which use noisy time series (i.e., monthly or quarterly) data. Results from such data may be inconclusive due to methodological problems, such as omitted variable bias. If various other events happen before and after each dispute, the constantly changing information environments may affect public opinion. Using coarsely measured time-series data also poses a fundamental problem because the theory assumes the rally effect as a short-term phenomenon.
In addition to these problems, the existing literature has an issue of “scope condition.” In this study, we consider all MIDs identified by the CoW dataset, while previous studies do not adequately consider militarized disputes short of wars. These studies may assume that these “lesser” conflicts, in terms of hostility and geopolitical consequences, do not lead to the rally effect and therefore are not subjects worth investigating. However, we should expand the scope of inquiry to these other disputes, which are not wars, because many MIDs should satisfy the conditions for a “rally event” that Mueller (1970) originally proposed. These conditions are: (1) the incident is international (i.e., involves two or more sovereign states), (2) it directly involves the political leadership, and (3) it is sufficiently “specific, dramatic, and sharply focused” to draw public attention. Many militarized disputes, even those that do not lead to wars, could conceivably satisfy these conditions, although there could be variations in the third condition. Focusing only on wars excessively limits scope conditions. 2 Using a large number of MIDs introduces a substantive variation in conflict intensity, leading to a more nuanced understanding of conditional effects.
Specifically, in our data, 46 survey sampling periods coincidentally include the onset of an MID. With these cases, which include 34,118 responses from 27 countries between 2008 and 2014, we find the following. First, on average, militarized disputes decrease the public’s approval, and increase the public’s disapproval, of political leaders. We also find theoretically relevant effect heterogeneity. Importantly, we find that political leaders of states that defend against foreign aggression will not see any change in public support, while those who initiate offensive conflicts will suffer a public backlash. We also find that the rally effects are partially conditional on hostility levels. More dramatic actions can attenuate the backlash in public opinion. Similarly, incidents that occur in the context of a heightened international crisis also attenuate the public backlash.
These results have significant theoretical and practical implications for foreign policy. Most importantly, our findings may imply that political leaders cannot rely on militarized disputes for public support increases, as they are generally penalized for such decisions. However, depending on the context, they may still take advantage of interstate disputes in their favor by silencing public criticism.
Rally 'Round the Flag Effects
This section reviews the existing literature on the rally 'round the flag effect. In addition to the central hypothesis on the overall effect, some intriguing hypotheses regarding conditional effects exist.
Overall Effects
The seminal paper by John Mueller (1970) is often considered the first systematic investigation of the rally 'round the flag effect on public opinion. Mueller expands on the theoretical work of previous scholars (e.g., Kenneth Waltz, Tom Wicker, Richard Neustadt) and finds that the “rally 'round the flag variable” produces a “sturdy” short-term boost in a U.S. President’s popularity after an international crisis (Mueller 1970, pp. 34).
Since Mueller, scholars have canonized the expectation that American presidents would see their popularity increase after such a dispute (e.g., Mueller 1973; Sigelman and Conover 1981; MacKuen 1983; Ostrom and Simon 1985; Hurwitz and Peffley 1987; Russett 1990). Lee (1977), Kernell (1978), Ostrom and Simon (1985), and Ostrom and Job (1986) corroborate Mueller’s prediction that presidential popularity increases in the range of 5–7 percentage points after international crises. Marra, Ostrom, and Simon (1990) and James and Rioux (1998) suggest that U.S. presidents enjoy higher approval ratings after “vigorous” responses to international conflicts. Marra, Ostrom, and Simon (1990), Jentleson (1992), Nincic (1997), and Jentleson and Britton (1998) find that the use of force in geopolitically significant regions for the U.S. generates particularly larger rallies compared to those in other regions.
However, more recent empirical works dispute that the rally effects associated with militarized disputes are negligible. Edwards (1990) finds that most uses of force do not result in consistent rally effects: only 24 of the 85 incidents are correlated with higher approval ratings, and 11 incidents are associated with lower approval ratings. 3 James and Rioux (1998) show that when the Soviets are not involved, the rally effect almost disappears. Baker and Oneal (2001) expound upon previous research and conclude that the increase in presidential popularity is just 0.1% on average after militarized disputes. The estimated effect should be within the margin of error because Gallup does not measure public opinion to such a high degree of accuracy (Baker and Oneal 2001, pp. 670). 4
Although the empirical findings are mixed at best, the expectation that an electorate could rally 'round the flag has emerged as a foundational theory with significant influence on extended scholarly interpretations. For example, the diversionary theory of war (James 1987; Levy 1989; Morgan and Bickers 1992) assumes that unpopular political leaders have incentives to initiate wars to mobilize public support. The diversionary theory has also gained widespread traction in the news media. For example, Burakovsky’s (2022) op-ed article claims that in February 2022, tensions in Ukraine coincided with an increase in public support for Russian President Vladimir Putin. Similarly, in Britain, a CNN analyst claimed that in the face of a debilitating pandemic-related “Partygate” scandal, Prime Minister Boris Johnson intentionally tried to divert public attention to the deteriorating security situation in Eastern Europe. Johnson’s decision to visit Ukraine was allegedly motivated by the expectation that a foreign policy crisis would strengthen support for his premiership at home (McGee 2022). The prevalence of such articles in the media may have potentially significant consequences for the conduct of foreign policy today by influencing voters’ attitudes.
Conditional Effects
What might explain such variations in the empirical findings? Public approval of political leaders can fluctuate in response to many factors, and not all potential “rally events” show the same characteristics. In the following, we introduce three hypotheses on conditional effects advanced by the existing literature. 5
First, whether the policy objective is defensive or offensive may explain variations in public support in the aftermath of a militarized dispute. The existing literature suggests the following hypothesis:
In the American context, Jentleson (1992) and Jentleson and Britton (1998) demonstrate that the American public is much more supportive of the use of force when the primary objective is to restrain aggression by other governments. In contrast, the public is less inclined to support the use of military force to achieve revisionist goals, such as creating internal political change in other governments. Similarly, Nincic (1997) notes that “protective” interventions for American geopolitical interests are more popular among the electorate than “promotive” interventions that seek to secure a net foreign policy gain proactively. On the other hand, Baker and Oneal (2001) claim that the rally effects will be greater (but still modest, 1.20% on average) for disputes in a small subset of cases, in which the United States was “both an originator and a revisionist” (p. 674).
Second, some studies examine the relationship between the intensity of a dispute and the magnitude of changes in public support for the executive. Specifically, we examine the following hypothesis.
Mueller (1970) theorizes that rallies are most common after international events that are “specific, dramatic, and sharply focused” (p. 21). Ostensibly, this criterion may imply that public support increases only after a rally event that is consequential enough (i.e., is characterized by highly intense uses of force) to increase presidential popularity (Mueller 1970, p. 20). However, as Lian and Oneal (1993) point out, only one-third of Mueller’s sample of 34 events involves the actual use of military force. Mueller perhaps assumed that the rally effect, albeit small, could occur in disputes that do not necessarily lead to using military force or even to war. Hence, we reiterate the importance of including all MIDs to re-examine Mueller’s conclusions concerning the relationship, if any, between the levels of force utilized in a dispute and the magnitude of changes in public opinion.
When it comes to the conditional effect of the rally, James and Rioux (1998) suggest a pattern that is opposite to Mueller’s (1970) expectation. Specifically, they find a significantly negative relationship between the level of force used in a militarized conflict and rallies with presidential approval. With crisis escalation, the rally effect dissipates, indicating that the American public may be quite cautious about using force. On the other hand, Baker and Oneal (2001) consistently observe large rallies after a U.S. decision to enter into a war, the most intense level of conflict defined in the Correlates of War dataset. 6 The average rally associated with an American entry into the war is almost 8% points. 7 Compared to wars, less intense uses of force, such as threats or displays of force, or use of force short of war, are correlated with minimal rallies in public opinion.
Regarding the international context, there is also disagreement on whether conflicts that occur during more hostile international environments create larger rally effects. Previous scholars have debated the validity of the following hypothesis.
Brody (1991) argue that a rally is more pronounced during a period of increased international tension because a collective sense of danger pervades the public. Lian and Oneal (1993) confirm that the public is more likely to support the use of force when the U.S. is already involved in a severe international crisis. On the other hand, Oneal and Bryan (1995) show no significant relationship between the severity of the international context and the magnitude of the rally effects. Furthermore, other studies show that prolonged wars tend to be unpopular with the public (Mueller 1970, 1973; Ostrom and Job 1986; Cotton 1986) and bring additional electoral misfortune to leaders in democracies (de Mesquita, Siverson, and Woller 1992; de Mesquita and Siverson 1995).
Baker and Oneal’s (2001) analysis may present a more nuanced picture. Although crises in more hostile international environments produce greater rally effects, the effect may be more strongly influenced by other covariates. For example, a severe international situation should strongly correlate with increased publicity. Their argument about the existence of other confounders is reasonable and could apply to most existing studies that test both the general and conditional rally effects.
Research Design
To test the hypotheses specified in the previous section, we use a combination of data on militarized interstate disputes (MIDs) and worldwide surveys. In this section, we first discuss several methodological problems inherent in previous studies. Next, we explain our data and research design for causal inference. Finally, we elaborate on our statistical methods.
Problems in Existing Studies
We argue that existing empirical tests of the rally effect suffer from two significant flaws: the lack of cross-national studies and an identification strategy for causal inference. We discuss these problems in turn.
Lack of cross-national studies
Undoubtedly, the U.S. is an essential case for testing the theory. For decades, most studies of the rally effects of MIDs have used polling data on presidential popularity in the United States. 8 James and Rioux (1998) argue that given the political and military preeminence of the U.S. in the international system, the American context readily provides valuable scholarly research. Furthermore, they caution that lower levels of autonomy in other nations’ foreign policies (that is, the executive has less autonomy in choosing when to use military force), data limitations, and other “idiosyncratic factors” hamper cross-national studies. Similarly, Most and Starr (1989), and James and Rioux (1998) suggest that non-U.S. leaders’ foreign policy decisions could be more a matter of opportunity than political willingness.
Is the U.S. an anomaly in that the theory is applicable only in the U.S. context, or is it one of many cases for which the theory has predictive power? To examine this question, some scholars suggest the importance of conducting similar studies in other national and cross-national contexts. In the seminal study, Mueller (1970) already suggested that data on national leaders of “such countries as Britain, Canada, and France” (p. 33) could be used to complete similar studies on the effect of conflicts on the popularity of political leaders. More than six decades have passed, but scholars still call for more cross-national studies. For example, in their study of the effect of military casualties on the popularity of incumbents in 10 OECD nations, Kuijpers (2019) notes that rally effects, for which researchers may extrapolate their arguments from U.S. findings, are not necessarily valid for other advanced industrial Western democracies.
We believe that the growing body of literature on rally effects in other nations reflects the importance of examining the theory cross-nationally. For instance, several studies examine European responses to terrorist attacks. Chowanietz (2011) examines European responses to terrorist attacks and argues that “rallies around the flag are the rule” (p. 673) in nations such as the United States, the United Kingdom, France, Germany, and Spain. However, Turkoglu and Chadefaux (2022) use survey data from 30 European democracies collected over 15 years to show that there is no effect of terrorism on public approval. Falcó-Gimeno, Muñoz, and Pannico (2022) examine the effects of terrorism on the electoral support of the incumbent in Spanish elections, suggesting that the effects depend on temporal proximity. Finally, Godefroidt’s (2022) meta-analysis confirms the rally effect but warns us about overgeneralization. She points out that “studies on non-Muslim violence or conducted in non-Western contexts remain virtually nonexistent” (p. 2).
While the latest series of scholarly literature avidly adopts cross-national designs, they primarily examine the rally effect associated with terrorism specifically, and not militarized interstate disputes. Consequently, the findings from these terrorism studies may not be directly relevant to testing various hypotheses derived from the classic theory of rally 'round the flag effect, which is about the consequences of interstate conflicts, not terrorism.
The lack of causal identification strategy
The second problem in the existing literature is that these studies are insufficient to make valid causal inferences. This issue is related to the common usage of coarse-grained time series data. These studies use varying units of time series analysis, ranging from quarterly data (e.g., Ostrom and Job 1986; Morgan and Bickers 1992; James and Hristoulas 1994) to monthly data (e.g., Lian and Oneal 1993; James and Rioux 1998) or data with 6-week intervals (Baker and Oneal 2001). However, using such a long interval would prevent researchers from eliminating other confounders.
Importantly, Russett (1990) argues that rallies are short-term phenomena. Consequently, allowing too much time between a government’s major response and the succeeding poll could fail to capture a rally effect. Oneal and Bryan (1995) acknowledge that economic changes between the two (monthly) surveys may influence the public’s evaluation of presidential performance. Similarly, James and Rioux (1998) recognize that a shorter time interval in observations is always desirable to minimize any “confounding factors” that can occur. They hypothesize that “more complete data and a shorter time frame should reveal that involvement in crises, in particular, does not produce significant effects” (James and Rioux 1998, pp. 786).
Natural Experiments using Worldwide Surveys
We address these two problems using cross-national and individual-level data. Specifically, we use cross-national survey data in addition to MID data and adopt the so-called unexpected-event-during-survey or UEDS (Muñoz, Falcó-Gimeno, and Hernández 2020). 9 Under a set of reasonable assumptions discussed in the next subsection, we can make a causal interpretation of the onset of an MID by comparing survey responses measured just before the incident and those measured just after.
Some recent studies use similar date-specific observational designs to examine changes in public opinion following terrorist attacks (Holman, Merolla, and Zechmeister 2022; Turkoglu and Chadefaux 2022). For example, Holman, Merolla, and Zechmeister (2022) use a similar design to show the lackadaisical rally of the British public behind Prime Minister Theresa May in the aftermath of the Manchester and London Bridge terrorist attacks. 10
Importantly, our study is the first to estimate the impact of militarized interstate disputes (not terrorism) based on natural experiments. We justify our focus on MIDs for the following reasons. First and foremost, interstate disputes are the primary and original subjects of inquiry in the literature on rally 'round the flag effects. Second, only by using MIDs can we test the effects that are conditional on “sides” (Hypothesis 2). In contrast, studies on terrorism allow us to examine the effects of receiving attacks, but not the effects of initiating attacks. Third, the effects of interstate conflicts may be substantively different from those of terrorist attacks. Importantly, as we introduced above, Godefroidt (2022) shows that researchers often examine the effects of Islamist terrorism on public attitudes in Western countries. Therefore, it is difficult to generalize the rally effect beyond the specific political and ethnic contexts of these terrorist attacks. Finally, compared to terrorist attacks, MIDs have a wide variation in hostility levels. This variation allows us to test Hypothesis 3.
To examine the public responses to the onset of each MID, we use two data sets. The first is incident-level data of the Militarized Interstate Disputes from the Correlates of War (COW) project. 11 The second is the Gallup World Poll, which contains comprehensive individual-level survey data covering more than 160 countries. 12 To the best of our knowledge, no scholar has yet examined the effects of MIDs on public opinion using such a global scale of individual-level data.
We compiled our data from these two large-scale databases as follows. First, we identified all country-specific and year-specific surveys in which the onset of a MID in a particular country coincidentally occurred during the survey sampling period in that country. 13 We then eliminated the cases in which the main outcome question (to be explained) was not measured.
Second, we selected a subset of these cases by setting a narrower “bandwidth” for analysis; specifically, 7 days before and after the beginning of an incident (excluding the date of onset). Bandwidth length is generally a trade-off: a narrower bandwidth tends to yield higher interval validity (because it minimizes the influence of other post-incident events and situations), but a wider bandwidth would consider more cases and observations (increasing external validity and statistical power, respectively). In our study, bandwidth selection is largely data-driven: in 90.0% of all cases after the initial screening mentioned above, the duration of an incident (excluding the date of onset) is within 7 days. Since this very short bandwidth covers almost all cases, we have no reason to expand it. Also, we chose not to further reduce the bandwidth to fewer than 7 days. We assume that a period of 7 days accounts for the maturation of the intensive news cycle that generally follows the onset of each dispute.
Third, if the onset dates of two (or more) disputes are less than 7 days apart, we excluded both from consideration. 14 This procedure aims to estimate the independent effect of each dispute on public opinion. To have enough statistical power, we excluded cases with fewer than 30 observations each before and after the incident’s onset. Because we cannot identify the treatment status of the respondents interviewed on the day of the beginning of an incident, we also excluded them from our analysis, following similar practices used by other studies that rely on date-specific observations (e.g., Goldsmith, Horiuchi, and Matush 2021).
After these screening stages, the total number of valid cases for our analysis is 46. The number of respondents is 34,118 in 27 countries. The data set covers 42 surveys administered between 2008 and 2014. Details of all cases are listed in Section A of the Supplementary Materials.
Statistical Methods
We use the following survey question for our main analysis: “Do you approve or disapprove of the job performance of the leadership of this country?” It is the most commonly used question to measure the approval rating of a political leader. For this reason, it is particularly suitable for testing the rally 'round the flag effect (see James and Rioux 1998; Newport and Saad 2021).
For our statistical analysis, we use a dichotomous treatment variable: 0 if a respondent was interviewed before the date of each MID’s onset and 1 if a respondent was interviewed after the onset. We assume that respondents in the treatment group (i.e., those who were interviewed after the MID onset) were “treated,” meaning they were aware of the onset itself. However, because this study is not based on researcher-administered true experiments, we do not have an adequate measure of each respondent’s “compliance status,” which might be obtained through a question asking for the respondent’s familiarity with the relevant MID. The possible existence of respondents who did not know the corresponding incident should attenuate the treatment effects toward zero. Therefore, we consider our estimates conservative. To put it differently, if most (if not all) respondents in the treatment group were fully aware of the incidents, we would expect even larger treatment effects (in absolute value).
The outcome variables are also dichotomous, corresponding to the three response choices: “Approve,” “Disapprove,” and either “Don’t know (DK)” or “Refused.” Specifically, Approve is 1 if a respondent approved of the leadership’s job performance and 0 otherwise. Disapprove is 1 if a respondent disapproved of the leadership’s job performance and 0 otherwise. Finally, DK/Refused is 1 if a respondent neither approved nor disapproved of the leadership’s job performance and 0 otherwise. We collapse the last two response categories (“DK” and “Refused”). We use this combined category for our analysis because it may indicate the unclear disposition of the respondents to express their attitudes due to insufficient information or ambivalent views. A political leader’s decision to get involved in a militarized dispute may send confusing or conflicting signals to citizens, so using this category as another outcome variable is substantively relevant.
We follow the statistical approach suggested by Goldsmith, Horiuchi, and Matush (2021) in their analysis of Gallup World Poll data. They use a linear probability model (LPM) that includes the dichotomous treatment variable, one of the three dichotomous outcome variables mentioned above, and case-specific fixed effects. This approach is justifiable for the following reasons. First, the “running variable” that deterministically shapes each respondent’s treatment status is discrete and takes only 14 different values (7 days before or after the onset date). Therefore, applying the standard regression discontinuity design, which requires the assumption of “continuity” (of continuous variables) (de la Cuesta and Imai 2016), is not suitable. Second, it is critically important to add fixed effects to take advantage of variations within each case. For this reason, running an ordered probit or logit model is not suitable due to the incidental parameter problem (Neyman and Scott 1948)
Two assumptions are required to make a causal inference based on our design (Muñoz, Falcó-Gimeno, and Hernández 2020). The first is temporal ignorability: the interview timing (that is, whether respondents are interviewed before or after the onset of MIDs) is independent of the respondents’ other characteristics that may influence how they respond to the outcome question. The temporal ignorability assumption is likely satisfied because Gallup uses random-digit dialing. Therefore, it is doubtful that any particular characteristic of the respondents is systematically correlated with the interview date. 15
The second assumption is excludability: whether a respondent is interviewed before or after a particular date affects the outcome variable only by the onset of an incident and the media coverage associated with it. Other potentially confounding events unrelated to the MID itself could occur on the beginning day of each MID, and they might influence the public’s approval of their political leader. We would find it hard to assume the validity of the excludability assumption if we focused only on a specific event within a particular country. Importantly, however, we use large-scale respondent-level data covering 46 cases to estimate the average rally effect in these cases. Therefore, we assume that certain types of unrelated events independent of the onset of an MID itself are unlikely to occur systematically in conjunction with interstate conflicts and affect outcome variables.
To test conditional effects (Hypotheses 2, 3, and 4), we construct various subsets of our data based on each of the theoretically relevant moderators. We then estimate the treatment effects for each subset and graphically present the results. 16
Results
In this section, we first show the results of testing the overall effect (Hypothesis 1) and a range of robustness tests. We then introduce the results of testing various conditional effects. 17
Hypothesis 1: Overall Effects
Figure 1 shows the results of testing Hypothesis 1 (the overall effects) with all 34,118 responses for 46 cases.
18
The dots indicate the average treatment effects, and the vertical lines represent their 95% confidence intervals. Effects that are statistically significant at the 0.05 level are highlighted in red. Average treatment effects using all the MIDs. Note: The treatment effects significant at the 0.05 level are highlighted in red.
The directions of the estimated effects are opposite of our hypothesis: The onset of an incident decreases the public approval of a political leader’s job performance by 1.44% points, while it increases the disapproval rating by 1.94% points, respectively. The effect on other responses is almost null (−0.49% points). Therefore, we reject Hypothesis 1.
The magnitude of the effect, in the range of 1–2% points, may be considered modest. For example, our findings do not differ greatly from Baker and Oneal’s (2001) finding of a treatment effect of 1.20% for the U.S. cases (albeit in the opposite direction). However, we could detect the statistical significance of such a modest effect because our sample size is much larger than any previous study.
Is the United States an exception?
As discussed in Section 2, most previous studies, including Mueller (1970), focus on the United States. Some scholars claim that rally effects are primarily a U.S. phenomenon (Kuijpers 2019). We can test this claim using our data. Specifically, we run a mixed-effect model with the treatment variable interacting with country-specific dichotomous variables as “fixed effects” and case-specific “random effects” (or random intercepts).
The results are presented in Figure D.1 in Section D of the Supplementary Materials. The treatment effect on Approval tends to be slightly larger, and the effect on Disapproval tends to be slightly smaller for the United States compared to most other countries. However, there is no clear difference across all the countries included in the analysis. There is no evidence that the effect is specific to the United States.
Robustness tests
We perform three robustness tests for this main analysis. First, as discussed, an important assumption for causal inference is temporal ignorability. With the random digit dialing used by Gallup for their sampling scheme, it is unlikely that the types of respondents are systematically correlated with the dates of the beginning of the MIDs. We test this assumption using a range of observable measures. The results are presented in Section B of the Supplementary Materials. We find some variables unexpectedly correlated with the treatment variable. However, as expected, adding these statistically significant variables does not change the results of our findings (Figure D.2 in Section D of the Supplementary Materials).
Second, we also believe that excludability is satisfied as long as we combine a large number of MIDs. It is difficult to consider other events that would happen systematically alongside MIDs. However, there is some concern that the MIDs are not completely exogenous and “unexpected” because foreign policy decisions are often preceded by public signaling of political intent. 19 Importantly, each dispute could occur as part of a broader succession of disputes between states. Therefore, there may be some trends among these disputes. It might be suggested that these patterns can be controlled to some extent by adding case-specific trend variables or general trend variables. However, because our “running variable” is not a continuous variable, with only 14 data points, it is difficult to capture trends that are robust to statistical specifications. 20
Therefore, instead of trying to measure and control such patterns before and after the onset of MIDs, we use the shortest possible bandwidth and re-estimate the treatment effects by leveraging the large number of observations we have for analysis. Namely, we only compare respondents who were interviewed just a day before the beginning of each MID and those who were interviewed just a day after. Although the numbers of cases and respondents decrease, this analysis is more suitable for satisfying the assumption that there is no systematic difference between the treatment and control groups. 21 The results are presented in Figure D.3 in Section D of the Supplementary Materials. Although some effects become statistically insignificant, the effects are substantively similar to those presented in Figure 1.
Finally, to examine the sensitivity arising from the possible inclusion of a highly influential case among the 46 cases, we perform “leave-one-out” tests. Specifically, we exclude each of these 46 cases sequentially and re-estimate the treatment effects. The results are presented in Figure D.4 in Section D of the Supplementary Materials. Although excluding the case of India in April 2011 would make the effect on the approval rating barely insignificant, all other effects on the approval and disapproval ratings are statistically significant. More importantly, excluding each of the three U.S. cases does not change the estimates substantially. Therefore, we can conclude that the results are fairly robust and do not derive from the specific political context of the United States.
Hypothesis 2: Effects Conditional on Sides
Figure 2 shows the treatment effects conditional on the countries’ “sides,” which indicate whether a country is on the defensive side (i.e., engaged in an interstate conflict initiated by a foreign country) or on the offensive side (i.e., engaged in an interstate conflict initiated by the respondents’ own country). If countries are on the defensive side of a conflict, neither the approval rating nor the disapproval rating changes substantially. In contrast, if countries are on the offensive side, public approval decreases by 2.70% points, while disapproval increases by 2.77% points. These effects are highly significant. The null effect on the “DK/Refused” category of responses (i.e., a decrease by 0.07% points) indicates that when the respondents’ country initiates an interstate dispute, such a foreign policy decision clearly shifts people’s attitudes, instead of making respondents more (or less) ambivalent in the evaluation of their political leaders. Average treatment effects conditional on sides. Note: The treatment effects significant at the 0.05 level are highlighted in red.
Again, we test the robustness of our findings by sequentially excluding each case. 22 The results are presented in Figures D.5 (for the defensive cases) and D.6 (for the offensive cases) in Section D of the Supplementary Materials. We observe highly robust results for offensive cases. All the coefficients for the approval and disapproval ratings are statistically significant. When it comes to defensive-side cases, there are some cases whose exclusion slightly changes the results, such as India in December 2012 and September 2013. But the directions of these sensitive cases are opposite to Hypothesis 1. That is, even when some cases are excluded, the approval rating decreases, and the disapproval rating increases. Put differently, these cases in India could be anomalies in which the approval rating increases and the disapproval rating decreases. 23 The percentage of other ambivalent responses decreases.
These results, in part, support Hypothesis 2, particularly the arguments previously made by Jentleson (1992), Jentleson and Britton (1998), and Nincic (1997). These previous studies predict that using military force for revisionist objectives would create weaker or null rally effects. Our findings corroborate that the public is strongly against the use of force for such purposes. This new finding has important policy implications. Political leaders are unlikely to succeed in diverting attention away from domestic problems by initiating military action against a foreign country. In other words, the diversionary theory of war (James 1987; Levy 1989; Morgan and Bickers 1992) has no empirical basis and therefore provides minimal justification for associated foreign policy decisions.
Our findings also have important theoretical implications. In our study, the defensive cases include those of Azerbaijan, Cambodia, Greece, India, Iraq, Lebanon, Pakistan, the Philippines, Ukraine, and Venezuela. It is worth noting that none of these cases involves the United States. Therefore, the vast majority of existing studies that use data from the U.S. could have limited predictive power across non-U.S. countries.
Hypothesis 3: Effects Conditional on Hostility Levels
What if a political leader chooses a less dramatic course of action? Hypothesis 3 refers to the treatment effects conditional on hostility levels. In the CoW dataset, each incident is given one of the following five hostility levels: (1) no militarized action, (2) threat to use force, (3) display of force, (4) use of force, and (5) war. Among the 46 valid cases for our tests, there is no case of war, which we discuss in the conclusion as an analytical limitation. However, given that two-thirds of the Mueller’s (1970) cases are incidents in which no militarized force was utilized at all, our sample supports a more general claim about the rally effect. We further note that other studies that consider wars, such as Baker and Oneal (2001), also suffer from extremely small sample sizes of conflicts of the highest intensity.
Importantly, there is no reason to believe that MIDs of lesser intensity cannot satisfy Muller’s concept of a “specific, dramatic, and sharply focused” event. As discussed earlier, there could be variation in the degree to which disputes satisfy Mueller’s criteria of a so-called “rally event.” Because it is a matter of degree, rather than a dichotomous classification, it is worth examining the rally effect conditional on this variation.
Specifically, we collapse the first two categories because there is only one incidence of “threat to use force.” Therefore, we estimate the treatment effects for each of the three groups of incidents: no militarized action (including a threat to use force), display of force, and use of force. The number of cases is 12, 9, and 25, respectively.
The results are presented in Figure 3. The right panel includes cases where countries use military force. The left and center panels show the treatment effects of countries that took no militarized action and those that merely displayed (and did not use) military force. These panels suggest some variation in the magnitude of the effects. Citizens’ critical views of their leader increase when there is no militarized action: the percentage of approval decreases by 3.59% points, and the percentage of disapproval increases by 3.66% points. Both effects are highly significant, but their effects are attenuated when a country displays its military force. The estimated coefficient for approval (−3.39% points) is still significant, but the estimated coefficient for disapproval (1.98% points) becomes insignificant. The insignificance is partly due to the small number of cases (i.e., 9). Most notably, the point estimates for the cases in which countries used actual force are very small, with 0.15% points on approval and 1.05% points on disapproval. They are statistically insignificant despite constituting the largest number of cases (i.e., 25) among the three categories. Average treatment effects conditional on hostility levels. Note: The treatment effects significant at the 0.05 level are highlighted in red.
These results may partially support Hypothesis 3. As the level of hostility increases, people become more supportive of the leader’s foreign policy decision, compared to the cases with less aggressive foreign policy decisions, for which we observe a backlash effect. However, more aggressive actions do not necessarily cause significant increases in support. Hence, our findings are inconsistent with Hypothesis 1. This finding is another reason to question the validity of the diversionary theory of war. Empirically, political leaders cannot increase their domestic support by using military force.
However, these results are not necessarily robust in the selection of cases. Figures D.9, D.8, and D.7 in Section D in the Supplementary Materials show the results of the leave-one-out tests. For cases with no militarized action, excluding the case of India in April 2011, would attenuate the effects toward zero (also see Figure D.4). In contrast, excluding the case of India in December 2012 would increase the magnitude of the effects (also see Figure D.5). We observe similar variations for cases in which countries display force. For cases in which countries use military force, the case of India in September 2013 seems to be highly influential. When this case is excluded, public reactions to their political leaders become more critical. In other words, India appears to have had a rally effect in September 2013 that boosted the popularity of the leadership. This case is classified as a defensive use of military force, and Figure D.5 also suggests that it was influential. More research may be needed to investigate influential and anomalous cases. In particular, India could provide an important context for understanding how the public responds to national security crises.
Although there may exist a small number of potentially influential cases, it is worth noting that excluding these cases nonetheless produces results that are opposite to the general theory of rally 'round the flag effect (Hypothesis 1), and the associated diversionary theory of war. For example, when we excluded the case of India in September 2013, the disapproval rating would increase among the cases where countries used military force.
Hypothesis 4: Effects Conditional on the Context
Finally, Hypothesis 4 evaluates the conditional effects of the international context. Existing studies do not clearly agree on how the international environment affects the rally effect. This hypothesis is perhaps difficult to test empirically, particularly cross-nationally, because leaders and citizens of countries could perceive heightened international tensions differently.
One way to test this hypothesis is to identify whether each MID is in the context of an “international crisis.” There are many possible definitions of international crises. For our analysis, we use the International Crisis Behavior (ICB) Project’s identification of major militarized crises. 24 To hold country-specific environments and conditions constant as much as possible, among our 46 valid cases, we find a pair of MID cases that satisfy the following conditions. (1) The cases concern the same country. (2) The cases have the same hostility level and “side.” (3) However, one MID case occurred during an ICB-identified crisis, while the other did not. We find that two MIDs involving South Korea satisfy all these conditions. In both cases, South Korea displayed military force and was on the belligerent side.
The results presented in Figure 4 suggest that the negative reactions of Korean citizens tend to be stronger when a foreign policy action is taken not within the context of an international crisis. This finding complements our observations in Figure 3 for Hypothesis 3. The findings suggest that when an incident escalates to a crisis level and/or a country takes more aggressive action, citizens’ criticism of the government’s decision may be attenuated. Still, as we discussed earlier, we find only the effects of attenuated criticism. There is no evidence of an increase in outright support for the government. Average treatment effects using two cases in South Korea, one in crisis and another not in crisis). Note: The treatment effects significant at the 0.05 level are highlighted in red.
An alternative way to test Hypothesis 4 is to compare the cases by leveraging variation in the number of days since the end of the previous MID (if any) within the same country. We admit that this method is not a sufficiently valid measure for increased international tension. But turnaround periods between MIDs could approximate incident frequency, which in turn could be an approximate proxy for the severity of an international crisis in which the corresponding country is involved. 25
The results are presented in Figure 5. These cases are divided into 3 groups, (1) a turnaround period of 2 weeks or less, (2) a turnaround period of more than 2 weeks but less than a year, and (3) a turnaround period exceeding a year (or no MID previously). We find that a more severe international context (i.e., with more frequent disputes) seems to attenuate people’s critical views of the government. Specifically, if the previous incident occurred more than 2 weeks ago, the effect on the approval rating is negative: −1.93% points if the turnaround period is within a year and −3.10% points if it is longer than a year. However, if the turnaround period is within 2 weeks, the effect on the approval rating is positive at 2.32% points, which is close to significance at the 0.05 level. Average treatment effects conditional on the number of days since the end of the previous MID (if any) within the same country. Note: The treatment effects significant at the 0.05 level are highlighted in red.
Conclusion
Our results paint a nuanced picture of public opinion after the onset of militarized interstate disputes (MIDs). Most importantly, we find that MIDs do not cause rallies in the public’s attitudes toward political leaders in the short term. Rather, the government’s decisions to engage in MIDs increase citizens’ criticism of the job performance of their political leaders. The increase in the disapproval rating and the decrease in the approval rating could be attenuated, depending on the context. Specifically, we find attenuated effects if a country engages in a conflict to defend itself against the belligerency of another country, or if the intensity of a conflict is relatively high (i.e., the use of actual force), and/or if an action is taken in the broader context of heightened international crisis. Nevertheless, except for some anomalous cases, such as those in India, we do not find any systematic evidence that people’s support for the government increases outright, as should be expected from the original theory on rally effects.
These findings have important scholarly contributions. First and foremost, the theory of the “rally 'round the flag” effect does not have a robust empirical foundation. Previous studies (e.g., Lee 1977; Kernell 1978; Ostrom and Simon 1985; Russett 1990) may have been susceptible to model specifications, or their findings may apply only to idiosyncratic situations in the United States. Second, our analysis presents the largest cross-national evidence that people tend to be averse to interstate disputes, as Jentleson (1992), Nincic (1997), and Jentleson and Britton (1998) argue. Belligerent attitudes in citizens, if any, could be anomalous rather than typical. Third, these sentiments against militarized conflicts could be attenuated depending on the exogenous circumstances. Citizens may be more reluctant to express their skepticism of the government when their countries are attacked by foreign nations, when the level of hostility is high, and when the case occurs during a heightened international crisis. Our findings on these conditional effects are preliminary. We call for further investigation into how various interstate conflicts skew public discourse in favor of incumbent political leaders and influence public opinion. In particular, examining whether heightened national security crises influence media coverage and suppress opposing views should be an important topic for future research.
Our findings also have policy implications. First, as we have repeatedly noted in this paper, we do not find empirical evidence implying that political leaders can increase popularity by engaging in interstate conflicts. Second, however, depending on the context of these conflicts, political leaders’ decisions could silence opposing views among citizens. This could be a matter of concern for democratic countries: political leaders may use international conflicts to their advantage. Leaders in authoritarian regimes could also use these conflicting situations to silence citizens’ discontent against their regimes. Therefore, our analysis suggests that, while the diversionary theory of interstate conflicts may be partially valid, it likely operates on mechanisms that are substantively different from what previous scholars have assumed.
In addition to examining the potential role of interstate conflicts in suppressing opposing views, there are several other avenues for future research. First, if data become available, we should examine the effects of wars (i.e., cases with hostility levels coded at 5). Our sample does not include such cases. The number of wars is small and did not coincide with any of Gallup’s survey periods. The data limitation is also partly due to the current version of the CoW dataset, which does not include incidents after 2014. With time, we hope to access more observations on the CoW dataset, making it more likely that multiple observations coincide with Gallup World Poll survey periods.
Analyzing the cases of war, once more data become available in the future, may provide further opportunities for theoretical refinement. Suppose we find the rally effects associated with war. In that case, we need to modify our arguments to indicate that aside from the most intense of conflicts (i.e., war), militarized disputes fail to create a rally effect. Alternatively, if we do not find the rally effects associated with war, we would be even more confident in rejecting Hypothesis 1, concluding that all militarized disputes fail to create a rally effect.
Second, to our knowledge, all previously published studies on the rally effect examine the effect on public attitudes toward political leaders (or governments) of countries that engage in conflicts. However, there may be spillover effects of conflicts in other countries that are not involved in the conflicts. An important new study by Fukumoto and Tabuchi (2022) finds that the Russian invasion of Ukraine increased the popularity of the ruling parties of Japan by approximately 4% points. Wars in other countries can be opportunities for leaders of “third-party” nations to gain popularity without becoming belligerent in any conflict.
Ultimately, advancing our understanding of the association between public opinion and international conflict should be an urgent priority for political scientists. In presenting our findings, we highlight the importance of critically examining a common belief that political leaders may take and justify decisions of war and peace for domestic political gain. Identifying how the general public responds to such political decisions, as well as the political communication associated with these decisions, has become increasingly relevant because modern societies are most vulnerable to disinformation campaigns conducted directly by politicians or through the mass media and social media networks. We should continue to examine the behavioral foundation of international relations within emerging contexts.
Supplemental Material
Supplemental Material - Natural Experiments of the Rally 'Round the Flag Effects Using Worldwide Surveys
Supplemental Material for Natural Experiments of the Rally 'Round the Flag Effects Using Worldwide Surveys by TaeJun Seo, and Yusaku Horiuchi in Journal of Conflict Resolution.
Footnotes
Acknowledgments
We thank Matias Engdal Chrisensen, Kentaro Fukumoto, Jean Hong, Kelly Matush, Megumi Naoi, Kosuke Imai, Kyosuke Kikuta, Daniel Simmons, Sam van Noort, and other conference participants for their useful comments.
Author’s Note
We presented earlier drafts at the 2022 Summer Meeting of the Japanese Society for Quantitative Political Science (Tokyo, July 8, 2022), and at the 2022 Annual Meeting of the American Political Science Association (Montreal, September 15–18, 2022). Seo initially wrote this manuscript as an independent research paper supervised by Horiuchi.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
