Abstract
In 2014, the most prominent anti-corruption investigation in Latin America called Lava Jato, exposed a Brazilian corruption scheme with reverberations in 61 countries, resulting in legal judgments for nearly 5 billion USD in reimbursements thus far. This article applies the synthetic control method on data from 135 countries (2002–2018) to test the hypothesis that Lava Jato impacts the Worldwide Governance Indicators in Brazil. The findings reveal that Lava Jato negatively affects control of corruption, the rule of law, and regulatory quality. There are signs of possible improvement in at least the corruption and the rule of law measures. This paper brings value to the criminological body of literature, notably lacking in the Global South.
The world faces renewed debates regarding the scale and cost of corruption to the global economy. In real-world settings, corruption is relatively high, and anti-corruption reforms are generally challenging to implement (Windsor, 2019). Cases involving senior executives involved in fraud show how difficult it is for regulators to expose injustice without self-disclosure or the corporation’s cooperation. For regulators, corporate fraud is difficult to prevent. It is even more difficult for law enforcement to address, as those who want to defraud usually use sophisticated methods (Araujo, 2020). Examples of the complexity in fighting corruption are Walmart’s corruption crimes in Mexico (Barstow, 2012; Foroohar, 2012), Siemens’ global corruption strategy in many countries (Berghoff, 2017; Klinkhammer, 2013), and the Odebrecht corruption scandal (Araujo, 2020). In the case of Wal-Mart, Barstow (2012) described an orchestrated corruption campaign to gain market dominance in Mexico and where the company had paid bribes in almost all parts of the country. For Siemens, the company made 77 payments totaling $18.6 million to enter into telecommunications contracts (Berghoff, 2017). The anti-corruption operation called Operação Lava Jato (Operation Car Wash), starting in 2014, exposed the Odebrecht 1 scandal. Investigators found that Odebrecht paid millions of dollars to authorities and politicians to enter into agreements with the Brazilian State oil company Petrobras (Araujo, 2020). Although the Lava Jato led to the conviction of defendants and raised the issue of corruption at a national level, critics argue that the investigative procedure or even the legal process was biased, used for political purposes, and disrespected the fundamental rights of the accused (Agência Brasil, 2016). Brazil is a country of continental dimensions, and its administration is divided into three levels of government: federal, state, and municipal. The country comprises 26 states and the Federal District, and at the municipal level by 5,570 municipalities. Furthermore, the country has tremendous economic and social disparities. In this context, the challenges and benefits obtained from possible improvements in public governance are enormous.
We apply Synthetic Control Method (SCM), a quasi-experimental design approach, to assess the causal impact of Lava Jato on governance indicators. Several empirical applications in other areas have implemented SCM (Abadie et al., 2015; Acemoglu et al., 2016; Bartos & Kubrin, 2018; Billmeier & Nannicini, 2013; Bueno & Valente, 2018; Cavallo et al., 2013; Pinotti, 2015). This article tested the hypothesis that Lava Jato impacted the Worldwide Governance Indicators (WGI) in Brazil. At the time of writing, this was the first paper to evaluate Lava Jato’s impact on WGI. Furthermore, it brings value to the criminological body of literature, notably lacking in attention to the Global South.
The Worldwide Governance Indicators and the Challenge of Measuring Corruption
Although policymakers and scholars have long discussed the concept of governance, a strong consensus has not been reached around a single definition of governance or institutional quality (Kaufmann et al., 2010). Over the past 3 decades, governance research has attracted politicians and academics from various economics, public administration, and political science disciplines. Although there is no consensus on the definition of corporate governance, the research needs some stylized facts (Polat, 2020). Although good governance is conceptually complex, several studies have sought to measure it (Bevir, 2009; Rhodes, 2007; Sbragia, 2000; UNDP, 2007; Williams & Siddique, 2008).
There is a debate in the public administration literature on practical/normative issues in this governance (Triantafillou, 2019). Aggregate indexes of the quality of governance and corruption levels have become popular in the comparative political analysis (Langbein & Knack, 2010). Transparency International’s Corruption Perceptions Index (CPI) ranks countries in terms of the degree to which corruption is perceived to exist among public officials and politicians. It is a composite index aggregated from numerous corruption-related data, including expert and business surveys carried out by various independent and reputable institutions. Kaufmann et al. (1999) attempted to improve on the approach of CPI. They proposed the WGI project to reduce the vast content of the numerous available data on governance quality into a smaller number of aggregate indexes (Langbein & Knack, 2010). The data reflect the views on survey respondents’ governance and public, private, and non-government organizations (NGOs) sector experts worldwide, based on several hundred variables obtained from more than 30 data sources (Kaufmann et al., 2010). Kaufmann and colleagues define governance as: the traditions and institutions by which authority in a country is exercised. This includes (a) the process by which governments are selected, monitored and replaced; (b) the capacity of the government to effectively formulate and implement sound policies; and (c) the respect of citizens and the state for the institutions that govern economic and social interactions among them. (Kaufmann et al., 2010, p. 4)

Areas and dimensions of governance.
WGI and other aggregate indicators are attractive to academics and practitioners because their broad country coverage provides general information about differences between nations. The WGI is a useful tool for measuring the public services’ quality, government independence from political interference, and policy development and implementation (WPP, 2010). Kaufmann et al. (2010) used the unobserved components model (UCM) to construct WGI instead of other possibly more straightforward methods. The researchers claim that this approach has benefits. First, the UCM method of placing data in common units allows some of the information to be stored in the underlying data. However, methods based on countries’ classification by definition only contain information on countries’ relative ranking, but not on the degree of differences between countries. Second, it is less sensitive to extreme outliers. Third, it provides a natural basis for weighting revised indicators based on their relative accuracy, rather than merely constructing unweighted averages as is done by most other cross-country composite measures of governance. Thus, UCM improves the accuracy of aggregated global indicators. Finally, the UCM emphasizes the uncertainty of aggregate management measures by formally formulating the aggregation issue as a signal extraction problem. According to this view, each corruption indicator should be considered, for example, as indicators of noisy or imperfect corruption. Combining these elements can lead to a more informative sign of corruption. However, even these aggregate measures are inadequate, and it is useful to summarize these shortcomings using standard errors and confidence intervals. Finally, Van de Walle (2006) also suggests some benefits from the WGI. First, data aggregation reduces the likelihood of systematic errors. WGI also has extensive coverage, which makes it the most comprehensive dataset. Moreover, having the data procedures, quality assessments, and the data itself hosted by the World Bank may increase its perceived trustworthiness.
However, as WGI measures have become widely used among policymakers and academics, they have attracted some criticisms. Acknowledging that concepts such as the rule of law or corruption are multi-faceted, Van de Walle (2006) notes a trade-off between the reliability achieved by aggregating indicators and the accuracy with the same indicators. This aggregation “comes at the expense of a loss in conceptual precision” (Knack & Manning, 2000, p. 11) and jeopardizes the usefulness of governance statistics. Van de Walle (2006) also argues that, given the nature of government, it is probably impossible to select relevant and comprehensive indicators for different types of governments. Putnam et al. contend that “because governments do so many different things, they have no single ‘bottom line’, like profit in the capitalist firm. This fact opens the possibility that different governments might be simply good at different things” (Putnam et al., 1993, p. 64). Besides, there is a debate that much of the public governance literature seems to be so enmeshed in the diagnosis of increasing societal complexity that they tend to push complex processes and solutions rather than considering more straightforward ones (Triantafillou, 2019). Overall, it appears that WGI has more benefits than disadvantages. WGI attempts to capture the inherent complexity of the concept of governance present in many different data sources and summarize information about governance’s various dimensions. Furthermore, averaging the data on the country-level control accounts for the possible idiosyncrasies between sources.
Like governance, the term corruption has also been defined differently by different scholars, 3 as fighting corruption is central to policy discussions worldwide. Concerned about reacting differently to corruption, Costa Rica (Wilson & Villarreal, 2017), Uruguay (Corleto & Piniero, 2017), and South Korea (You, 2017) are pursuing a series of progressive reforms they hope will eventually lead to dramatic shifts toward greater governmental integrity (Stephenson, 2020). However, legislators have been struggling to adopt effective anti-corruption strategies, particularly between jurisdictions (Goel & Saunoris, 2020). Guidance is missing, in part, due to competing influences (Dimant & Tosato, 2018; Goel & Nelson, 2010; Lambsdorff, 2006; Treisman, 2000, 2007) and the failure to measure corruption properly (Donchev & Ujhelyi, 2014; Sampford et al., 2006; Williams & Siddique, 2008). However, robustness studies have helped researchers and decision-makers propose more effective strategies addressing corruption (Goel & Saunoris, 2020; Jetter & Parmeter, 2018; Seldadyo & de Haan, 2006; Serra, 2006).
The World Bank defines corruption as “the abuse of public office for private gain” (World Bank, 1997, p. 8), and Transparency International labels it as “the abuse of entrusted power for private gain.” 4 Aiming to define the type of corruption practiced by persons of high-level status and income in the society, Sutherland (1939) coined the term “white-collar crime,” described as a “crime committed by a person of respectability and high social status in the course of his occupation” (Sutherland, 1983, p. 7). This definition has been revisited and modified over the years as researchers attempt to determine “the extent to which it occurs in the population as a whole” (Rosoff et al., 2020, p. 18). The economic dimensions of corruption also can be used to differentiate between “grand” and “petty” corruption (Uslaner, 2008).
Corruption is difficult to measure because it is “hidden, and extremely difficult to capture with confidence, accuracy, or a minimal level of resources” (Trapnell, 2015, p. 11). Researchers have tried to measure corruption in four dimensions, each with its advantages and disadvantages (Donachev & Ujhelyi, 2014; Goel et al., 2012; Rohwer, 2009). Goel and Saunoris (2020) describe these dimensions as corruption experience, corruption perceptions, corruption ratings, and corruption prosecutions. Independent of the category, there is no perfect data on corruption, and scholars and governments must use their best available information (Trapnell, 2015). Corruption experience is an ideal measure. However, these data are typically only available in limited surveys that include partial samples of the population and countries. There are two problems with experience as an empirical measure: the uncertainty of the measure and the risk of capturing potential rather than actual corruption threats. Besides, it is costly and time-consuming to collect these data (Goel & Saunoris, 2020). Corruption prosecutions are another measure of exposed corruption and are a complex measure of corruption (Goel & Nelson, 2011, 2014). Nevertheless, this measure partly captures the strength of enforcement and cannot measure undetected corruption, which might partially be captured by perceptions or ratings-based (Goel & Saunoris, 2020).
Corruption perception measures use several aggregate indicators to detect corruption. Potential problems with such indicators are that respondents may not directly experience corruption and the lack of comparability over time (Goel & Saunoris, 2020). Perception indicators are also inherently biased because individuals’ beliefs are subjective (Olken, 2009). In this regard, it is helpful to note the observation made by Olken (2009) that perceptions are biased because individuals’ beliefs are biased and suffer from the perceptions’ convergence problem (Cabelkova & Hanousek, 2004). These scholars attest that people’s perceptions of corruption tend to converge, as they repeatedly hear opinions of corruption from the same sources (media and people in their social circle). Furthermore, acknowledging this bias and understanding how it shapes corruption experts’ perceptions in different countries allows this bias to be adjusted downward in interpreting the results (Bitterhout & Simo-Kengne, 2020). A possible way to address the bias problem is using various corruption perception indices (Bitterhout & Simo-Kengne, 2020) as the control of corruption index from the World Bank. However, Kauffman et al. (2006) argue that, even with these limitations, corruption perceptions remain the most efficient way of measuring corruption. Corruption ratings are based on consistent benchmarks to assess the prevalence of corruption in jurisdictions over time. In this way, they can overcome comparability problems over time, even if criticism remains due to a lack of personal experience of corruption by raters (Goel & Saunoris, 2020). Furthermore, studies of corruption have shown that perceptions of public sector corruption and experience of corruption have been “highly consistent when comparing samples of those who have and those who have not recently experienced…corruption” (Charron, 2016, p. 167), at least in the context of European states.
Lava Jato and Criminal Justice in Brazil
Lava Jato is a Brazilian anti-corruption investigation related to the Petrobras’ corporation, a Brazilian oil company, that has been under investigation by a federal judge since March 17, 2014. Lava Jato is one of the biggest corruption scandals in history (Londoño, 2017), uncovering corruption networks involving billions of dollars across 61 countries and resulting in charges and convictions against 210 people, including politicians, top-level corporate executives, and two former presidents. 5 In summary, Petrobras, a Brazilian state oil and gas company and one of the world’s largest companies, was at the corruption scheme’s center. The corruption worked as follows: when Petrobras issued a call for proposals, it also solicited bribes of around 1%–2% of the contract value. These amounts would be paid from suppliers to Petrobras employees for charging the surcharge, and intermediaries would launder the money before distributing it to Petrobras employees and politicians, who would use the funds to finance election campaigns. Politicians from the parties involved indicated people in their network to assume executive positions and responsibility for selecting Petrobras suppliers. With the formation of a cartel, the engineering firms that would win the contracts were rotated, forging a bidding process without real competition with other firms (Moro, 2018). As a result, the cartel was able to artificially increase the prices charged to Petrobras.
Based on the perceived public scandal of Lava Jato, we tested the hypothesis that Lava Jato had an impact on WGI in Brazil. For instance, we would expect that in the short term (1 or 2 years), the Lava Jato scandal reduced WGI. In the long term, the indicators tend to improve, moved by the adjusted perception of accountability from the investigation, and, eventually, exceed pre-shock levels. There are two main reasons to support this hypothesis. First, given that Lavo Jato drove significant news coverage of corruption in Petrobras and across Brazil’s government, especially in a time with increased news dissemination through social media, public perceptions of corruption will be significant (Charron & Annoni, 2021). This is further supported by the fact that several corruption indicators in Brazil decreased after 2014, 6 except for voice and accountability. The uncertainty resulting from political instability in the country due to Lava Jato’s corruption accusations against political leaders, such as former Brazilian presidents Luiz Inácio Lula da Silva, Dilma Roussef, and Michel Temer (Sedlmeir, 2019), has shown significant variation and reached its lowest rating in 2017. Next, assuming that the investigation results in actions against the responsible actors, continued media coverage of the investigation may facilitate increased perceptions of accountability in the medium to long term if it keeps up pressure for the government to take the corruption seriously (Jacobs & Schillemans, 2016).
Data and Method
To empirically examine if Lava Jato affected WGI in Brazil, we use a synthetic control design, taking a weighted combination of the 135 donor pool countries that optimally fit Brazil’s WGI trend pre-intervention period proxy for a country’s reaction to such an event. SCM (Abadie et al., 2010; Abadie & Gardeazabal, 2003) is often used in empirical research in economics, political science, and other disciplines. The idea that combinations of unaffected units often provide a more appropriate comparison than a single unaffected unit is the synthetic control method’s base (Abadie, 2015). The synthetic control method formalizes comparison units’ choice using a data-driven procedure and opens up precise quantitative conclusions in comparative case studies (Abadie, 2015). Under the right conditions, synthetic controls have significant advantages as a design method for social science research.
Establishing a credible counterfactual, identifying what would have happened if Lava Jato had not been introduced, is difficult. To analyze the effects of Lava Jato on the development of WGI, we compare Brazil’s performance with outcome-specific synthetic control groups. Synthetic control groups consist of weighted averages of countries not linked to Lava Jato. In this paper, we use synthetic control methods (Abadie et al., 2010) to compare Brazil to a plausible group of control countries without Lava Jato. For each World Bank’s six composite indicators of WGI (control of corruption, voice and accountability, political stability and absence of violence/terrorism, government effectiveness, regulatory quality, and the rule of law), we construct a separate synthetic Brazil since the relative influence of observable and unobservable characteristics is likely to differ across outcomes. This approach provides a data-driven way of obtaining an optimized weighted average of the control countries to track Brazil in terms of WGI outcomes before Lava Jato. Resulting synthetic Brazil is then used to simulate the country’s hypothetical trajectory in the absence of Lava Jato. For each outcome variable, we create a separate synthetic Brazil. This approach is similar to other studies that have used synthetic control methods with multiple outcomes (Bohn et al., 2014; Bove & Elia, 2014; DeAngelo & Hansen, 2014; Fletcher et al., 2015; Quast & Gonzalez, 2016; Rieger et al., 2017).
Multiple SCM are applied to identify Lava Jato’s causal effect on each of the World Bank’s six composite indicators of WGI: control of corruption, voice and accountability, political stability and absence of violence/terrorism, government effectiveness, regulatory quality, and the rule of law. It is used for those indicators the percentile rank among all countries, ranging from 0 (lowest) to 100 (highest). This article focuses specifically on the control of corruption, which is an aggregate of 22 sub-sources. Corruption control measures different aspects of corruption through a globally comparable method (Cary et al., 2014). The analysis is utilized in a country-level panel dataset containing annual data spanning 2002–2018. The year 2002 was chosen as the beginning year because this was when the WGI started to be computed annually. Before then, the information was added biannually. Additionally, we calculate the lagged version for each of those indicators.
SCM allows us to evaluate the net impact of Lava Jato’s intervention on perceptions of corruption in Brazil against the counterfactual scenario in which the country lacked the anti-corruption prosecution. This quasi-experimental design is an extension of “difference-in-differences” models, aiming to estimate an intervention’s causal effect by computing the distance between two-time series after that intervention (Bartos & Kubrin, 2018). The major attraction of this technique stems from the ability to address severe econometric issues such as heterogeneity. When the number of pre-intervention periods in the data is large, the correspondence between the pre-intervention results helps control the unobserved factors and the heterogeneity of the observed and unobserved factors in the result of interest. As a result, only the same units in the observed and unobserved determinants of the outcome variable produce similar outcome variable trajectories over long periods. Once the case of interest and the synthetic control unit behave similarly for long periods before the intervention, a discrepancy in the outcome variable after the intervention is produced by the intervention itself (Abadie et al., 2015).
SCM offers a unique toolset to contribute to the governance and corruption literature beyond the methods most often employed, like case studies. While there have certainly been other investigations following major white-collar scandals, these often occurred in national and temporal contexts that make strong comparisons to Brazil in the 2010s difficult. For instance, the “Mani pulite” or “clean hands” investigation in Italy in the 1990s was similarly extensive, involving hundreds of politicians, bureaucrats, the business community, and political parties after an extensive network of kickbacks and bribes were uncovered in the government contracting process (Vannucci, 2009). However, major differences in timing that predate the WGI, national political context, and in WGI measures even today in Italy make Mani pulite alone a less apt comparison when measured against the robustness of several dozen nations as synthetic comparison cases. While the contexts of cases are invaluable, we can better control for significant differences that otherwise would exist between just two nations using SCM.
This method also has advantages compared to other approaches. Both the SCM and the difference-in-difference method focus on the difference between treated and untreated units concerning an intervention (Galiani & Quistorff, 2016). However, they differ because the SCM assigns untreated weights that are not all the same. Instead, the control units most similar to the treated unit in the pre-treatment period receive higher average weights. Then, the unit of interest’s counterfactual is estimated for the post-treatment period from the assigned weights. SCM also avoids extrapolation biases present in time series results, limiting the linear combination coefficients that define the synthetic control between 0 and 1 (Abadie et al., 2015). Another advantage is the possibility of inference through tests with a placebo (Abadie & Gardeazabal, 2003; Abadie et al., 2010). Finally, SCM predicts these unobserved variables considering the observed and unobserved variables’ linear relationship in the pre-intervention period (Abadie et al., 2010). Some scholars explain the mathematical terms regarding the SCM and inference methodology (Abadie et al., 2010; Cavallo et al., 2013; Galiani & Quistorff, 2016).
Brazil’s synthetic control group is a weighted combination of donor pool countries that optimally fits Brazil’s WGI trend for the pre-intervention period (from 2002 to 2014). We expect to select a better comparison unit than any individual country by fitting the synthetic control groups over the pre-intervention time series containing 13 years of corruption and governance indicators before deploying the Lava Jato investigation. The strategy to match on a long (n = 13) pre-intervention time series would reduce the likelihood of identifying a spurious effect if compared with synthetic control group models matched on fewer pre-intervention observations (Abadie et al., 2010; Bartos & Kubrin, 2018; McCleary et al., 2017). Conversely, fitting the models on longer time series that exhibit relevant white-noise variation results in the optimization routine less likely to converge on a perfect approximation of pre-Lava Jato. However, it is much less likely to identify a spurious effect than models fit on smoother or shorter pre-intervention time series (McCleary et al., 2017).
It is worth considering that having multiple SCMs for multiple outcome measures could raise a practical concern of overfitting. The use of multiple different synthetic controls to fit each variable can overfit what is known about the model. We tested the data with a smaller training period to validate the SCM and reduce overfitting concerns based on Abadie (2021). The author proposed to split their pre-treatment data into a training period and a validation period.
Methodology Specification
Following Abadie et al. (2011), let
First, the Non-Restricted Donor Sample method (NRDS) populates the “donor pool” with all the 187 countries in the database. It is then included in the “donor pool” just the countries not listed in the Lava Jato investigation. According to Villar and Papyrakis (2017), this second method is named the Restricted Donor Sample method (RDS). We expect that excluding from the donor pool all countries in Lava Jato would avoid contamination by contributing to a treated donor pool country. Therefore, it includes the remaining 135 countries in the donor pool from which Synthetic Brazil is constructed. Table 1 displays the weights of the donor pool countries using the RDS for the three variables that produced significant results: control of corruption, the rule of law, and regulatory quality. Table 2 presents the WGI’s mean values during the pre-intervention period (2002–2014). We can see that synthetic Brazil RDS values provide a better approximation to Brazil for all the six WGI. For this reason, all the following analyses are based on RDS.
Synthetic Brazil RDS.
Mean of Worldwide Governance Indicators (2002–2014).
Note. NRDS = Non-Restricted Donor Sample method. RDS = Restricted Donor Sample method.
We employ the data-driven approach for assigning donor pool weights 8 to minimize the distance between Brazil and its counterfactual indicators trends throughout the pre-intervention time series (Abadie et al., 2010, 2015; Bartos & Kubrin, 2018). When a gap emerges between Brazil and its synthetic version after Lava Jato started, the difference between the two-time series is interpreted as the causal effect of Lava Jato on the indicators examined. The quality of the match between Brazil and Synthetic Brazil across the pre-intervention period will determine the causal interpretation gaps.
We use the “Synth” routine 9 to fit the models. The approach considers the root mean squared prediction error (RMSPE) term to describe the pre-intervention fit’s quality. We assume that no effect beyond the matching error can be identified if the gap between Brazil and its counterfactual that emerges post-Lava Jato is within the range of the pre-intervention RMSPE. However, if the gap post-Lava Jato is out of the range of the pre-intervention RMSPE, it doesn’t necessarily mean the estimated effect is of practical significance. When the precision of the pre-intervention fit between Brazil and Synthetic Brazil is excellent, “a post-intervention gap that is small relative to the observed variation in the pre-intervention time series can result in an effect size that is an order of magnitude greater than the pre-intervention RMSPE” (Bartos & Kubrin, 2018, p. 701). As a consequence, we can identify smaller treatment effects when the pre-intervention fit is more precise.
One of the disadvantages of SCM is that it does not give standard errors to assess whether the results are statistically significant. To solve this problem, we use in-sample placebo tests 10 to identify erroneous inference. The rationale for placebo experiments is to test whether the intended treatment effect can be completely random. This approach provides a type of randomization inference (Abadie et al., 2010; Abadie & Gardeazabal, 2003; Bartos & Kubrin, 2018; Fisher, 1922; McCleary et al., 2017). The treatment condition to each donor pool country is iteratively reassigned and constructed a synthetic control group. The countries are then ranked by a ratio of the 2015–2018 gap to pre-2015 RMSPE. In other words, what we want to know is whether Brazil’s treatment effect is extreme, which is a relative concept compared to the donor pool’s placebo ratios. We would expect, if Brazil is the highest among the collection of countries, the estimated effect is greater than the unidentified variation observed in the donor countries. However, if Brazil does not have a high rating, the estimated impact is not significant with the noise displayed by countries in the donor pool. For instance, it wouldn’t be worth focusing the analysis on results that the estimated effect is not significant because these results would suggest relatively small magnitude effects.
We also evaluate whether an estimated effect is sensitive to changes in synthetic Brazil. This approach is iteratively applied, excluding the donor pool. In each iteration, the country that contributes the most considerable weight to the synthetic model is removed until all of the original donor pool countries with non-zero weights are excluded from the matching algorithm. This approach will result in a synthetic Brazil composed of a different set of donor pool units than the original synthetic model. Suppose the original effect persists in sign and magnitude once all of the original donor pool countries have been excluded. In that case, it is possible to claim that this effect is insensitive to synthetic Brazil’s composition changes. As a result, the interpretation of the Lava Jato operation’s impact on corruption and governance indicators would not change, even if there were substantial changes in synthetic Brazil.
Results
To estimate the impact of Lava Jato’s investigation on the WGI in Brazil, we applied SCM
11
for each governance indicator. A large treatment effect would occur if the two trajectories of the outcome variable for Brazil and synthetic Brazil were quite similar before the intervention and diverge abruptly when the intervention arises. Figure 2 displays the courses of Brazil (treated unit) and each synthetic control unit
12
for the six composite indicators of WGI: Control of corruption The rule of law Regulatory quality Government effectiveness Voice and accountability Political stability and absence of violence/terrorism

Trends in worldwide governance indicators: Brazil vs. Synthetic Brazil.
In this case, Figure 2 shows that Lava Jato appears to affect 13 corruption, the rule of law, and regulatory quality. For those indicators, the gap that emerged after Lava Jato was more significant than the model’s pre-intervention RMSPE. However, Brazil does not have high gaps after Lava Jato for the indicator’s government effectiveness, voice and accountability, and political stability and absence of violence/terrorism, suggesting relatively small magnitude effects. There are possible explanations for those outputs.
The control of corruption trajectory presented in Figure 2 shows a large decrease after Lava Jato’s enactment. Control of corruption captures the perceptions of how public power is exercised for private gain, including both petty and grand forms of corruption, and the perception of the extent to which the state is captured by elites and private interests (Kaufman et al., 2010). This trajectory can be understood as one of the effects of the exposure of grant corruption cases resulting from the widespread dissemination of Lava Jato by the media that helped to unveil that the problem of corruption was underestimated in the previous years. Moreover, the disclosure of corruption scandals also reveals a particular organization’s weakness in having instruments that effectively prevent and punish corruption acts. There is a deterioration in the control of corruption in the short term. Then, in the medium and long terms, the indicator would rise due to the increase in society’s trust in the control bodies, the improvement of the public and private sectors’ internal instruments to fight corruption, and the sufficient punishment of those involved in illegal acts. It seems that the indicator had a reversal trend in 2018, probably provided by the perception that the institutions accountable to detect, investigate, and punish acts arising from corruption are working. However, it is necessary to have data for the following years to confirm this trend. Another point to consider is the distribution pattern of control of corruption during the years. Considering the mean, quartiles, skewness, and kurtosis, it seems that the distribution is stable and normal. 14
The rule of law and regulatory quality trajectories are also very similar to synthetic Brazil for the pre-period, and they deteriorate after Lava Jato arises. The rule of law captures the perceptions of the extent to which agents trust and abide by society’s rules, the likelihood of crime, and violence and the quality of contract enforcement, property rights, police, and courts (Kaufman et al., 2010). Lava Jato seems to have reduced society’s confidence in the legal structure and compliance with the law. Perhaps this effect is explained by the reality shock caused by the exposure of an endemic corruption scheme impregnated in the government and private entities’ relations. Regulatory quality captures perceptions of the government’s ability to formulate and implement sound policies and regulations that encourage private sector development (Kaufman et al., 2010). It seems that Lava Jato similarly affected this dimension of governance. Because these dimensions of governance should not be thought of as being somehow independent of one another, we assume that in the short run, Lava Jato directly deteriorated society’s trust in the government, eroding the capacity of the government to formulate and implement policies effectively. Since the creation of regulatory agencies between 1996 and 2005, the regulatory system has been involved in a constant debate on improving its legal framework. However, the most significant changes in the system occurred beyond the timeframe of this research. Only in 2019, with the publication of Law 13,848, 15 which addresses the management, organization, decision-making process, and social control of regulatory agencies. In 2020, there was another advance in the regulated sectors with decree 10.411, 16 which establishes the regulatory impact analysis framework. As we can see from Figure 2, the indicators of government effectiveness, voice and accountability, political stability, and absence of violence/terrorism don’t present significant magnitude effects.
There was little to no apparent impact on the WGI indicators for government effectiveness, voice and accountability, and political stability. One potential reason for this may be that while Lava Jato highlighted deep corruption in the construction and oil industries and received a great deal of public attention, it does not directly affect citizens’ daily activities. Because Petrobras is a public-owned firm whose operations are seen as important but distinct from traditional notions of governance, this may explain Lava Jato’s low impact on government performance. The voice and accountability indicator reflects the extent to which a country’s citizens can participate in their government’s election and freedom of expression, association, and media freedom. There is no strong reason that Brazil does not have those characteristics or that these measures would be especially material to the present case. It is plausible to assume that anti-corruption campaigns, like Lava Jato, would not substantially impact free speech and free press in the country. The indicator of political stability and non-violence/terrorism reflects the likelihood of destabilizing or overthrowing a government unconstitutionally or violently, including politically motivated violence and terrorism. There is no evidence of a high probability that the Brazilian government will be destabilized by unconstitutional means. As a consequence, we expect a small effect of Lava Jato on this indicator.
The government effectiveness indicator reflects the perception of the public services’ quality, the quality of public services, and their level of independence from political pressure. Taken together, education, health, and the environment are critical areas that affect people’s everyday lives (Parker et al., 2004). Anti-corruption efforts in these areas can have a more direct and short-term impact and receive more political support than programs in other sites that are further removed from people’s daily activities.
Figure 3 shows how the gap between treated and synthetic control outcomes changes over time. The control of Brazil’s corruption trajectory is very similar to its synthetic unit for almost the entire pre-Lava Jato period. Once Lava Jato arose in 2014, however, the control of corruption trajectory in Brazil depreciates at a much higher rate than in synthetic Brazil, suggesting a large negative effect of Lava Jato on Brazilian’s control of corruption. Our results suggest that for the entire 2014–2018 period, control of corruption reduced 20 points, a decline of approximately 33%. A similar reduction occurs in the rule of law and regulatory quality. The rule of law decreased 10 points, a decline of 18%, and regulatory quality reduced 15 points (27% of decline).

Worldwide governance indicators’ gap between Brazil and synthetic Brazil.
The indicators for government effectiveness, political stability, and voice and accountability did not show a good fit in the pre-intervention period. For instance, their effects in the post-intervention should be interpreted with even more caution. Government effectiveness, and political stability, declined 20%, 12%, respectively. Voice and accountability had a slight increase of 6%. However, it is not worth focusing the analysis on results that indicate that the estimated effect is not significant, as these results suggest effects of relatively small size.
Figure 4 shows the permutation test results from the gaps associated with each iteration to control units in the sample. Following Abadie et al. (2010), countries with a poor fit for the pre-treatment period (regions with Pre-Lava Jato MSPEs more than five times higher than the MSPE for Brazil) were excluded, and the treatment to each country was randomized by re-estimating the model and calculating a set of MSPE values for the pre- and post-treatment period. In other words, the treatment to each country was reassigned, putting Brazil back into the donor pool each time, estimating the model for that “placebo,” and recording information from each iteration. We intend to know whether Brazil’s treatment effect is extreme, a relative concept compared to the donor pool’s placebo ratios. Figure 4 signs that Brazil is in the tails of some distribution of treatment effects to control corruption, the rule of law, and regulatory quality.

Permutation test: WGI gaps in Brazil and 174 countries.
Significance and Sensitivity Tests
We performed in-sample placebo tests for the indicators control of corruption, the rule of law, and regulatory quality as they showed a good fit in the pre-intervention period. The in-sample placebo tests determine whether the estimated effects of Lava Jato are statistically significant, relative to the unidentified annual variation observed in countries that did not experience Lava Jato. Suppose the test identified more than 13 countries in the donor pool that produced more significant treatment effects than Brazil. In that case, the likelihood of identifying an impact equal to or greater in magnitude than Brazil is greater than 0.1 and would not be significant. Figure 5 displays Brazil’s ratio of the post-intervention gap to pre-intervention RMSPE relative to the donor pool countries. Control of corruption gives an exact p-value of 0.058 (eight out of 136), which is higher than the conventional 5% most journals want to (arbitrarily) see for statistical significance. The rule of law (15 out of 136; p = ∼.11) is not statistically significant. The change in the gap size between the treated and the synthetic regulatory quality index has statistical significance. Brazil ranks particularly highly for regulatory quality (five out of 136; p = ∼.04), suggesting that Brazil’s estimated effect appears significant. Figure 6 presents the RMSPE’s highest ratios for control of corruption. It points out that Samoa, Kazakhstan, and Croatia have a low pre-RMSPE, which signals an excellent fit in the pre-intervention phase and a large post-RMSPE. As a result, these countries have pretty large RMSPE ratios. 17

In-sample placebo test RMSPE ratios ranked.

RMSPE and ratio for Samoa, Kazakhstan, and Croatia.
To determine whether the estimated control of corruption, the rule of law, and regulatory quality effects are sensitive to changes in Synthetic Brazil’s composition (i.e., different donor pool weights), the donor pool country with the greatest weight (ω) was iteratively excluded until all of the original donor pool countries with non-zero weights have been removed. The version of synthetic Brazil that results from this procedure comprises a set of donor pool countries that are entirely different from our original model. Suppose the estimated impact of Lava Jato on Brazil’s indicators persists under both compositions. In that case, it is possible to be confident that those estimates are not dependent on certain donor pool countries’ contributions to synthetic Brazil. If the interpretation changes under synthetic Brazil’s new composition, the estimated effect depends on certain donor pool countries’ contribution, and the finding should be interpreted cautiously.
The results of our sensitivity test for control of corruption, the rule of law, and regulatory quality are displayed in Figure 7. In addition to Brazil and synthetic Brazil using the RDS method (as seen in Figure 2), Figure 7 also shows a series of alternative specifications for Synthetic Brazil. It is included an additional alternative Synthetic Controls by excluding from the donor pool communist countries (Cuba, Lao, and Vietnam) and countries with a population in 2002 of less than one million. 18 The effect is sustained for all the alternative specifications for control of corruption, the rule of law, and regulatory quality. This suggests that the valid causal interpretation of the Lava Jato effect on control of corruption, the rule of law, and regulatory quality is independent of the validity of including specific countries in our donor pool. Thus, control of corruption and the rule of law appear to be independent of particular countries’ contributions from the donor pool.

Synthetic control methods.
Overfitting may affect the post-intervention trajectory’s prediction power in the case of many units in the donor pool and a small number of pre-intervention periods (Abadie, 2015). Although the data has a long (n = 13) pre-intervention time series, there is always the trade-off between the desire to have an adequate fit and concerns about over-fitting. We apply the procedure proposed by Abadie (2015). First, we divided the pre-intervention periods into a training period and a validation period. Then, we use data from the training period to identify predictor weights to be used in the validation period. Table 3 presents the results in terms of squared RMSPE for the different synthetic control methods.
RMSPE for the Different Synthetic Control Methods.
Note. A: Pre Lava Jato RMSPE squared; B: Post Lava Jato RMSPE squared; B/A: Ratio (Post/Pre Lava Jato).
Discussion
To summarize the findings, the initial synthetic control estimates suggested decreases in control of corruption, the rule of law, and regulatory quality after Lava Jato’s enactment. The effect survives both significance testing (randomization inference) and sensitivity testing for the indicators control of corruption and regulatory quality. The rule of law survives for sensitivity testing but doesn’t pass significance testing. There is no evidence of a statistically significant robust change post-Lava Jato operation for the remaining variables (government effectiveness, political stability, and voice and accountability).
This study is not without limitations. Although no other country implemented an anti-corruption task force that is wholly comparable to Lava Jato within our analysis time frame, 19 a diverse body of country-level anti-corruption reforms has been implemented worldwide. At least some countries have likely employed nationwide judicial investigations into political corruption that are comparable, in some part, to Lava Jato. Suppose Synthetic Brazil is constructed with a donor pool unit that partially experienced a Lava Jato-like intervention. In that case, both trends will reflect the impact of the shared aspect of Lava Jato. The gap would then reflect Lava Jato’s effect on the outcome indicators beyond what was caused by the shared aspect of Lava Jato, producing a more conservative estimate of the impact.
Some issues related to Lava Jato remain to be understood. For example, even years after Lava Jato, Brazil’s corruption control remains low, posing credibility challenges. Although the operation led to the conviction of defendants and raised corruption at a national level, critics argue that investigative procedures or even the legal process are biased and were used for political purposes. The findings suggest that Lava Jato increased corruption perception in the short term and can increase corruption control in the long term. Perhaps the persistence of the low control of corruption is related to criticisms about the conduct of the Lava Jato process or even about the alternative punishments implemented, such as house arrest and reduced sentences resulting from a plea bargain.
Concluding Remarks and Policy Implications
This study represents the first systematic analysis of Lava Jato’s impact on WGI throughout Brazil after the Lava Jato operation’s deployment. With country-level panel data from 2002 through 2018, a multiple synthetic control group design was employed to approximate Brazil’s indicators if Lava Jato did not exist. The findings reveal that Lava Jato affects corruption, the rule of law, and regulatory quality. Simultaneously, the results suggest that government effectiveness, political stability, and voice and accountability appear not to be impacted by Lava Jato. Furthermore, the results show Lava Jato has a largely negative impact on the indicators control of corruption, the rule of law, and regulatory quality. Control of corruption reduced 20 points, a decline of approximately 33%. The rule of law decreased 10 points, a drop of 18%, and regulatory quality reduced 15 points (27% of decline). In the short term, the Lava Jato scandal reduced those indicators. While there was some slight improvement in the measures for control of corruption and the rule of law in the last year where data is available, it remains to be seen whether these will improve in the long term with a U-shaped recovery.
We would expect in the short-term deterioration in the control of corruption and an expectation that, in the medium and long terms, control of corruption would raise by the increase in society’s trust in the control bodies, improvement of the internal instruments of the public, and private sectors to fight corruption and the sufficient punishment of those involved in illegal acts. It seems that this turning point may have been in 2017–2018, when corruption control increased. However, the progress or regress in the anti-corruption campaign is often determined by the extent to which corruption fights back. If we look at events in 2019 and 2020, the prospects for improving corruption control are not very bright. The Brazilian Supreme Federal Court’s decision to invalidate the provisional execution of the sentence in the second instance enormously weakened Lava Jato and allowed criminals, including former President Lula, to be released. The Supreme Federal Court’s decision was even more controversial because the existence of four instances of trial, peculiar to Brazil, associated with the excessive number of appeals, results in delay and prescription, leading to impunity. Other proposals like the guaranteed judge and the law of abuse of authority just contribute to increasing uncertainty and complexity in a legal framework that is already complex. When available in the future, the inclusion of more recent WGI would help confirm our insights.
Another point of consideration is that there is little doubt that our countrywide analysis masks significant variation at the local level. It is also worth determining whether Lava Jato’s impact on corruption perception varies across Brazil’s 26 states and the Federal District, each with different socioeconomic, demographic, and criminal justice profiles.
This analysis has implications well beyond Lava Jato and Brazil. Although Lava Jato is specific to Brazil, it has reverberations in 52 countries in the sample. Furthermore, the country’s steps to fight corruption are closely watched by other countries also confronting similar challenges in grant corruption and white-collar crime. Lava Jato is an informative case demonstrating how corruption or perceptions of corruption can be influenced. Corruption in the country is systemic and is entrenched in the public and private sectors. To change this reality, the powerful must be afraid of committing acts of corruption. Investigations may serve this purpose. Moreover, the sufficient punishment translated in the imprisonment of the convicted corrupt, return of the deviated values, and an additional payment of fines generate the deterrence effect and demonstrates unequivocally that the country does not corroborate with impunity. Although more research is necessary, initial findings suggested that Lava Jato decreased perceptions of control of corruption in the short term and may increase perceptions of corruption control in the medium/long term. A crisis can open opportunities to take actions never done before (Petersilia, 2016). A real change in managing the state should consider four crucial reforms: political, judicial, administrative, and tax. These reforms are areas in which it is necessary to establish a well-planned strategy for action, which considers all stakeholders and affected, lest corrupt actors react negatively and seek ways to maintain the status quo.
In the context of political reform, proposals from reducing the number of congress members, their salaries and benefits, and assessing the real need for the Electoral Fund should be debated. However, there is no evidence that such matters will be addressed in the short and medium terms. Concerning the judiciary, it would be necessary to evaluate how the courts use their budgets, the resumption of prison-based convictions in the second instance, in addition to the need to revisit the salaries and benefits of civil servants. Four instances of trial, peculiar to Brazil, associated with the excessive number of appeals, results in delay and prescription, leading to impunity.
In the administrative reform, a more explicit discussion of the state’s role, the simplification of procedures, and the revision of its institutional design, costs, and governance instruments should be at the heart of the debates. The proposed administrative reform that the government of President Jair Bolsonaro has defended seems to be addressing, at least in part, these issues, like proposals to lean the number of government careers. It is worth highlighting the government’s effort to reduce the number of political nominations in senior management positions. However, it seems incipient discussions to improve public servant allocation, public servant evaluation, public manager support for decision-making, end of progressions by the length of service, end of management compensation incorporations, and low-performance regulation dismissal. Furthermore, a reform should be necessary that impacts the three levels of powers, states, and municipalities. The benefits provided to the judiciary servants are far from reasonable. For example, judges and prosecutors have 60 days of vacation, compulsory retirement as “punishment,” and attendance leave.
Finally, tax reform is urgent to simplify and optimize its tax burden, which is the most complex in the world. According to the World Bank Open Data, Brazil spent the longest of 235 countries worldwide to prepare, file, and pay (or withhold) taxes with 1958 hours (2019 statistics) compared to a world average of 234 hours. 20
Besides the fact that this paper is the first to investigate the impact of Lava Jato on WGI, other research strategies could be implemented. One approach would be to come up with a single synthetic control for all six outcome measures. For example, in examining a crime reduction program, Saunders et al. (2015) looked at the changes in several different outcome variables of interest using a single weighted synthetic control. Additionally, although RSMPE seems to be good enough to assess the model performance, Parast et al. (2020) indicate that Absolute Standardized Mean Difference (ASMD) could be a better indicator of fit and a more sensitive test. Further investigations could include ASMD in the analysis. Alternatively, the approach proposed by Kinn (2018) seems promising. The author offers a general scheme designed to predict the counterfactual by minimizing the trade-off between underfitting (bias) and overfitting (variance). Based on structural and reduced-form machine learning approaches, the framework seems suitable for single treated time series and a high dimensional pool of control time series that may exceed the number of time periods.
It is challenging to make any causal statement about the impact of anti-corruption measures on governance. The assessment of anti-corruption measures’ impact is challenging because it is hard to measure the dependent variable and a relative dearth of data directly measuring corruption. But SCM offers new possibilities and contributions to the study of corruption and white-collar crime across national contexts.
Footnotes
Acknowledgments
We thank Kenneth Sebastian León, Anne Alvesalo-Kuusi, and two anonymous reviewers for their constructive comments and suggestions that have greatly improved our article. We also thank Bradley J. Bartos for his kindness in sharing his knowledge. All views and remaining errors are our own.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
