Abstract
Performance indicators have had to endure severe criticism. They are said to lack accuracy, encourage gaming and ultimately fail to improve performance. Yet, despite their well-documented weaknesses, performance indicators abound in governance. This article asks under which conditions performance indicators can improve performance outcomes, despite these proven weaknesses and dysfunctions. Our case study is the stress test of the European banking system, a high-profile performance indicator used for risk regulation. Based on interviews with risk managers in Belgian banks as well as staff at the European Central Bank, the European Banking Authority and the National Bank of Belgium, we find that the process of calculating the stress test improves performance outcomes in itself. It does so by fostering banks’ capacity to self-regulate, tying into Foucault’s notion of governmentality. As such, practitioners and academics should not only pay attention to how performance results can be used, but also examine how the process of calculating the performance indicator might be designed to improve performance outcomes latently.
Introduction
Since the global financial crisis, stress testing has become part and parcel of regulators’ toolkits for monitoring and maintaining financial stability globally. In Europe, the European Banking Authority (EBA), in close coordination with the European Central Bank (ECB), now conducts system-wide stress tests to simulate the ability of the big, so-called systemic, banks to deal with an adverse scenario, making it easy to rank banks from worst to best ‘performer’.
At first sight, the success of stress tests is easy to explain. As awareness for risks increases, the demand for instruments that measure and control risks is growing (Power, 1997). Performance indicators’ success hinges on their ability to simplify, standardise, compare and control complex systems. The European Union (EU)-wide stress test delivers on this demand for control. The stress tests are designed to show at a glance the banks that would weather a crisis.
However, decades of research have also exposed the adverse effects of performance indicators. They are said to lack accuracy, encourage gaming, demotivate workers, be biased towards what is quantifiable and ultimately fail to substantially improve performances (Berten and Leisering, 2017; Bevan and Hood, 2006; Bouckaert and Balk, 1991; Davis et al., 2012; Pollitt, 2018). Research has also identified the misuse of performance information as a key factor hampering performance improvement (Taylor, 2011; Van Dooren et al., 2015). Some scholars go so far as to claim that ‘evidence-based’ and ‘rational’ policymaking is a myth because policy work is fundamentally political (Boswell, 2015). Adverse effects are primarily at play when highly incentivised performance indicators are used (Bevan and Hood, 2006). The stress tests have not been spared this criticism, especially in financial media. Stress tests are said to not make sense economically, be biased towards certain banks and be little more than communication exercises to reassure financial markets (Cecchetti and Schoenholtz, 2016; Dowd, 2015; Elliott, 2016).
To be sure, there are studies that show how performance measurement positively affects performance outcomes (Boyne and Chen, 2006; Nielsen, 2014; Walker et al., 2011). Yet, we do not have a clear understanding of why performance indicators improve performance outcomes in some cases and fail to do so in others.
Despite these differing and seemingly contradictory effects of performance indicators, regulators continue to promote them as indispensable tools for regulation. Over the past decades, the use of performance indicators has proliferated in (global) governance (Davis et al., 2012). The ambition of this article is thus to explain the enduring success of performance indicators, and explore under which conditions they can improve performance outcomes, despite their proven weaknesses and dysfunctions.
We find that performance indicators have important latent functions that have so far been under-studied in the literature. We borrow the concept of a latent function from Merton’s (1968) functional sociology. Merton distinguishes between manifest functions (intended and positive effects), dysfunctions (unintended and negative effects) and latent functions (unintended positive effects). The performance literature typically deals with the manifest functions and dysfunctions of performance information. In this article, we take a closer look at the latent functions.
Our case study studies the latent functions of the stress test. First of all, we show that the process of calculating the stress test latently improves performance outcomes. To complete the stress test in a timely fashion, banks have professionalised their internal risk-management systems, investing in enhanced data quality, improved information technology (IT) systems and better coordination between risk domains. This inadvertently improves their performance outcomes. Second, we find that previously recognised latent functions, such as the ritualistic and symbolic function of performance indicators (Boswell, 2015; Power, 1997), are much more important for performance outcomes than they are often ascribed to be. Finally, we find that certain dysfunctions, such as the inaccuracy of results, do not necessarily hamper performance outcomes.
Manifest functions of performance indicators
Indicators abound in public governance. New Public Management (NPM) reforms in particular have led to the dissemination of indicators in all corners of government (Van Dooren et al., 2015). NPM aggregates a plethora of policy principles that all in some way intend to improve public sector performance by making it more efficient and goal-centred (Van Dooren et al., 2015). It does so through the principles of disaggregation, competition and incentivisation (Dunleavy et al., 2005). For performance-based regulation to work, a key condition is that performance needs to be measured, and this is where performance indicators enter the picture. Many of the incentives that NPM promotes can only be applied when quantitative performance indicators are available.
Performance indicators easily found their way to the regulators’ toolbox as they provide information at a glance. They provide an objective measure of which organisations are reaching targets and who is underperforming. This information can then be used to improve performance outcomes by fostering competition between organisations, allocating resources according to performance and increasing accountability (Braithwaite, 2014; Kagan, 1995; Levi-Faur, 2005). Moreover, the simplicity of performance indicators allows for more succinct communication with actors inside and outside government, tapping into an agenda of transparency and accountability, at least in theory (Sarfaty, 2011).
Dysfunctions of performance indicators
Despite the aforementioned noble intentions, and high hopes, an increasing number of critical voices cite the paradoxical and dysfunctional effects of NPM and performance indicators, calling for a post-NPM reform (Christensen and Fan, 2016; Klenk and Reiter, 2019; Mikuła and Kaczmarek, 2019; Reiter and Klenk, 2018). After more than three decades of performance measurement in public policy, even sympathetic analysts like Hood and Peters (2004) or Dunleavy et al. (2005) acknowledge the adverse effect of NPM reforms, especially performance indicators (Pires, 2011). Although in some (predominantly developing) countries, NPM reforms are still playing out, most advanced countries have come to realise that NPM has not fostered more effective or efficient public organisations. Instead, the NPM themes of disaggregation, competition and incentivisation led to siloed public bodies impeding collective action, perverse quasi-market mechanisms and an obsession with intermediate organisational targets overshadowing service delivery and effectiveness (Dunleavy et al., 2005). In line with this, performance indicators specifically received a number of criticisms as well.
A first dysfunctional effect is that performance indicators may lead to tunnel vision. They tend to focus upon easily quantified dimensions of performance, thereby narrowing down the focus of policymaking and political debate to a small and often unrepresentative aspect of policy (Bevan and Hood, 2006; Pidd, 2005; Power, 1997; Termeer et al., 2013). Doig, McIvor and Theobald (2006) add to this that an overreliance on scores and rankings might overlook the fact that the phenomena they intend to depict are moving targets in terms of progress and direction. In the stress test, easily quantifiable risk areas such as credit and market risks have been addressed substantially, while areas that are more difficult to quantify, and difficult to pin down and define, such as operational risk, are less developed.
Besides this, performance indicators can create perverse incentives and encourage ‘gaming’ and cheating. There are many empirical examples of how data are manipulated (Bevan and Hood, 2006; Hood and Peters, 2004; Pollitt and Talbot, 2004; Smith, 1995; Stone, 2002). For instance, hospitals will cancel appointments or schedule fewer follow-up meetings to cut down waiting lists, creating an illusion of efficiency. Indicators are also said to stifle curiosity and diminish learning opportunities (Radin, 2006). Moreover, performance indicators may lead to goal displacement, where organisations focus on the indicators rather than the underlying objective that the indicators are supposed to measure (Bohte and Meier, 2000). Furthermore, as O’Neill (2002) and Power (1997) have shown in their work on audits, performance indicators often obscure what is actually happening in the workplace, fuelling suspicion and mistrust, undermining professional ethics, and generating a host of unforeseen problems.
Finally, the information that performance indicators produce is often not even used or applied in decision-making (Johnston, 2004; Mol and De Kruijf, 2004; Pollitt and Talbot, 2004; Taylor, 2011; Walshe et al., 2010). Frequent causes are the insufficient quality of the performance information and the lack of important data, but also cultural or institutional barriers (Hoogenboezem, 2004; Van Dooren et al., 2015). De Vries (2010) adds that performance measures are usually a-contextual and unable to reveal anything substantive about the quality of politics. Performance indicators are just more red tape and paperwork, wasting away in binders and computer folders.
This leaves us with a puzzle: with so many dysfunctions from research and practice being documented, under which circumstances can performance indicators actually improve performance outcomes? We claim that the answer is to be found in the latent functions of performance indicators.
Methodology
We aim to explain under which conditions performance indicators, such as the stress test, can improve performance outcomes, despite their proven weaknesses and dysfunctions. We operationalise this by studying the contribution of the stress test to risk performance according to regulators (the National Bank of Belgium (NBB), the ECB and the EBA), regulatees (the Belgian banks) and intermediate organisations (consultancy firms). We aim to understand how these actors make sense of and give meaning to the stress test as a tool of governance. We selected the EU-wide banking stress test because it is an eminent example of a performance indicator that has placed itself at the centre of risk regulation in the wake of the 2008 financial crisis. Moreover, it is a fairly recent indicator (only a decade old) and, as such, its design is still subject to yearly changes. We are interested in how the stress test has developed over time and how the different iterations of the stress test have affected performance outcomes, according to those involved.
We conducted 45 conversational interviews with 33 people. The interviews lasted 75 minutes on average. We did a first round of 18 interviews in Belgian banks and four in consulting firms in 2015/2016; a second round of eight interviews was done at the ECB in 2017; and in a third round, we conducted nine interviews in banks, one in a consulting firm, three at the EBA and two at the NBB. We selected the four Belgian banks that were involved in the stress test and contacted the Chief Risk Officer. Then, we used the snowball method to assess which other key people in the bank were involved in the stress test, and planned interviews with them. We selected our respondents at the ECB, EBA and NBB through desk research. We also attended two seminars on stress testing, 1 where we were able to speak to people (new as well as previously interviewed respondents 2 ) in a more informal setting. We continued interviewing until we had gathered sufficient information to formulate an answer to our research question.
We transcribed the interviews verbatim, and used ‘Nvivo’ software to code and analyse the data systematically. We took an interpretive approach to our analysis (Schwartz-Shea and Yanow, 2012; Yanow and Schwartz-Shea, 2006). This means that we focused on how respondents made sense of the exercise and the performance outcomes. Our research strategy also follows an abductive logic (Timmermans and Tavory, 2012). This means that we started from an empirical finding and then went back and forth between our research field and theory. Our results section reflects this abductive logic. Empirical findings and theoretical implications are not separated, but alternate in the text. We used a semi-structured topic guide during all interviews. After each round of coding, we revisited our topic guide and added new theoretical concepts to go back into the field. Our topic guide allowed us to discuss the same set of topics with all respondents and still act in a responsive way.
In order to stay close to the experience of the respondents, we first used emergent codes, close to the text, to code interview transcripts (Drisko and Maschi, 2015). Our goal was to study how the stress test affected performance outcomes. Subsequent rounds of coding therefore focused on identifying how various elements of the stress test affected organisational behaviour and how the stress test was experienced by the various respondents. The resulting coding process clearly showed links between the latent functions of the stress test and changes in banks’ behaviour. Based on this analysis, we gained an overview of how the stress test affected performance outcomes, which is presented in the following.
The latent functions of the EU-wide stress test
Our interviews gave us three key understandings of how performance indicators affect performance outcomes. We present our findings and theoretical interpretation simultaneously so as to allow the reader to follow our abductive analytical process. First, we show how a common dysfunction of performance indicators – inaccurate measurement – does not hamper performance outcomes. Second, we corroborate and complement the existing literature on the latent ritualistic functions of indicators, showing that these can importantly affect performance outcomes. Finally, we show how the process of calculating the performance indicator can have a larger impact on performance outcomes than (the use of) the performance information itself. In calculating the stress test, banks made internal changes that improved long-term performance outcomes.
The numbers are not right
A common dysfunction addressed earlier in this article is that, ultimately, performance indicators do not provide accurate performance information. Results are often said to be inaccurate, biased or gamed. In this section, we examine how the actors involved perceive and deal with this apparent dysfunction.
Banks have unique assets in their portfolio that justify a unique way to calculate the risk weight of those assets. However, when given too much freedom, banks would end up with different risk weights even for very similar assets, gaming the system to their advantage. As such, the stress-test methodology introduced caps and floors to somewhat level out the differences between banks’ internal models. However, this common methodology was said to stand in the way of accurately reflecting banks’ individual risk, raising questions about using the stress test to assess banks’ performance. When we mentioned the EBA’s common methodology to stress-testing teams, respondents sighed and started to shake their heads. A lot of bottled-up frustrations flowed freely, as a respondent noted: ‘What you see in the EBA stress test is that you are put in a corset in terms of methodology. This is necessary to be able to compare banks, but it does not make sense economically.’ Although the stress test makes a good effort at treating banks’ risks and assets equally through the common methodology, the exercise sometimes lumps very different things together at the cost of accuracy. This supports the critical voices. The stress test might not paint a very accurate picture of each bank’s actual performance, which might lead to unjust performance evaluation.
However, some nuance is required here. A goal of the European stress test was to do away with national bias and establish a European Level Playing Field (LPF). Although the stress test compromises on accuracy, without the LPF, the stress test would, most likely, not be taken seriously at all. As many respondents pointed out, the early (2009, 2010) stress tests – where the common methodology was only a few pages long – gave banks substantial discretion in their calculations, which was often used to game results to banks’ advantage. While the stress-testing exercise today is frustrating to banks (which are concerned above all with having an accurate result for their bank), all respondents agreed that, overall, the results paint a fairer picture of banks’ performance vis-a-vis each other. As such, we find that both regulators and regulatees agree that the stress test is dysfunctional, in the sense that it does not provide a completely accurate calculation of banks’ performance under risk, but it does minimise gaming, which leads to an overall better assessment of banks’ performance.
Rituals of verification
Although risk teams in banks were sympathetic towards the detailed rulebook and the LPF, they remained particularly frustrated about the granularity and intensity of the exercise. A respondent in a bank commented somewhat jokingly: Risks are very specific, your clients can be pharmacists, and pharmacists are not butchers, it’s a specific market, so the model needs to be specific. People with car loans in [one region], that’s different from loans in [another region]. And each model depends on the behaviour of your clients, so you need behavioural parameters. That’s what the internal models are for. And then what does the ECB do? They just add a buffer. But it’s the same everywhere I guess. Engineers do this too, they make complicated calculations about how much cement they need and it’s like 2.3658987 and eventually they’re told, let’s just take four. Everything is four. Always extra buffers.
Besides a message of rigour, the simplicity of the exercise also worked to its advantage. The stress test can basically be presented as a ranking of banks in a crisis situation, making it easy to explain and disseminate to a wider public. This raised awareness that European supervisors were ‘taking control’, they were measuring banks’ health and setting clear capital goals.
3
This added value of the stress test was picked up by respondents as well. During an interview, a respondent confessed: It’s a very visible exercise. It helps to explain to people what it is I do. They’ve heard about it, seen it in the news. It gets more attention from a wider public. This is not just in De Tijd [a financial newspaper], it’s even on Het Journaal [the daily evening news]. People say they’re frustrated with the results; they say they’re not credible because they’re low. Now there are two reasons for that. One is technical: with a bottom-up stress test, you’re going to get downwards-biased results. The other – and people underestimate this – is that you can’t see stress testing in isolation. If you look at what happened after the crisis in the US, the FED committed to endless liquidity, and the government committed to a floor under the economy within weeks. And then they said to the banks ‘we’re going to do a tough stress test’. And if you’re a bank asking for 20 billion dollars and they see you’re operating mostly in a country where there is a serious commitment not to let the economy slip away, and liquidity risk is non-existent, it’s an interesting business proposition. While in Europe, the ECB was cautious, for good reason, and governments retreated because they had maxed out on expenditures, and anyway the fiscal stance is more conservative, and then we stressed the banks. Imagine a bank going to the markets and asking 20 billion euro. So, it’s obvious that the results were mild. This is not because the people that do this are incompetent, or captured by banks, or intellectually weak or whatever. It would have been irresponsible to come out with a capital request of 100 billion in such a situation.
To be sure, just saying that banks are healthy in the stress test is not enough; it needs to be a credible statement. In the early 2009 exercises, banks scored well and faltered shortly after. In order to remain a credible exercise, regulators had to make sure that behind the scenes, banks were cleaning up shop. The stress-testing exercises conducted by the ECB and EBA contributed to this in a rather unexpected way. We elaborate on this in the next section.
More than just a ritual: governmentality
As mentioned, banks are required to fill out extensive templates with over 20,000 granular data points, over several risk categories. To do so, banks need to access granular data from all subsidiary branches in a short amount of time, be able to reconcile data from different risk departments and explain in excruciating detail how various macroeconomic variables will affect their assets. This did not merely serve the ritualistic or symbolic purposes stated earlier. Rather, it also, and more importantly, encouraged banks to improve their self-regulation. A respondent in a bank explained, for instance, how chief executive officers (CEOs) approved higher budgets for risk departments to improve their IT systems in order to successfully complete the stress test. These improvements in the IT systems are then used beyond the stress test to improve banks’ day-to-day risk management. Better IT systems help banks complete the stress test faster, but they also help banks detect problems and risks faster in their day-to-day business. Another improvement along these lines is that the stress test brought people together over different departments. A stress-test coordinator in a bank said: It’s a good experience to have, also for our internal stress tests. Because it’s so intensive, you really need to go over everything, line by line. And you’re also sitting at the table with so many people. That is also very important, this interaction between the different groups. Because when we do internal stress tests, it’s not as thorough, and we’re not sitting at the table with so many people. Here, it’s an important and rich exchange of thoughts and methods that is very valuable to think about stress testing in general.
Theoretically, we tie this to Foucault’s (2011) notion of governmentality, used to describe power that is exercised not by directly regulating behaviour, but by steering how individuals or organisations self-regulate. In this concept, Foucault brings together the notion of governing (governer) with modes of thought (mentalité). The government does not explicitly act upon an organisation; rather, the organisation acts upon itself. As such, the term is often described as the ‘conduct of conduct’, the state-steering of self-regulation (Lemke, 2011). This emphasis on self-regulation can be seen as characteristic of the transition from liberalism to neoliberalism; as Renou (2017) observes, the apparent withdrawal of the state actually marks a new kind of interventionism. Individuals and organisations are encouraged to take responsibility for themselves. Performance indicators typically act as tools of governmentality manifestly by using performance information to make certain outcomes desirable (as demonstrated in Renou’s work). However, we additionally find that performance indicators act as tools of governmentality by making certain practices and behaviours desirable and even necessary. While organisations can often readily game outcomes, gaming actual behaviour is much more of a challenge.
To be sure, the latent power exercised by the EBA and ECB is not against the interests of banks. Foucault is adamant that coercion is not necessarily bad. Moreover, this coercion does not mean that organisations are stripped of all their liberties and act as brainwashed ‘puppets’. On the contrary, power, as it is discussed by Foucault, can result in an ‘empowerment’ or ‘responsibilisation’ of subjects with agency capacities (Bevir, 2010; Lemke, 2011). This empowerment and ‘responsibilisation’ is noted in the stress test as well. Supervisors do not simply hand banks knowledge about what is healthy or risky. The EBA’s common methodology includes predefined categories, but banks still have the room to object or disagree (at least in theory). They are encouraged to think for themselves. Banks picked up on the learning experience that they had through the stress test.
5
As a risk director noted: We read papers and we follow workshops and we try to keep up, but it’s not always easy to find time. The stress test forced us to look at things that we had been neglecting. So, it’s a good learning experience. It’s a new kind of learning, not just from books, but learning as you go.
Conclusion: the latent functions that make performance indicators persist
Over the past decade, stress testing has become part and parcel of banking regulation in the EU. The EU-wide stress test calculates how banks would fare in a hypothetically plausible yet adverse stress scenario. Many comparable indicator regimes in national and international governance have been criticised for generating dysfunctional effects. On a theoretical level, we therefore asked under which conditions performance indicators, such as the stress test, can improve performance outcomes, despite their proven weaknesses and dysfunctions. The manifest objective of performance indicators is to measure performance so as to use this information for learning, steering and control, or accountability (Van Dooren et al., 2015). Besides these manifest functions and the often-accompanying dysfunctions, we argue that indicators fulfil important latent functions as well, which have been largely overlooked so far. Based on interviews with stress-testing teams in Belgian banks and the NBB, and consultants and officials at the EBA and ECB, this article took a closer look at the latent functions of performance indicators, and how they can contribute to performance outcomes.
First of all, we found that what is commonly seen as a dysfunction of a performance indicator need not negatively affect performance outcomes. Like many other indicators, stress tests seem to face validity issues. These seem to challenge the manifest goals of performance indicators, that is, accurately measuring performance in order to steer performance outcomes (Van Dooren et al., 2015). However, dysfunctions such as inaccurate measurement can serve an important role in furthering the overall objective of improving performance outcomes. Compromising on accuracy proved to be a necessary part of a trade-off to ensure a level playing field, which was key to the overall credibility and legitimacy of the stress test, allowing it to improve performance outcomes.
Second, our fieldwork corroborates and complements earlier findings in the literature that indicators fulfil important ritualistic functions (Boswell, 2008, 2015; Power, 1997). They can signal that governments are dealing with a problem thoroughly, and that accountability mechanisms are in place, instilling trust. In this article, we showed that these latent functions can also be key to actually improving performance outcomes, a quality that is often overlooked. By stating that an organisation is performing well, and that regulators are on top of the situation, organisations are given the necessary room to actually work on improving performance outcomes. This mechanism is known as performativity (MacKenzie, 2006). As such, these ritualistic functions of performance indicators should not be brushed off as a pleasant side-effect of performance indicators; rather, they should be more widely recognised as key factors in allowing organisations to improve performance outcomes.
Finally, we showed that the process of calculating a performance indicator can latently improve performance outcomes in itself. Calculating the stress test inadvertently caused banks to professionalise their risk departments, which improved banks’ internal risk management, in its turn, improving long-term performance outcomes. We explain this mechanism by drawing on Foucault’s theory of governmentality (Foucault et al., 1991): regulators improved performance outcomes by operating on a latent level, by educating and configuring habits, aspirations and beliefs. In the process of calculating the performance indicator, banks updated IT systems, increased communication across departments and improved internal processes. Although they initially only made these changes as part of the process of calculating the performance indicator, they ended up actually improving their internal risk management, and thus their performance outcomes. Banks’ performance outcomes are not improved because they learned from the performance information itself; on the contrary, they actually find the performance information to be invalid. Rather, they improved their performance outcomes by revising internal management systems to be able to calculate the performance information. This mechanism has been largely overlooked in the literature so far, warranting more scholarly attention. The process of how the performance indicator is calculated might be a key factor in explaining why performance indicators succeed in some cases and fail in others.
To conclude, the manifest goal of performance indicators is to produce performance information that can be used to improve performance outcomes by allowing organisations to learn from or reflect on this information, or by rewarding and penalising over- and underperformers. However, an increasing number of critical voices show that, in many cases, performance indicators fail to improve performance outcomes because of dysfunctions such as gaming or manipulation for political power plays (Davis et al., 2012; Dunleavy et al., 2005; Hood and Peters, 2004; Van Dooren et al., 2015; Van Thiel and Leeuw, 2002). The main contribution of this research is that it shows new, latent, ways in which performance indicators affect performance outcomes. We show how the process of calculating performance indicators, which is often overlooked in the literature, can in itself improve performance outcomes – regardless of the results of the indicator, and their use or validity. Where a key criticism of performance indicators is that performance information is often inaccurate, biased or invalid, we thus rebut with the afterthought that using numbers that do not count can latently help to improve performance outcomes.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship and/or publication of this article: This research was funded by the Research Foundation-Flanders (FWO) through a PhD fellowship grant to the corresponding author, grant number 11X9818N.
