Abstract

This, the first issue of 2017, features themes and debates that are by now familiar to readers of this journal. There are two articles on the ways institutional settings affect evaluation practice; two methodological explorations: a visit to the frontier worlds of Process Tracing and Bayesian Updating; a meta-evaluation of a consultancy’s naturalistic/constructivist evaluation studies; and an equally contrasting pair of articles on evaluation capacity building, one with the Irish school system and the other with a pan-African initiative to build (and evaluate) research capacity.
It is often difficult to separate out evaluation as a practice from the institutional arrangements that surround it. This is a strong theme in the first two articles in this issue. In the first Kristin Reichborn-Kjennerud and Signy Irene Vabo explore the extent to which national audit bodies (so-called Supreme Audit Institutions – SIAs) precipitate organizational learning in the agencies and departments audited. The authors argue that whilst performance audit resembles evaluation, in the SIA case there is an added dimension because these audits are backed by the possibility of parliamentary sanction. Existing research on the topic is suggestive but inconclusive. A feature of this article is that the empirical investigation – surveys and interviews of Norwegian auditee respondents – has an explicitly ‘theoretical point of departure’. In addition to research on organizational change, the authors also draw on how we understand institutional relations, deliberation and theories of action. It is unsurprising that the effects of audit pressure on organizations are not straightforward. As the authors recognize, there is always the possibility of ‘symbolic’ compliance. Indeed, the extent to which performance audits are seen to be of value and useful by auditees partly ‘depended on perception of the extent to which auditees’ comments were taken into account’
Stijn van Voorst considers the evaluation of European legislation, a relatively recent expectation of the European Commission in addition to the more well-established evaluations of expenditure programmes. There has been considerable variation in the extent to which this has been taken forward and in the quality of evaluations produced by different Directorates-General (DGs) of the European Commission. Van Voorst focuses attention on differences in capacity to evaluate legislation within DGs as a possible explanation of variation in performance of legislative evaluations. The author adapts a model of organizational capacity to evaluate developed in Danish local government; and uses fuzzy-set QCA (fsQCA) to analyse interview and documentary data drawn from 17 DGs responsible for major acts of legislation. This use of fsQCA is justified both because of the method’s ability to work with small samples and with ‘combinations of causal conditions’. The conclusions that are derived from this analysis are modest: those DGs with a long tradition in evaluating expenditure programmes and with higher budgets ‘tend to attach more importance to legislative evaluation and invest more means in legislative evaluations’. The author identifies the kinds of future analysis needed to further unpick variance across DGs. As with Reichborn-Kjennerud and Vabo, Van Voorst underpins analysis with a well-worked theoretical framework which strengthens analysis as well as interpretation of results, especially if causal inference is to be drawn.
Barbara Befani and Gavin Stedman-Bryce open up fascinating new horizons for the world of ‘impact evaluation’. Whilst impact evaluation has continued to be a prominent feature of the evaluation landscape, we are today as likely to be concerned with a programme’s ‘contribution’ as with traditional concerns for ‘attribution’. After all, few programme managers let alone evaluators now believe that the intervention of interest is solely responsible for outcomes and impacts. Yet as Befani and Stedman-Bryce point out, how to bridge the gap between data and ‘contribution claims’ remains unclear. This article ingeniously weaves together two recent imports into evaluation methodology – Process Tracing and Bayesian Updating – to suggest ‘how to collect data and assess the strength of such data towards (or against) a contribution claim’. They exemplify their argument with the evaluation of a health advocacy campaign in Ghana. The authors argue that their approach to ‘quali-quantitative’ methods complements theory-based evaluation approaches such as Contribution Analysis, system-based evaluations and Realist Evaluation. Their hybrid methodology, which they label ‘Contribution Tracing’, focuses on the probative value of causal claims, i.e. ‘the power of specific items of evidence to increase or decrease our confidence in a specific claim’. Thus the authors complement the four tests of Process Tracing with systematic estimates of probability based on Bayesian Confidence Updating. For the practising evaluator this both adds the promise of rigour to the way qualitative evidence is used and as importantly the ‘evaluator is thus forced to be transparent about their assumptions and confidence on the existence of the claim, and to “declare” its observable implications’. Even more ambitiously the authors sketch out the possibility of a ‘trial’ process through which stakeholder juries can make their own judgments of the probative value of contribution claims. As is often the case with innovative articles, one is left with the impression that this journey is only just beginning!
Shivaun O’Brien, Gerry McNamara, Joe O’Hara and Martin Brown take us back both to this issue’s institutional theme and to matters of evaluation capacity – albeit in a very different institutional setting, that of ‘post-primary’ schools in Ireland. As the authors note ‘self-evaluation has become a key quality assurance mechanism for schools internationally’ as are various forms of capacity building for schools undertaking such self-evaluations. Various studies have identified the main kinds of support that are now widely in use ranging from manuals to training; and from comparable indicators to web-based tools and online forums. O’Brien et al. focus on an innovative action-research approach both to implementing and evaluating capacity building in Irish schools. Self-evaluations in Ireland consider teaching and learning broadly but with an emphasis on numeracy and literacy with ‘improvement targets’ linked to the OECD’s PISA scores. The model of school support that the authors discuss centres around a ‘critical facilitator’ role – a variant on the better-known concept of ‘critical friend’. This role was operationalized along structured lines with three meetings with each school team; tasks to be completed by team-members including data collection. Participating schools were generally positive about the critical facilitator role and the researchers ‘identified the key characteristics of the critical facilitator role, highlighting the importance of their external status, their credibility, and their use of pressure and support in working with SSE teams’.
The next article revisits the capacity building theme – in this case research capacity – and does what we always encouraged the evaluation community to do: reflect on the experience of undertaking evaluations. This is what in the theatre would be called an ensemble piece as is made clear by the extensive authorship from Kenya, Malawi, Côte d’Ivoire, Tanzania and the UK. Sonja Marjanovic, Gavin Cochrane, Enora Robin, Nelson Sewankambo, Alex Ezeh, Moffat Nyirenda, Bassirou Bonfoh, Mark Rweyemamu and Joanna Chataway reflect on ‘the experience of conducting a real-time, theory-driven evaluation of a complex health research capacity-building intervention’. This initiative funded by the Welcome Trust ‘aimed to build sustainable research capacity in Africa at institutional and network levels through African ownership and control of capacity-building efforts’. The team followed a ‘theory-driven, real-time approach’ and the article describes how high-level frameworks were prepared, starting points assessed, interim reports reviewed and emerging lessons disseminated. The discussion of the team’s experience working as ‘independent evaluators’ but in an engaged way in University settings, raises interesting parallels with the experience of O’Brien, McNamara, O’Hara and Brown in Irish schools – what should consortium members do when they are asked for specific advice about underperforming students? Of course there are inevitably many parallels when evaluation roles are expanded and tested in the field. But what is striking here are the parallels in two evaluations of such different ambition and scale!
Tracey Phillips and Jacques de Wet are interested in the quality of evaluation results when working in a ‘naturalistic/constructivist paradigm’. There are points of connection here with Befani and Stedman-Bryce’s preoccupations. They are both concerned with making the most of qualitative data but from a somewhat different philosophical and methodological perspectives. Phillips and de Wet have been ‘developing a framework for assessing rigour in naturalistic research’ in the development sector. The framework builds on Guba and Lincoln and related traditions which assert that it is indeed possible to be rigorous when following naturalistic modes of enquiry. Indeed the authors follow Lincoln and Guba’s 1985 ‘trustworthiness’ criteria of Credibility, Transferability, Dependability/Auditability and Confirmability, consistent with the ontological and epistemological assumptions of naturalistic and constructivist practice. The core of this article is the application of the framework through a meta-evaluation of five studies conducted by a firm of consultants which followed a naturalistic approach in their own work. Data was gathered through interviews conducted with lead evaluators and with informants from client organizations. It is unusual for consultants to open themselves up to this kind of scrutiny and creditable that they did so. However, given the publication of many evaluations these days, the scope for learning from such meta-evaluations is considerable. Evaluators and the commissioners of evaluation would have much to learn from the process, independent of the evaluation paradigm concerned.
Happy New Year!
