Abstract

In these days of electronic alerts and keyword searching, it is easy to stick to our predefined paths: reading only what we already know will be of interest! Yet it is often articles read by accident, at the edge of our peripheral vision that ignites the spark of innovation. Of course I recognise that this reflects the inevitable bias of a journal Editor. However, I would like to think that this first issue of 2019 may convert others to this view. We cover evaluation in public health, conflict resolution and education; and international organisations, national evaluation systems and regulatory evaluations. Yet there are bridges to be built across different domains and evaluation approaches – for example, the uses of theory on which all articles depend to a degree; uses of new and established data sources; and how to move up and down the micro, meso and macro scales when evaluating both large-scale and small-scale systems. I was struck by how Mike Coldwell’s article focused on the relatively micro level of school leadership; his discussion of ‘underlying’ dimensions of context raises similar issues from a different perspective as do Graham Moore and colleagues when subjecting public health evaluation to a ‘systems lens’.
There have been repeated calls for evaluators to take advantage of the opportunities presented by ‘big data’. This interest in using new data sources is already being translated into evaluation practice. Javier Fabra-Mata and Jesper Mygind have evaluated Norway’s role in the peace process between the government of Colombia and the FARC guerrilla movement; or more specifically ‘… Norway’s engagement in, and contribution to, the process as one of its facilitators’. An important part of this evaluation relied on Twitter analysis. The authors discuss the rationale for choosing this particular social media platform for part of the evaluation – a key decision that would be faced by any evaluator using social media. One of the criticisms of some approaches to big data in the wider research world is a claim that the availability of such data obviates the need for theoretical frameworks. It is therefore noteworthy in this case that an analytical framework was developed based on what is known from research about ‘trust dynamics’ and mechanisms ‘derived from the literature on trust, negotiations and conflict resolution’. The description of ‘data extraction’ from available Twitter feeds outlines the kinds of decisions that have to be made when collecting, cleaning and analysing data. The authors conclude with their reflections on the kinds of ‘lessons of broader applicability’ that evaluators can extract from this experience that follow from these kinds of social media analyses.
Since the beginning of the 20th century, public health in the United Kingdom has provided one of the most productive settings for the exploration and further development of complexity and systems thinking in evaluation. The current state of the art in the domain is exemplified in this issue by Graham Moore, Rhiannon Evans, Jemma Hawkins, Hannah Littlecott, G.J. Melendez-Torres, Chris Bonell and Simon Murphy, themselves among the leading protagonists in the complexity/public health/evaluation debate in the United Kingdom. From 2000 onwards, the UK’s Medical Research Council (MRC) has issued ‘guidance’ to those wishing to evaluate complex interventions, inevitably a challenge given traditional medical reliance on randomised control trials. Subsequent revisions of MRC guidance – in 2008 and on process evaluation of complex initiatives in 2014 – began to shift from linear formulations of complex interventions to include understandings of the complex systems – and contexts – within which these interventions are situated. In this article, the authors argue for ways in which MRC guidance needs to be further revised to accommodate complexity through a ‘systems lens’. They do this within a conceptualisation of interventions as ‘events within systems’ and follow through the implications of this thinking for the kinds of evidence, theory and evaluations needed; and the importance from a systems perspective of interdisciplinary cooperation and knowledge co-production with stakeholders.
Many individual evaluations do not stand-alone, rather they are integrated into evaluation systems that may themselves be national and sometimes also sectoral, regional or institutional. Žilvinas Martinaitis, Aleksandr Christenko and Lina Kraučiūnienė consider the way evaluations in Lithuania are (or are not) used in relation to an understanding of the Lithuanian evaluation system. In doing this, they take as their starting point two authors who have presented interesting theorisations of evaluation systems in an institutional setting previously published in Evaluation – Højlund (20.1 and 20.4); and Raimondo (24.1). Martinaitis and colleagues propose an idealised typology of evaluation systems drawing on the knowledge management literature. They are interested in ‘why most evaluations in Lithuania are geared to improve implementation and management systems’. Theorisation of evaluation systems would suggest that when initiated externally (in the case of Lithuania by EU accountability demands), evaluations would most likely be used to ‘legitimise interventions rather than improve their implementation’ – which is not the case in Lithuania. The authors are also interested in why evaluations do not feed into better polices, which again might be expected on the basis of past research. Martinaitis, Christenko and Kraučiūnienė analyse the evaluation of EU Structural and Investment Funds in Lithuania in line with their typology of evaluation systems and knowledge management. An interesting exercise in using an idealised framework as a heuristic is to interrogate theorised consequences of different types of evaluation systems.
Steffen Eckhard and Vytautas Jankauskas are also interested in evaluation systems and their study of stakeholder influence across 24 UN organisations builds on some of the same ideas about evaluation institutionalisation as the preceding article. The authors acknowledge that ‘despite its ambitious intention to be a scientific and data-based inquiry, evaluation is an inherently political activity’. Hence, they are interested in how stakeholders engage in ‘interest contestation’ and have the potential to exercise influence over evaluation processes and results. From this starting point, Eckhard and Jankauskas develop a ‘resource-based’ taxonomy of political power that attributes different kinds of resources to different stakeholders. The two main stakeholder groups considered are governments, who might be expected to govern in international organisations (IOs); and international public administrations that run these agencies on a day-to-day basis. Although there is considerable variety across UN organisations, it should come as no surprise that international administrations appear to exercise the greatest evaluation influence in just over half of the UN organisations considered. What is more thought-provoking is the suggestion that following a ‘principal–agent’ logic, governments use evaluation in order to bolster their control, while administrations ‘seek to deliberately escape member states’ control to increase autonomy’. An interesting supplementary explanation surely, that goes beyond a belief in the merits of ‘independence’, is to explain the growth of ‘independent evaluation offices’ in IOs reporting directly to Boards made up of government representatives.
As one of many cognate evaluative approaches that intersect with the world of programme and policy evaluation, Regulatory Impact Assessment (RIA) – assessing whether regulations are likely to achieve their intended results – will be familiar to many evaluators. Deborah Shmueli, Michal Ben Gal, Ehud Segal, Amnon Reichman and Eran Feitelson look beyond the normal scope, scale and complexity of particular RIAs. In order to ‘evaluate the regulatory framework for earthquake preparedness in Israel’ the authors developed a more inclusive Regulatory System Scan and Assessment (RSSA) methodology. Although building on more conventional RIAs, the authors argue that a more systemic assessment approach is needed to take account of ‘scores of organizations, laws, regulations, and policies relating to many areas of professional expertise’ that are linked to earthquake preparedness. Furthermore, in order to ensure that the RSSA strikes a balance between so many different interests, the methodology also involves stakeholders of implicated organisations. Shmueli and colleagues argue that regulatory systems are ‘sub-optimal’ if there are either tensions, designated as ‘disruptive friction’ at the public interface with responsible bodies, or ‘gaps’ not addressed by present regulatory systems. The RSSA is therefore designed to identify sources of friction and gaps, prioritise their importance and incorporate ‘self-reflective’ elements by including ‘more general procedural values, such as public engagement and attention to cultural diversity’. The authors see prospects for applying the RSSA to other regulatory systems, although this requires further research.
Mike Coldwell as mentioned earlier focuses his attention at a different scale than many of the articles in this issue: his interest is in context and theory-based evaluations applied to professional development in schools. However, the author’s interest in ‘context’ and theory are not far removed from the ‘systems’ concerns of Moore and colleagues, nor indeed the interest in institutional contexts in articles about evaluation’s institutionalisation in Lithuania and the UN. The bridge here is Coldwell’s awareness that ‘observed contextual factors’ can lead to an over-simplistic understanding – in particular with the more linear depictions of theory in evaluation as logic or path models. Coldwell suggests that ‘observed contextual factors have underlying features, which can help explain how they act to influence the programme or initiative..’; and ‘that context works to aid explanation, rather than acts as something to be controlled for.’ He identifies a set of six ‘underlying factors’ that he argues are common to context in many fields of evaluation. For example, he posits that context needs to be seen as dynamic rather than as a static input factor, that programme actors have agency and can influence as well as be influenced by programmes, and that contexts have a particular relationship with interventions and that these relationships also help explain programme effects. Coldwell usefully highlights some of the common characteristics (and occasional sources of confusion) between core ideas in different schools of evaluation – citing among others, the work of Stuffelbeam, Greene and Pawson. Over recent weeks, I have participated in more than one discussion in which ‘context’ elides into ‘mechanism’; where colleagues have perhaps too easily migrated between notions of context, institution and structure, and where the decision to define the unit of analysis – what to include or consign to the ‘context box’ – raises the same questions of systems boundaries at a micro level that have been signposted in this issue when evaluating at meso and micro levels.
