Abstract

Strengthening the ‘practice’ dimension of what we publish is an ongoing preoccupation of the journal. This issue includes a ‘visit to the world of practice’ that describes how evaluation is understood and applied by one of the major institutions that shape evaluation practice in Europe, the European Parliament. This is one of a number of articles on institutional practice in the journal pipeline. Unsurprisingly, evaluation practitioners often start their consideration of evaluation practice from the practitioner perspective – how we evaluators practice or should practice our craft. It is salutary to be reminded of the influence of institutional practices whether directly through commissioning evaluations or through the many other ways institutions shape the wider evaluation ecosystem. A second innovation in this issue that takes us a little closer to evaluation practice is a new ‘Review’ feature. This is a ‘roundup review’ of 2023 evaluation blogs, podcasts and webinars. We hope this will become a regular (six-monthly) feature. The aim here is to heighten awareness among our readers of the many informative and sometimes heated online debates that often nowadays gives an early warning of the way the evaluation field is evolving. The bulk of this issue continues to showcase innovative thinking in evaluation theory and methodology as well as practice.
Evaluation capacity building (ECB) has been the subject of much investigation, mainly at a single organisational level. As Charlotte Laubek and Isabelle Bourgeois observe, ‘little is known, currently, about how EC [Evaluation Capacity] might apply in interorganizational initiatives’. On the other hand, many policy challenges require inter-organisational collaboration. Indeed, the authors argue that contemporary challenges such as climate change, social inequality and pandemics make this more necessary than ever. This exploratory study considers ‘how building capacity to do and use evaluation across organizations and sectors can support interorganizational initiatives and lead to potential interorganizational learning . . .’. Laubek and Bourgeois build on and extend existing EC frameworks especially those developed by Brad Cousins and colleagues. These emphasise individual capacities and competencies including enabling skills such as stakeholder involvement; organisational capacities including organisational resources, processes and policies; and capacities to use and to learn from evaluations. The authors apply their framework to four case-studies drawn from central government, healthcare and community organisation in Denmark and Canada. They underline the importance of inter-organisational evaluation capacities to enable stakeholders to learn from each other and work together. The role of stakeholders is becoming more central in many aspects of evaluation, as the shift from a governing to a governance model – long-established for example in innovation studies and political science – becomes more pervasive also in the field of evaluation.
Marijn Faling, Greetje Schouten and Sietze Vellema draw on their experience providing ‘strategic support through action research’ when setting up a Monitoring and Evaluation (M&E) function. The setting is a major Dutch development initiative (2SCALE) ‘that focuses on incubating inclusive agribusiness fostering food and nutrition security’. This large-scale ‘partnership’ initiative is described by Faling and colleagues as doubly complex: both complex in its many programmes across sub-Saharan Africa, which are themselves situated in many different, complex and uncertain environments. According to the authors, such doubly complex settings are especially ‘prone to paradoxes’ – the presence and persistence of ‘contradictory logics’ and ‘competing demands’. Paradoxes have been widely observed in management studies – which the authors draw on – as well as in other fields such as organisational development and psychotherapy. Reflective practices such as action research are well suited to surface paradoxical or contradictory dynamics. The authors identify five main paradoxes and their contradictory logics. These concern the ‘purpose’, ‘position’, ‘permeability’, ‘method’ and ‘acceptance’ of M&E systems in the 2SCALE programme setting. For Faling, Schouten and Vellema, paradoxes have a positive value – to be balanced and accommodated rather than rationalised. (This is consistent with Greco and Berti (2023) who warn that organisational paradoxes can only be regarded as ‘generative’ – that is, positive – if organisations develop ‘adequate response capacities’.) Identifying and working with paradoxes was critical to the M&E ‘design choices’ that the authors describe. They also argue more ambitiously that thinking in terms of paradoxes ‘is essential for M&E to be successful’.
An afterthought about language. Faling, Schouten and Vellema frame their article in terms of ‘action research’ and ‘complexity’. They do not use terms such as ‘capacity’ or ‘capacity development’ even though the discussion of the design, management and support for a new M&E function speaks directly to capacity-related ideas. The earlier article by Laubek and Bourgeois was more explicitly bound by the literature, language and discourse of evaluation, drawing on a considerable body of work on evaluation capacity and capacity development. As evaluation is an interdisciplinary field, it is reasonable to expect that different researchers and practitioners will draw on different concepts, methods together with associated language. So does it matter? Perhaps. For example to remain well-informed, we all have to engage in a constant process of translation in order to draw together strands of knowledge and experience that use different disciplinary language and grammar. The increasing use of systematic review methods by some evaluation researchers risks ignoring whole swathes of work that doesn’t match pre-selected search terms, which in turn risks undermining the usefulness of such reviews. The black-box not being opened here is whether evaluation will ever be able to don the clothes of a ‘discipline’ without restricting the possibilities of learning however messily from across other disciplinary boundaries. These dilemmas will not go away. Indeed, the next article describes itself as a ‘formative evaluation’. Faling and colleagues could have described their work as formative but coming from an action research tradition did not, which does not make their article any less interesting! A topic deserving further discussion.
Gabriel Sidman and Carlo Carugi discuss the use of geospatial analysis in a formative evaluation for the Global Environment Fund (GEF). This was part of a ‘comprehensive evaluation’ preparing for GEF’s ‘replenishment’, in anticipation of the Fund’s next phase of activity. The authors note that formative evaluation ahead of results is necessarily limited, for example, to assessing programme design and targeting of programmes. The appropriateness of incorporating geospatial data depends on the extent to which the programmes being introduced or planned include a spatial dimension. This was so for this programme where existing geospatial datasets were used ‘to assess [. . .] thematic and geographic relevance of project site selection in GEF food systems integrated programs’. Geospatial factors also coincide with other ‘priority factors’ able to support a ‘multi-criteria’ spatial analysis. Inevitably other ‘non-spatial factors’ were also relevant, which is why, the authors suggest, ‘it is important to triangulate findings from spatial analysis with other qualitative and quantitative methods’. Generating ‘indicators relevant to program goals’ was facilitated by the availability of relevant datasets (e.g. on environmental degradation). Sidman and Carugi argue that incorporating geospatial data into such a formative evaluation adds a quantitative dimension that reduces reliance on subjective evidence, and insofar as data sources are ‘produced by outside parties’ for different purposes, can ‘also add transparency to an evaluation’.
Meenakshi Fernandes, Katharina Eisele and Irmgard Anglmayer overview the role of evaluation in the European Parliament (EP). Evaluation in support of policy-making is usually associated with ex-post evaluations initiated by the ‘executive branch’ of government. In the legislative context of the EP, the focus is understandably on the process of lawmaking. However, a distinctive feature of EU lawmaking is that it is a joint prerogative of the EP, the European Council (that represents EU member states) and the European Commission. The authors describe how these joint responsibilities are laid out in the 2016 ‘Inter-institutional Agreement (IIA) on Better Law-Making’. They also discuss the way institutional and inter-institutional relations in the EU that shape the lawmaking process have evolved and continue to do so. The authors describe the EP’s various evaluation products at ‘agenda-setting’, ‘preparatory’, ‘legislative’ and ‘post-legislative’ stages. Anyone who has struggled to understand the differences between reports such as ‘Cost of Non-Europe’ and ‘European Added Value’, or between ‘implementation appraisals’ and ‘European Implementation Assessments’ will find this article invaluable. It is also interesting to compare the key role of evaluation in the EP compared with other Parliaments worldwide. In these terms, this article contributes also to ongoing debates about what can be done to counter-balance the risks for evaluation of ‘administrative capture’. Finally, it is noteworthy that Fernandes, Eisele and Anglmayer are all part of the European Parliamentary Research Service (EPRS), which carries out some but not all evaluations initiated by the EP. How different institutions and agencies balance internal, outsourced, advisory and peer-review roles is also a topic of interest. Such configurations inevitably shape evaluative practice more generally and also feed into debates about what kinds of evaluation ‘independence’ are appropriate in different institutional settings.
Astrid Brousselle, Megan Curren, Bronwyn Dunbar, James McDavid and Rik Logtenberg are committed to evaluation practices that ‘can contribute to the design and implementation of policies and programs that, in turn, contribute to a better future’. This commitment widely shared in the evaluation community stems from an awareness of the urgency and scale of the ‘polycrises’ we face in terms of climate change, biodiversity, public health and demographics, to name but a few. The authors argue that not only do these crises interact but have led to commitment by government (and other governance actors) to new programmes on biodiversity and climate change that also promote positive impacts on human systems – and in the Canadian context among others, reconciliation with indigenous people. All of which highlights for Brousselle and colleagues, the limitations of existing assessment approaches. The crises we face create ‘a need for systematic consideration of a range of policy and program impacts that extend beyond the program’s core objectives’. According to the authors, the Planetary Health Framework was identified by Canadian local government officials they were working with as a possible basis for such an assessment tool. Although various health and environmental assessment tools exist, the authors argue that there is a ‘need for simpler, more accessible, and engaging [assessment] processes’. On this basis two ‘rapid assessment tools’ were developed and piloted, ‘co-created’ by researchers and local government officials.
One may remain agnostic about using any one set of assessment tools in the many, rapidly changing settings in which contemporary crises are manifest across the world. However, an important message that comes through from this Canadian case is that it is no longer sufficient to simply argue that ‘evaluators must do more’. What Brousselle, Curren, Dunbar, McDavid and Logtenberg demonstrate is that this commitment has to be translated into tools and methods that can be adapted and made operational across programmes and contexts by officials and citizens as well as by practicing evaluators.
Marko Nousiainen and Lars Leemann tackle a familiar but thorny problem: how to evaluate the effects of large-scale programmes intended to have ‘effects’ on individual’s lives and experiences. The setting is European Social Fund programmes in Finland that aim to promote social inclusion for the most disadvantaged citizens. Programmes that aim to have more direct effects on individuals and communities in terms of experience and personal identity are always challenging to evaluate. At one level, this is a straightforward mixed method realist evaluation that examines CMO configurations, that is, mechanisms and outcomes in their context. The challenge for many such evaluations is where these ‘theories’ come from. Nousiainen and Lars Leemann, who emphasise the experience of social exclusion, build on ‘capability’ and ‘self-efficacy’ theory. Successful outcomes here are in terms of human agency: experiencing greater sense of control, of ‘being an active agent in one’s own life and in society’. These theories were ‘tested’ in four exemplary cases that collected participants own narratives in the form of ‘small success stories’ and then ‘used to depict the CMO-configurations of social inclusion interventions’. The study also conducted a small-scale survey that included an existing scale for measuring social inclusion to validate the results of the qualitative case-based analysis. The survey was applied at two moments, at ‘baseline’ and at ‘follow-up’, in order to identify changes that could be attributed to programme inputs. As is often the case, quantitative survey results were ‘not as positive as the qualitative findings suggested’. However, subgroup analysis proved more illuminating.
And so to the ‘roundup review’ of blogs, podcasts and webinars anticipated and introduced at the beginning of this Editorial. When exploring with Tom Aston, what form this new feature in the journal might take, the challenge of translating a dynamic and rapidly changing online world into print was clear. Even making it possible for hyperlinks to be available – at least for those with online access – was not unproblematic. As Editor of a peer-reviewed journal, I was aware that the transient nature of online material was at the other end of the spectrum in terms of the cautious, anonymised peer-review process that distils regular articles. Although this is not a peer-reviewed feature, it is curated by Tom Aston – but ultimately it is up to readers to reach their own conclusions about the usefulness of the signposted material. We would be pleased to hear suggestions of current and forthcoming online resources for inclusion in future reviews. Consistent with the often personalised and expressive character of the blogosphere, it is difficult to imagine a review such as this as being anything other than selective and maybe sometimes even idiosyncratic. Although as Aston notes in his abstract, he reviewed ‘hundreds of blogs and dozens of podcasts and webinars’ the feature is explicitly described as a ‘personal’ review. Indeed, it includes blogs by Aston himself which themselves refer to or comment on other online resources. On the other hand, it seems inevitable that a ‘roundup review’ such as this would be produced by someone who is themselves an active and engaged digital native! This feature will undoubtedly continue to evolve but already this first review is likely to stimulate many ideas and responses.
