Abstract

Evaluation is often described these days as ‘theory driven’ or at least ‘theory informed’. That is certainly instantiated in the first four articles in this issue. Very different theories but very potently deployed. The next four articles follow a more analytical and methodological tradition.
Øyunn Høydal asks us to reconsider the often-assumed divide between on one hand knowledge producers such as evaluators and researchers, and on the other, those who use this knowledge in the policy community. Instead of the ‘two communities’ assumption, Høydal argues that ‘policy is part of knowledge production in the same way as knowledge production is part of policy’. This ‘coproduction’ framing builds on Peter Dahler-Larsen’s understanding of influence at different ‘moments’ in an evaluation process – from initiation through to reporting of results. Norwegian civil servants in six public agencies were surveyed and interviews were also conducted. The author’s findings raise interesting questions about knowledge brokering. This is often thought of as an independent, external and mediating role. Øyunn Høydal highlights how this role is also fulfilled by civil servants who mediate between political demands for answers and validation, and what can legitimately be said on the basis of evaluation studies.
As so often in evaluation, it is the taken for granted and quotidian that challenges us most. Ghislain Arbour takes on one such: the notion of ‘evaluation frameworks’. The author sets out to bring order into the diverse and often confusing picture of how frameworks are understood. Rather than to aim for a common understanding, Arbour instead provides a model that ‘acknowledges a variety of types of framework and offers a conceptual rationale for their inclusion’. According to Arbour, this ‘allows frameworks to be differentiated from non-frameworks, allows for identification of frameworks [ . . . ] and allows users to efficiently harness the potential of different framework types’. Arbour’s model encompasses theories, concepts and intellectual insights; a normative or value element; the ‘object of evaluation’ which can encompass both ‘interventions’ and those who evaluate or manage evaluation process; as well as institutional including regulatory and accountability aspects. The authors’ discussion of how this model can be applied to a host of different frameworks is thought provoking: as a good model should be!
Jone Martínez-Palacios and Igor Ahedo apply ‘critical deliberative theory’ to the evolving notions of gender, encompassing gender identities and self-expression, as well as diverse sexual orientations. For the authors, this is set against a ‘common goal’ of democratisation. Evaluation from a ‘Gender+ Perspective’ sets out to ‘compare the learnings of the critique of deliberative democracy and the feminist view of evaluation’. This builds substantially on Maria Bustelo’s conceptualisation of Evaluation from a Gender+ Perspective (EG+P) ‘as substantially political and incorporating an inclusive perspective that includes social justice’. The authors also deploy contemporary reformulations of deliberative theory that takes on board notions of ‘exclusion’ and inequality more than ‘early deliberative theory’. Martínez-Palacios and Ahedo emphasise the coherence of the EG+P and these more recent formulations of deliberative theory. This coherence is demonstrated by applying both sets of ideas in a case study of the evaluation of a law on equality between men and women in the Basque Country.
Sophia Rodriguez and Jeremy Acree theorise the practice of evaluation as ‘truth-telling’. The article follows from a specific empirical evaluation case which one of the authors led. That evaluation was distinctively focussed on ‘normative perspectives’ around a sensitive group – transnational migrant youth, the authors’ aims are broader. They see lessons for evaluation practice in this case as having more general relevance. The authors draw on Foucault’s theories of ‘biopolitical governmentality’ to illuminate evaluation’s ‘truth-telling’ role. This is an explicit and reflexive riposte to the more common ‘techno-rational’ or ‘instrumentalist’ response of evaluations in such settings. Although arguing from a particular empirical starting point, the article remains substantially theory focussed: theorising about the norms and practices of evaluation, and in its frequent reliance on technologies of quantification and measurement. In contrast, evaluation as ‘truth-telling’ has to challenge universalising concepts. There have been many articles advocating the need for new expressions of ethical evaluation behaviour in this journal. It is encouraging to see an attempt to ‘destabilise evaluation practice’ that centres on a specific evaluation in a way that demonstrates the consequence of different ethical and normative frameworks. Agree or not, the discussion of ‘paradoxes’ in evaluation certainly has resonance well beyond the particular case of newly arrived migrant youth in North America.
When public funds are dispensed, taxpayers and public agencies habitually ask questions about ‘benefits’, ‘impacts’ and ‘outcomes’. This is problematic in many spheres of public life, nowhere more so than in the ‘public cultural sector’. Kim Dunphy, John Smithies, Surajen Uppal, Holly Schauble and Amy Stevenson are optimistic about the possibility of measuring the outcomes of ‘cultural engagement’. This contrasts with a widespread belief in the intangibility of the effects of most cultural production, consumption and participation. The authors develop a ‘schema’ of what they argue to be the main ‘outcomes’ of cultural engagement encompassing knowledge, aesthetics, a sense of belonging and understanding of others. This schema builds on existing policy and research literatures and was ‘field tested’ and refined with over 350 stakeholders from Australasia, Asia and Europe. From the standpoint of evaluators who in one way or another are always interested in ‘valuing’, culture and cultural engagement could be seen as an ethical, theoretical and methodological test-bed for the limits of our craft. How indeed do we measure the public good especially when, as is often the case, the ‘good’ is on the verge of being commodified?
The realisation that most evaluation findings are probabilistic is a problem for many users and commissioners of evaluation. For evaluators also, how firm can our judgements be; if we depend on theories, can our theories be ‘demonstrated’? Having introduced a ‘diagnostic’ perspective in a previous issue, Barbara Befani turns her diagnostic eye to Bayesian Updating (BU), an increasingly favoured approach to judging probabilities. She argues that BU is well suited to demonstrate success beyond vague generalisations and avoid risks of confirmation bias by evaluators committed to their favoured theories. This is a demanding and technical article that many readers of this journal will shy away from. However, Befani offers fascinating insights into what frontier analytical methods are now attempting to deliver, and how much this differs from traditional ‘frequentist’ understandings of probability. Discussions of ‘how much evidence is enough?’ and how to work with ‘packages’ of evidence are challenges that will resonate with many evaluators of whatever persuasion.
Eran Raveh, Yuval Ofek, Ron Bekkerman and Hertzel Cohen are concerned with how evaluators can use large amounts of ‘unstructured’ electronic data. They argue that often evaluators and other researchers still approach big data using traditional methods such as content analysis. Instead the authors advocate state-of-the-art automated tools such as data-mining, text analytics, machine learning and data visualisation while still recognising the centrality of human judgement and interpretation for decision making. They suggest not only that newer data analytic tools are intuitive, efficient and user friendly, but that they allow for causal and trend analysis at increased levels of granularity. Raveh and colleagues ambitiously ‘propose an end-to-end methodology for analyzing vast amounts of text . . .’. The proposed approach is applied in this article to performance data but according to the authors, ‘the proposed system can be implemented successfully on any corpus of documents . . .’.
Considering the potentially disturbing possibilities of the new of approaches outlined by Raveh et al. should not diminish the importance of well-established approaches that fulfil a purpose. This in essence is the argument of Roger Slade, Peter Hazell, Frank Place and Mitch Renkow. Their focus is on ‘good practice’ when evaluating the impact of policy research in development settings. The authors review a substantial body of existing research albeit confined to a small number of international bodies and in particular the CGIAR (Consultative Group on International Agricultural Research). While recognising the complexity of policy making and the importance of many different qualitative as well as quantitative methods, Slade and colleagues are themselves focused on quantifiable socio-economic impacts, within a counterfactual logic and further elaborated with the help of cost-benefit analysis or other efficiency measures. Following their review of existing practice, the authors systematise this kind of evaluation in terms of six main steps. This starts with a ‘Theory of Change’ understood as a more elaborate form of logical-framework. The authors acknowledge both the strengths and limitations of these kinds of evaluations. However, they hold to the view that for an ‘assessments of the efficiency or economic value’ of the kind of policy research they are interested in, quantitative economic assessments continue to be central. As a subtext in this article, Slade et al. describe how ‘story telling’ and ‘outcome stories’ are beginning to take hold in this evaluation domain. The evolution and differentiation of evaluation paradigms in international development agencies such as the CGIAR is probably just beginning and worth keeping an eye on in coming years.
