Abstract
The systematic review undertaken for the Group Analytic Community by Sheffield University is an excellent piece of work ‘of its time’, but it may not speak to everybody in the field. Some of the reasons for this are the precise methodology of such reviews, concerned with an exclusively rationalist model for selecting and appraising evidence in a framework, which excludes other schools of thought (particularly the social sciences, including critical theory, anthropology and economics). As many of those who work in group analysis have backgrounds in these and allied disciplines, rather than biomedical science, there is a risk of excluding much useful and scholarly collaboration with adjacent disciplines unless we hold an open mind about such methodologies. As group analysts, we are in a strong position to observe and criticize the ‘evidence-based hegemony’ when it becomes closer to dogma than science.
This article is based on a talk given to a Group Analytic Society conference entitled ‘Can Group Therapy Survive NICE?’ held in London on 29 January 2010.
Introduction
My own background is first as a GP then a psychiatrist by profession, for which I also trained as a group analyst. I qualified about 15 years ago, and now work part time at the Department of Health with the National Personality Disorder Programme as their clinical advisor. My clinical work is in Berkshire, where I started as consultant psychiatrist to a small non-residential therapeutic community (Winterbourne) in 1994. I have been involved in the Department of Health’s Personality Disorder work since 2002, as a member of one of the expert advisory groups which were responsible for the policy implementation guide ‘Personality Disorder: No Longer a Diagnosis of Exclusion’, for which I specifically set up a process to hear the views of service users. I was also a member of NICE’s Guideline Development Group for their guidelines on Borderline Personality Disorder, which was published in January 2009.
I been asked to talk about the systematic review that the team from Sheffield University have written and produced. From first impressions it has the elegance and clarity of a beautifully produced document. It has been meticulously planned, it is precise, it is orderly, and I think, it is stunningly systematic. So, I think in a way it is what we might think of as a ‘state of the art’ systematic review.
So how can one disagree with such an object of perfection? Well that is exactly what I am here to talk about. I am not going to criticize the report, but concentrate my comments on the scope of the review and the context within which it was required to be commissioned.
The headline of what I am going to be saying is that, although it is very good in so many ways, it does not quite ‘reach all the parts’ that I would want to go as a group analyst. Having discussed these matters with Nick Benefield, I think it is safe to say that some of these are also the places we would want to go with personality disorder and the work of the national team: it is doubtless necessary at the moment, but it is not sufficient. This may, in due course, generalize to other areas of mental health in public policy. ‘New Horizons’ was recently launched, and public mental health may see changes we cannot currently foresee; it is being put on a new track, which might or might not endure a new government in a few months.
Comments on the Review
I will start at the end of the review, with the recommendations. Here I have boiled them down to a single word or short phrase each, and categorized them as zero, one or two (see table 1). Zero indicates that I disagree or am unsure; one that I quite agree with it; two that I agree so much I am going to spend a bit more time talking about it later.
Broad responses to the recommendations
So, first to comment on each of the eight recommendations:
‘Aptitude treatment interaction’ is in an unarguably rich area for research: ‘who does best in what sort of treatments?’ It is the question classically asked by Roth and Fonagy in What Works for Whom? A Critical Review of Psychotherapy Research (2004). Certainly in our therapy groups I expect we would all like to think the way we select and prepare people for groups ensures that those who are in them are the people who get most out of the therapy. But I am sure that it is realistic to think a more rigorous and systematic approach could improve this.
‘Economics’, by which I mean heath economic research. This will be covered in more detail later, as the global tightening of belts and demand for transparent justification of every penny of public money spent makes it crucial. A useful buzz word, or phrase, I have recently heard is ‘Social Return on Investment’ (SROI): the very language captures how economic imperatives may be judged in a wider context than the simple ‘value for money’ that has been a recent imperative.
‘Further Research’ —in the spirit of motherhood and apple pie, this is something that is difficult to disagree with. However, the framework of that research should be critically examined—and have a wider scope than the ‘RCT or nothing’ nihilism that is currently so prevalent. I suspect we have friends in positions of high authority to support this—but I will expand on this later.
‘Good practice guidelines’ —at its most basic, this simply means writing down what good practice involves. In the therapeutic community (TC) field, we have done this exercise in considerable detail, and have established core standards (the essentials of being a TC), commissioning standards (to meet the requirements of those who purchase services) and a value base (10 statements which lie behind the treatment philosophy, which can be independently evaluated). We now have a network of nearly one hundred TCs where communities undertake an annual visit to each other and discuss the standards, and how they are or are not meeting them. For those who want to go further than that, there is a detailed and high level scrutiny process with which they can also be awarded accreditation (by the Royal College of Psychiatrists). A process like this could well be an option for the group therapy, or at least practices undertaking group therapy, although it is harder to see how it would work with individual group therapists working independently.
‘Homogeneity’ and ‘Identification of groups to study’ —I read this recommendation as being that groups should be for similar types of problems in order to be subject to ‘product evaluation’. However, I do have serious reservations about this call for uniformity, as I think one of the main features of groups we use for therapy is they need to have a variety of different people, and attitudes, and thoughts, and behaviours, and cultures, and probably many other things beside, in order for the group mix to ‘do its work’. It is more than simply exposing individual members to different ways of seeing the world and so imposing some leverage on the members who do not share some common facet of the group’s humanity. This is about the old Foulkesian maxim of the group constituting the norm from which the individuals deviate. But in short—I am not happy about the idea of increasing the level of homogeneity in groups in order to be able to research them.
‘Service user involvement’. This is an area I will cover in more detail later. If we believe it to be an offshoot of ‘consumerisation’ of public services, I believe we have missed a major opportunity. It shares with consumer movements the intention to get mass support behind its ideas, but as psychotherapists we are in a powerful position to cut through the common adversarial argy-bargy, and start to develop a more collaborative way of working. I think it would be true to say ‘you ain’t seen nothing yet!’ when it comes to service user involvement.
‘Systematization’. This does worry me, as the unacceptable face of modernity. One of the phrases we bandy about in the Department of Health, and something we rail against, is what we call ‘the industrialization of therapy’. In this, all therapeutic effort and the systems surrounding it become a mechanical process that does not take account of human agency, the complexities of real life, let alone the very thought of a dynamic unconscious—or irrationality as a rational response to emotional circumstances. If we go down this road, I fear we are set on a path to becoming obedient Daleks in perfectly defined hierarchies where imagination, spontaneity, playfulness and normal human ‘fun’ are all programmed out of us.
Reflections
We live in a time where the concept of ‘evidence based practice’ has taken on a life of its own. Some might say that the idea has been fetishized, and is now used as a talisman for a relentless form of modernization that ‘takes no prisoners’ and threatens much orderly and well-established practice, which happens to lack the formal ‘evidence qualifications’. To illustrate this I have five examples.
The first is ‘football score mentality’ —where the ‘results’ of randomized control trials are bandied about by the cognoscenti as if they are football results: a competition between different types of therapy to see which one is ‘better’. For example I heard a conversation recently that went something like ‘CBT 4, MBT 2; MBT 6, DBT 4’ —although those are made up numbers it is this mentality whereby therapies are scoring points in league tables that many important policy makers and commissioners with resources take all too seriously. The fundamental point, which surely underlies this, is that the ‘Dodo Bird Verdict’ (where all therapies have some benefit, and what they share is more interesting than how they differ) is more important. What therapies share, which is probably something about the nature of the relationships and unconscious processes, seems to be the key point. Maybe it is pre-verbal, pre-oedipal and beyond simple rational analysis—but this is not a discourse currently acceptable to the mainstream.
The second pointer about our current evidence-based world that I would like to criticize is the tendency to reify many complex matters into three letter acronyms—what I call the ‘alphabetti spaghetti therapies’. We currently have IPT, CAT, MBT, CBT, DBT and doubtless numerous others in vogue. So, every therapy has to be branded, come up with a snappy three letter acronym so it can play a part in the jostle between different therapies to be the brand leader. Is this the way we want to go in psychotherapy? It has reduced therapies to mere techniques and manuals—and does no credit to the importance and under-considered common factors, and how best to use them. It also reduces the development of therapies to a marketing and branding exercise, where the ones that will come to the fore are those that have a better commercial production department, than a meaningful and rigorous therapeutic core.
A separate point is that where we are now with the ‘evidence based everything’ movement is that this is one moment in history. Although things are like this at the moment, with a ferocious enthusiasm for a certain type of evidence, they were not like this 20 years ago, and they may well not be like this in 20 years time. It is clear where the need for evidence has come from—particularly about the need to justify all we spend with full transparency and accountability in the face of global competition for resources—but I think a realization will soon dawn that this is necessary but not sufficient to define what humans need for good mental health.
Another point in criticism of the current fads is what I call the ‘syllogism error’. This is where ‘there is no evidence that this therapy is effective’ is taken by commissioners and other to mean ‘there is evidence that this therapy is not effective’. This is logically not true, although the pressure of competition between therapies—and presumably pressure of time on those who do commissioning work in this area—makes these conclusions easily arrived at. But this simple logical fallacy could so easily sign the death warrant to any particular therapy, not through its lack of effectiveness—but through its lack of acceptable evidence. So, good therapies—or promising developments of existing therapies—could easily be lost to future development because of this error of logic.
The final point I want to make about the general context of evidence at the moment is its hierarchical nature. In group analysis we think of things as a matrix of interconnected relationships that interact in complex ways, at several different levels. This is how the world works in general—there is a matrix of influence rather than a single and linear hierarchy of order and control. Strong negative feelings may well be evoked by imposing a hierarchical mechanism on professionals who feel they should be part of a network of influence.
This all leads to what I call the ‘tyranny of evidence’. Perhaps the simplest illustration of this is NICE itself: as power is so heavily concentrated in its processes and its decisions. What it writes in the guidelines is so strongly seen as ‘the last word’ that other influences are overshadowed. Not only is it power but also I think glory—in some ways the English system of NICE is the pride of the world. The British Medical Journal often talks about it in these terms and I was on a train coming to the talk today opposite an East Asian couple with a sheaf of NICE papers. I was trying to read upside down to see what it was about—and soon saw that they were part of the fact-finding mission for the Chinese delegation examining the English NICE system. So, it seems that NICE is a mechanism and technology that is taking over the world in many ways—and the English are leading it. Although that should be something to be proud of, I also wonder if we should fear it as we might Pandora’s box. To coarsely oversimplify the issue, it may bring a linear and hierarchical way of thinking that is antithetical to what we understand to be core therapeutic values.
The Hierarchy of Evidence
So, thinking of hierarchies, let us look at the ‘evidence hierarchy’. Table 2 shows a version which is commonly used.
Evidence Hierarchy
This so clearly gives the message that ‘type I evidence is better than type II’, that ‘type II evidence is better than type III’ and so on that why would any clinical researcher seeking to do his or her best ever aim at less than type I?
But maybe the matter is not so simple or linear. Consider this quote:
Evidence hierarchies attempt to replace judgment with an over simplistic pseudo-quantitative, assessment of the available evidence. Decision makers have to incorporate judgments, as part of their appraisal of the evidence in reaching their conclusions. (Rawlings, 2008)
This is not the critique of a disaffected researcher who has received a lower than hoped-for rating in the five yearly research assessment exercise for university funding. It is the august Harveyian oration, given in 2008 to the Royal College of Physicians. Its title was De Testimonio. On the evidence for decisions about the use of therapeutic interventions (2008). The distinguished speaker was Professor Sir Michael David Rawlings—who has been the chairman of NICE since its inception in 1999. The monograph is freely available on the worldwide web and is easily understandable, coming across as a balanced and non-dogmatic view of how evidence is built.
The position taken by Sir Michael shows that those in the highest positions in the social hierarchy of the ‘evidence-based world’, see the nature of evidence as a more subtle matter than a simple hierarchy where one type of research ‘trumps’ another. From my own experience on the NICE guideline development group for borderline PD, I can see how this wide scope is reflected in the full guideline, for which all types of evidence are considered: but it does not carry through to ‘the headlines’, or indeed the recommendations, which form the much more widely available guideline summary, which is the only document the vast majority of relevant people will look at.
This leaves one wondering where the expectation of balance and open-mindedness, as espoused by Sir Michael, has been lost on the way to the commonly perceived wisdom that ‘only randomized trials are good enough’, as epitomized by the NICE guideline summaries. A phrase which is rather out of vogue, but which seems to fit, is ‘dumbing down’.
Health Economics and Social Return on Investment
In this it is important to consider that a full economic appraisal of any publicly funded activity should not just consider health, but other systems of public policy as well. Wider ‘social returns’ are measured and add to the benefit of treatments than just improved health outcomes. This fits well with the more holistic and public health-related focus of the new mental health policy ‘New Horizons’ 1 .
Let us consider it as a thought experiment, using intensive group therapy, perhaps a therapeutic community, as the intervention, and the commonly used measure of ‘QALYs’ as the evaluation tool. Figure 1 shows a graph as we might expect of a person with severe borderline personality disorder. The top of the Y axis (100) represents a ‘perfect’ quality of life, and the lowest point (zero) represents death. The normal way to measure this is with a simple instrument like the five question ‘EQ-5D’, which is standardized by ‘citizens’ juries’ which place a value between zero and 100 on all the different permutations of the five ‘axes’. But for the purposes of this thought experiment we are making an intuitive estimate of where in the range between dead and having a fully engaged, fully creative and fulfilled life, people might be in the lifelong course of untreated and treated borderline personality disorder. Consider the x axis as age going through adulthood from perhaps 20 to 80 years old.

Lifelong Course of Borderline Personality Disorder
In figure 1, I have charted what might be a typical life course of someone with a moderate to severe borderline personality disorder. There are spells of slightly better quality of life than average and there is a general trend to increase over the lifespan. The graph prematurely hits zero for about 10% of the population, who commit suicide—but this would be a different example and calculation. The relapsing and remitting pattern is what might be clinically expected, and there is some research evidence to support this general pattern. This is demonstrated in the line (A) in the figure.
If, however, this trajectory is interrupted by a spell of therapy—this could be group therapy, or a therapeutic community programme, or a discrete episode of individual therapy (marked (C) on the diagram), the trajectory is changed. A common long-term clinical observation is that it would look something like line (B). This shows a generally better ‘quality of life’, still with vicissitudes, but a sharp dip for a matter of weeks or months soon after therapy finishes. For this period, an individual may have a worse quality of life than if they had never had the therapy. This is represented on the diagram where line (B) spends a certain time below line (A). We have had a small number of people returning to our therapeutic community after, say, 30 years since their first therapy, for a ‘top up’ —but the norm is that people get on with their lives afterwards and are maybe not ‘cured’ in a medical sense—but seek help in much more rational ways and generally get more satisfaction from their lives. Their ‘longitudinal quality of life’ is perceptibly better—for them and those around them.
There are clearly many caveats and cautions about estimating these two trajectories, but this is a ‘thought experiment’, or an initial model, which could be refined and developed through suitable research. The required research would be very different to an RCT—exploratory nature, cohort design, longitudinal and long-term, and with the intention of iteratively refining the model. In effect, this means trying to get ever closer approximations to representing ‘what really happens’ with lines (A) and (B). In contrast, an RCT would be a snapshot of where a statistically suitable number of people are (in their two randomly assigned groups) at point t0 (baseline) then t1 or t2, or indeed t3 (outcome). In a study design, these would be relative to time of treatment, not to age as the graph might suggest. It would simply be measuring the single dimension of the distance between the two lines—with treatment or without treatment or between two different treatments.
Making a longitudinal model adds another dimension—and what was just the measurement of a line, becomes the measurement of an area. The area between the two lines represents the increase in quality of life someone has. The unit for this measurement is the ‘quality adjusted life year’ (QALYs) and is the benefit of a treatment over somebody’s lifespan for a certain economic cost: one year extra life at 100% quality would be 1 QALY; two years of improving it from 25% to 75% would also be 1 QALY; as would 10 years of increasing somebody’s life quality from 60% to 70% (which is perhaps a more realistic expectation of psychotherapy). When NICE make difficult decisions about expensive pharmaceutical treatments or surgery, they generally allow NHS costs per QALY to go up to £20,000, and sometimes to £30,000. If this thought experiment of borderline personality disorder is at all accurate, the number of QALYs gained through such a treatment (represented by areas D plus E plus F minus G on the graph) is 11.75. If the full economic cost of a full treatment (excluding any cost-offset benefit) is £100,000 (and most current treatment programmes in non-residential therapeutic communities would be well below this), the cost of a QALY would be £100,000 ÷ 11.75 = £8,500.
This should surely be evidence enough to convince policy makers that the treatment is worth doing—indeed, it is a treatment they cannot afford not to offer when it saves so much money in the long run? (which is the equally salient ‘cost-offset’ question). However, the current funding crises mean that commissioners and managers can only concentrate on their budget for this year, and in their own areas of direct control—but, hopefully, more sophisticated accounting and whole-system based balance of benefit analyses will become accessible for general use in such mental health settings in the next few years.
Service User Politics
One problem with outcome measures in all forms of psychotherapy is the question of whether they are measuring what matters. What matters to a researcher, in order to show a good effect size perhaps, may have little meaning to a psychotherapist or a service user. This was illustrated rather vividly at a recent conference where there was a mixed audience of clinicians, researchers and service users. A small huddle of academics was making its way up the stairs to a meeting. I was in a different huddle of people including some service users who asked ‘who are that lot?’, and somebody else commented that they were the ‘old boys outcomes club’. This is an attitude that does prevail in service user circles—that the ‘old boys outcomes club’ measure what researchers are interested in measuring, rather than what matters to service users.
An example of a more user-friendly measure, which is mentioned in the DH New Horizons mental health strategy, is the ‘Mental Health Recovery Star’. This is a set of ten diagrams covering different aspects of a person’s life which service users rate for themselves—there are careful descriptions of each level of each area of difficulty. It has a very simple and graphical representation, which allows anybody using it to see their progress, and the areas on which they need to work. It is much more orientated towards the current concepts of ‘recovery’ than to symptomatology or diagnosis, as are more familiar to psychologists and psychiatrists.
Deeper than this though, there is a need for a different nature of relationship between service users and professionals. Many service users are disillusioned, angry and feel that they have not been properly heard, that they have been disrespected and that their needs are not being met by a hierarchical and ‘expert-driven’ system of care. Some clinicians are seen as remote and inaccessible, or even arrogant and pompous. This difference in the nature of relationship is one that has long been central in therapeutic community practice, in the approach of having a ‘flattened hierarchy’. However, it is also very relevant for group analysts because of the non-authoritarian nature of analytic groups, and particularly the way in which therapists move between different positions in relation to the group—the old teaching of ‘therapy in the group, of the group, and by the group’.
This should place us, as group analysts, in an excellent position to be able to work collaboratively with service users. However, while I do sense that mental health’s ‘service user movement’ is gathering momentum the balance of power still has a considerable way to move before more equality and partnership can be enjoyed in most clinical relationships. In the words of the rock band Bachman–Turner Overdrive, on their 1974 single, ‘You ain’t seen nothing yet’.
In an organization such as the IGA, involving service users in strategy as well as practical matters is vital. This might include teaching and communicating to outsiders about how groups work, and it could be a very effective method of demonstrating its relevance and power. I understand that the intention for it to start is with the follow-up conference to this one, later in the year.
Whole System Thinking
One of the things we notice in the world of personality disorders is that considering people as individuals is insufficient to explain much of what happens. We need to consider the individual in their relationships, particularly in their family, in how they relate to services, in how they relate within their neighbourhood and locality, possibly their profession and many other groups of which they are a part, and the mesh of interconnected relationships that they have. This is a variant of the figure and ground concept, which all group analysts will know from their training and is central in the practice of group analysis.
Our dominant hypothetico-deductive model depends on measurement, prediction and control, and this is based on a world where deviations from the norm are distributed in a Gaussian way. However, many would argue, and I number myself among them, that the way human and biological systems work in the world is not of this nature: things happen in a random way that is governed by chaos theory and Mandelbrotian distributions. In this way sudden disjunctions can come about, extraordinary events can happen and it is not possible to predict anything beyond the immediate future. This system is much closer to meteorology and weather than to a simple mechanical system.
It has been described in economics as ‘the Black Swan effect’ and it provides a vivid illustration of how the collective human actions in bringing about the banking crisis of 2008 were not predictable or possible to conceive without invoking chaos theory. Human development could be construed in a similar way: as a chaotic process with so many variables and so much complexity of interaction that the ultimate destination of somebody’s development, that is to say their personality, is not predictable, although it is understandable.
In closing I would want to make two extra recommendations to the reports and would suggest the IGA considers them.
The first would be to make a research strategy that, for the current moment in time, focuses on experimental design and measured outcomes. This is much as recommended by this report, but a new research strategy also needs to go wider, and give importance to historical, sociological, philosophical, anthropological, critical theory, art and uncertainty theory perspectives. This is likely to embrace the interests of many group analysts who are not primarily psychological and statistical in their orientation, and will also contribute to a deeper base for group analytic thinking.
The second recommendation would be to involve every group analyst in the ‘research effort’. This should not frighten everybody, but be an extension of people’s natural curiosity that is a necessary part of being a therapist (and perhaps a human being). It would involve them in their own area of expertise, and the seeds for this could be planted in their training. Keeping aware of and interested in research would be expected once in regular practice—through systems of continuing professional development.
