Abstract
Geraghty’s recent editorial on the PACE trial for chronic fatigue syndrome has stimulated a lively discussion. Here, I consider whether the published claims are justified by the data. I also discuss wider issues concerning trial procedures, researcher allegiance and participant reporting bias. Cognitive behavioural therapy and graded exercise therapy had modest, time-limited effects on self-report measures, but little effect on more objective measures such as fitness and employment status. Given that the trial was non-blinded, and the favoured treatments were promoted to participants as ‘highly effective’, these effects may reflect participant response bias. In non-blinded trials, the issue of reporting biases deserves greater attention in future.
I read with interest Geraghty’s (2016) editorial on the PACE trial for chronic fatigue syndorme (White et al., 2011) and also the PACE investigators’ letter of response (White et al., 2017). Here, I will comment briefly on several issues raised in the response letter, all of which highlight important themes for future research into chronic fatigue syndrome (CFS) and/or clinical trial methodology more generally.
To summarise, the PACE trial investigated the effectiveness of three interventions for CFS: (1) graded exercise therapy (GET), which focused on gradually increasing patients’ activity levels; (2) cognitive behavioural therapy (CBT), which addressed what were seen as patients’ ‘unhelpful cognitions’ about their illness and their fears about exercise; and (3) a novel treatment, Adaptive Pacing therapy, which encouraged patients to restrict their activity levels (White et al., 2011). Participants were randomly assigned to one of these three treatments, or to a control, no therapy condition. Each participant also received at least three specialist medical care consultations. One year after treatment allocation, the primary outcome measures – self-rated fatigue and physical function – showed improvement in all groups, but significantly more so for the CBT and GET groups. Using a definition based largely on these primary measures, the investigators concluded that 22 per cent of patients in the CBT and the GET groups had ‘recovered’ following treatment, but only 7–8 per cent in the other two groups (White et al., 2013).
At the outset, some clarification is needed on the current status of the PACE debate. The PACE researchers state they have, ‘… repeatedly addressed the criticisms made in the editorial of the methods and analyses used in the PACE trial’ (White et al., 2017: 3). They then refer the reader to various letters to editors, commentaries and a website that provides answers to selected questions. However, in these documents, readers can neither raise new questions nor query any of the answers given. In short, they present only one side of the debate. The time has come for these researchers to engage in more direct academic dialogue with critics.
Scientific procedures and standards
Geraghty’s editorial charged the PACE investigators with failing to adhere to ‘accepted scientific procedures and standards’. In defending themselves against this charge, the investigators note their adherence to the CONSORT guidelines for conducting randomised trials (Schulz et al., 2010). They also note the various ethics reviews and procedures the project underwent, the high status of the journals they published in and the sheer number of papers that were published.
Some authors have raised concerns regarding PACE’s ethics procedures, in particular about whether the participants were fully informed of the investigators’ financial interests (see, for example, Tuller, 2015). However, here, I focus on one departure from protocol that may have significantly influenced conclusions about treatment effectiveness.
The authors emphasise that they published a trial protocol prior to data analysis (see White et al., 2007). A published protocol is desirable because it ensures that researchers do not alter their dependent measures after they have seen the data, in ways that might unduly favour the study hypotheses. However, to be of benefit in this way, the protocol must be followed. The investigators made several major ad hoc changes to their dependent measures – for example, they considerably altered the definition of ‘recovery’ (White et al., 2013). The CONSORT guidelines specify that authors should identify and explain any subsequent changes to outcome measures (Schulz et al., 2010). However, in a recent paper, we examined the explanations provided for these changes and found them to be either insufficient or based on inappropriate extrapolation from normative data (Wilshire et al., 2016). We also found that the changes operated to favour the study hypotheses: they served to increase the apparent rates of recovery by a factor of three and to yield a significant treatment effect where there would otherwise have been none. This is a significant cause for concern.
A more accurate statement is that many relevant procedures and standards were adhered to, but there were some significant departures, sufficient to undermine several key conclusions of the study.
Researchers’ therapeutic allegiance
Another charge raised by Geraghty was that the researchers had significant professional and personal investment in two of the treatments: CBT and GET. Geraghty is referring to the ‘researcher allegiance effect’: the finding that, in studies examining more than one treatment approach, the treatment(s) favoured by the researchers tend to outperform other treatments (Luborsky et al., 1999, 2002; Munder et al., 2012; see also Wilson et al., 2012 for a discussion of this issue in relation to the Triple R parenting programme). Several factors may contribute to this effect, but one is likely to be the manner in which the non-favoured, ‘comparison’ treatment is conceptualised and implemented. Often, when a treatment is used as a comparison condition, it is implemented in a weaker form than when it is used clinically (Cuijpers et al., 2012; Munder et al., 2011). In the context of the PACE trial, the comparison treatment, Adaptive Pacing Therapy, was a novel intervention designed especially for the trial. None of the primary investigators believed in its effectiveness, and none had expertise in its delivery, so its failure to yield successful outcomes is not particularly surprising.
The best way to address the researcher allegiance effect is to include primary investigators in the research team that specialise in the comparison treatment approach and to charge these persons with the design and supervision of those treatment sessions. However, there are also other much simpler steps that researchers can take. The first is to recognise the problem. The PACE researchers’ defence against Geraghty’s claim indicates that they believe themselves to be entirely impartial, which, of course, cannot be the case.
The next step is to ensure that therapists present all treatments to participants as equally likely to lead to improvement. This is especially important when the primary outcomes are self-report measures, since these measures can be strongly influenced by patients’ expectations (Hróbjartsson et al., 2014). Unfortunately, in PACE, CBT and GET were promoted to patients during therapy as highly effective. For example, CBT participants were told that CBT was ‘a powerful and safe treatment which has been shown to be effective in … CFS/ME’ and that ‘many people have successfully overcome CFS/ME using cognitive behaviour therapy, and have maintained and consolidated their improvement once treatment has ended’ (Burgess and Chalder, 2004: 123). GET participants were told that ‘in previous research studies, most people with CFS/ME felt either “much better” or “very much better” with GET’ and that GET was ‘one of the most effective therapy strategies currently known’ (Bavinton et al., 2004: 28). No such information was given to the remaining two groups. In any rigorous trial, researchers need to at least acknowledge this potential confound and consider its possible impact on results.
Finally, one simple additional measure that can be taken is to share data as widely possible, so that researchers with different perspectives can examine it. I therefore urge the PACE researchers to share their (appropriately anonymised) data willingly and as widely as possible. This is the best way to demonstrate that they are aware of the issue of investigator bias and are willing to take steps to address it.
Our recent reanalysis of the PACE trial data on rates of recovery demonstrates just how powerful a data sharing approach can be. We were able to demonstrate that apparently minor, late changes to the definition of recovery impacted very substantially on the observed rates of recovery and on their final conclusions about the effectiveness of the different treatments (Wilshire et al., 2016). When we defined recovery according to the original protocol, we found that recovery rates were consistently low and not reliably different across treatment groups. The investigators appear to have been entirely unaware of the extent of this problem.
Specific problems and limitations
A researcher’s enthusiasm for a particular treatment can also lead them to overinterpret their findings or overlook limitations. One limitation of the PACE trial – which has been pointed out by critics, but never fully acknowledged by the investigators – is that the treatment effects were almost entirely limited to self-report measures. Most of the objectively measurable outcomes did not yield significant treatment effects, for example, fitness and employment status did not differ across treatment groups when measured an entire year after trial commencement, and although mean walking distances were higher after treatment with GET than after medical care only, this difference was small (approximately 30 m, less than 10% of the baseline walking distance 1 ), and no such benefit was observed for the CBT group.
Again, the problem here is that, in a non-blinded study, self-report measures are highly vulnerable to response bias. The size of this bias is not trivial. A recent meta-analysis of clinical trials for a range of disorders calculated that when participants were non-blinded to treatment allocation, self-reported improvements associated with treatment were inflated by an average of 0.56 standard deviations relative to comparable blinded trials. Importantly, no such inflation was observed when the outcomes involved objectively measurable indices (Hróbjartsson et al., 2014). Therefore, in order to securely demonstrate the efficacy of any intervention within a non-blinded design, researchers need to show that self-reported improvements are supported by evidence based on more objectively measurable outcomes.
It would be unreasonable to expect the PACE investigators to solve the problem of participant response bias single-handedly. But in a trial of this size and importance, we can reasonably expect them to take some simple measures, such as balancing the information they provide to different participant groups about effectiveness. We can also reasonably expect them to minimise – or at the very least acknowledge – potential sources of bias. And we can reasonably expect researchers to acknowledge and discuss potential red flags, such as a lack of agreement between self-reported improvements and objectively measurable outcomes. Instead, the PACE investigators did not even report the most worrying results until several years after publication of the main findings, and when they did, they dismissed them as unimportant (see especially, Chalder et al., 2015; McCrone et al., 2012).
The critique by Geraghty also raises some concerns with the long-term follow-up assessment, which was undertaken at least 15 months after completion of the treatments (i.e. at least 2 years after trial commencement; Sharpe et al., 2015). About three quarters of participants completed this assessment, and at this stage, the differences between the treatment groups on self-report measures were no longer statistically reliable. The PACE investigators do not consider this to be a matter of concern, because many patients received supplementary CBT or GET after the main trial had been completed, and therefore the randomisation had not been maintained. They reasoned that, since the mean ratings in the CBT and GET groups did not significantly drop over the follow-up period, at least the treatment benefits were ‘maintained’. Such a within-group comparison is of course meaningless, given that around one-quarter of participants were lost to follow-up, and these losses are unlikely to be random.
Recently, my colleagues and I calculated the long-term follow-up results for the sizeable number of patients who did not receive a substantial dose of CBT or GET after the trial. In this subsample, there were no significant differences between treatment groups at follow-up even on self-report measures (mean self-rated physical function scores for the CBT and GET groups were 64.2 and 62.5, and for medical care only, 62.6; for self-rated fatigue, the figures were 17.9 and 18.7, respectively, vs 18.7 for medical care only). This finding is not trivial. If patients who have received CBT and GET are indistinguishable from other patients when tested at least 15 months after treatment, the practical value of these treatments is limited. But more importantly, this finding raises further suspicions about the mechanisms underlying the self-reported effects obtained at the primary, 52-week endpoint. The limited duration of these self-report effects is fully consistent with an explanation in terms of participant reporting bias.
One might argue that the standards I describe here represent an ideal scenario and that many, if not most, behavioural studies fall short of them. However, few of those studies wield the power of the PACE trial when it comes to influencing policy and perceptions about this disabling illness. A false-positive conclusion in this context could significantly impact not only on patients’ current treatment options but also on future research that could potentially yield better treatments. In sum, extra vigilance is required in this situation.
Conclusions from PACE and directions for future research
The PACE investigators conclude their recent published defence by expressing their hope that research of this nature will continue in the future, in spite of the criticisms. I am puzzled by this statement. The £5 million PACE trial, which assessed more than 600 participants, was designed to provide ‘definitive’ evidence of the effectiveness of CBT and GET for CFS (Walwyn et al., 2013: 2). Findings from the trial showed that CBT and GET – as delivered here – can have modest effects on patients’ self-reports of fatigue and/or physical functioning, at least if those reports are elicited within several months of trial conclusion. The size of these effects did not exceed what might be expected from reporting bias alone, and they were no longer evident at long-term follow-up. These treatments do not improve more objectively measurable aspects of functioning, such as physical fitness or employment status. Finally, there was no evidence from the trial that patients can recover from CFS as a result of either of these treatments.
In my view, the implications of these findings are quite clear: there is no need to pursue the question further. CBT and GET are simply not effective enough as treatments for CFS (if at all). We need to do better. It is time to begin the search for entirely new treatments.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship and/or publication of this article.
