Abstract
The Diagnostic and Statistical Manual of Mental Disorders (DSM) has been criticized because evidence suggests that it lacks appropriate validity, reliability, and clinical utility, and the Hierarchical Taxonomy of Psychopathology (HiTOP) has been offered as a solution to these criticisms. Our goal in the present study was to compare clinician perceptions of these systems for the conceptualization of outpatient clinical cases. A sample of actively practicing clinicians (N = 143) rated three clinical vignettes using both diagnostic systems and then rated the two approaches on seven indices of clinical utility. HiTOP was favored for overall clinical utility score as well as utility for formulating effective interventions, communicating clinical information to the client, comprehensively describing psychopathology, describing global functioning, and ease of application. There was no preference between HiTOP and the DSM for communicating with other mental health providers. The DSM was not favored for any clinical utility outcome.
Keywords
The standard for diagnosing mental disorders in the United States is and has long been the Diagnostic and Statistical Manual of Mental Disorders (DSM, Fifth Edition; American Psychiatric Association [APA], 2013). The influence of the DSM spans the globe and cuts across various professional disciplines; it is referred to by researchers, clinicians from multiple orientations, policymakers, criminal courts, and third-party reimbursement entities (APA Publishing, 2017; Kawa & Giordano, 2012).
A classification system such as the DSM accomplishes many goals. It is fundamental to the study and treatment of psychopathology because it provides a common language for researchers, clinicians, educators, and students to communicate with one other. It serves as a central tool for the conceptualization of psychiatric cases, development of treatment protocols, dissemination of research, and education of future mental health workers. It also allows for interdisciplinary discourse, insurance reimbursement, and a foundation for common language among the general public.
Despite its widespread influence and impact, the DSM has been criticized for lacking appropriate validity, reliability, and clinical utility (Kotov et al., 2017; Widiger et al., 2019; Widiger & Simonsen, 2005). Identified limitations of the DSM approach to diagnosis are numerous. To start, most diagnoses are partly or fully polythetic, meaning that the diagnosis is made if a certain number of diagnostic criteria are endorsed from a longer list. Defining psychological disorders in this way produces significant heterogeneity within diagnostic groups because of the multiple permutations of symptom combinations that yield a positive diagnosis. This leads to the loss of valuable clinical information about a client because it eliminates details about what symptoms the client is experiencing, the onset and chronicity of the individual’s symptoms, and the severity of those indicators (Carragher et al., 2015). As an example, an individual can meet DSM-5 criteria for borderline personality disorder (BPD) by endorsing five or more of a possible nine diagnostic criteria. Two clients might share the experience of only one of the nine criteria and still be grouped under the same diagnostic label of BPD, despite what are undoubtedly distinct clinical symptom sets that could also vary in severity and course.
Further, the dominance of the DSM has led to a proliferation of “disorder-specific” treatment protocols. As evidence, Division 12 of the American Psychological Association evaluates and maintains a list of empirically supported treatments for specific psychological diagnoses (e.g., family-focused therapy for bipolar disorder). Despite the disclaimer that the resource is nonexhaustive, the list includes almost 90 interventions (American Psychological Association Division 12, n.d.). This approach to treatment poses significant training burdens for clinicians and inadequately addresses comorbidity within clinical cases.
Indeed, an additional consequence of the DSM system is that comorbidity is the norm in clinical and community samples (Conway et al., 2019; Kotov et al., 2017; Ruggero et al., 2019). This excessive co-occurrence of diagnoses challenges the notion that the DSM’s categorical diagnoses are discrete entities, and it complicates empirical and clinical work. For example, clinical trials that employ exclusion criteria barring potential participants with comorbid diagnoses greatly limit the generalizability of the studies (e.g., Zimmerman et al., 2019).
Another concern with the DSM system is that most patients are categorized with ambiguous “unspecified” or “not otherwise specified” diagnoses because they do not meet often arbitrary diagnostic thresholds for specific disorders (Achenbach, 2015; Carragher et al., 2015; Kotov et al., 2017). These unspecified diagnoses have reduced utility in research or clinical practice, as they provide minimal information about an individual’s clinical presentation.
Many criticisms of the DSM are rooted in the manual’s assumption that mental disorders are discrete categorical phenomena, each with putatively distinct etiologies, courses, and treatments. Thus, each disorder has a threshold for meeting diagnosis, but few of these thresholds have been demonstrated empirically. On the contrary, evidence has repeatedly indicated that psychopathology exists on continua that include a full range of severity as well as adaptive functioning (Achenbach, 2015; Carragher et al., 2015; Eaton et al., 2011; Kotov et al., 2017). Taken together, these concerns have led some researchers in the field to conclude that the DSM’s categorical diagnostic system is inadequate (Conway et al., 2019; Kotov et al., 2017; Widiger et al., 2019).
The Hierarchical Taxonomy of Psychopathology (HiTOP) has been offered as a solution to many of the criticisms of the DSM (Kotov et al., 2017; Sauer-Zavala, 2022; Widiger et al., 2019). HiTOP is an empirically derived system for classifying all variants of psychopathology (Kotov et al., 2017) along continua that reflect patterns of covariation. Broadly, this model conceptualizes the structure of psychopathology as a hierarchy: At the bottom of the hierarchy are psychiatric signs and symptoms, and at the top are increasingly broad dimensions such as internalizing and externalizing factors (Caspi et al., 2014; Kotov et al., 2017; Lahey et al., 2012).1 Importantly, HiTOP does not differ largely from the DSM in what constitutes signs and symptoms of mental illness; instead, HiTOP represents an evidence-based reorganization of these lower-level constructs. This reorganization is designed to increase specificity of diagnostic identifiers while reducing artifactual comorbidity across diagnostic groups and heterogeneity of symptoms within them. A clinical transition toward HiTOP will allow existing treatments to be extended or modularized to target the more homogenous constructs of the hierarchy, and several researchers have offered explicit guidelines for its incorporation into clinical practice (e.g., Hopwood et al., 2020; Ruggero, 2021; Ruggero et al., 2019; Stanton et al., 2019).
HiTOP represents a synthesis of available evidence and builds on a sturdy, long-standing foundation of research into the predominantly dimensional nature of psychopathology across the life span (Achenbach, 1966, 2015; Carragher et al., 2015; Krueger & Eaton, 2015). The developmental psychopathology approach has even successfully implemented this understanding into practice (Rutter, 2013; Rutter & Uher, 2012). Nonetheless, this dimensional approach challenges many long-accepted ideologies in the classification of psychopathology.
It is therefore not surprising that HiTOP’s recommendation to replace the DSM and its categorical diagnostic system has been met with criticism and doubt by other scholars in the field (Haeffel et al., 2022; Tyrer, 2018; Wittchen & Beesdo-Baum 2018; Zimmerman, 2021). One scholar likened the movement away from the DSM to “throw[ing] out the proverbial baby with the bathwater” (Zimmerman, 2021, p. 71), and others stated that “the HiTOP consortium is writing checks their taxonomy cannot cash” (Haeffel et al., 2022, p. 272). A common sentiment in these critiques is a lack of evidence that the HiTOP dimensions are clinically useful or that clinicians will be able to apply them. Critics have also hypothesized resistance from clinicians to adopt a novel system.
Importantly, there is no data regarding clinician perception of the full HiTOP model. Data pertaining to the clinician perception of and satisfaction with HiTOP-aligned measures such as the Achenbach System of Empirically Based Assessment (ASEBA), Minnesota Multiphasic Personality Inventory (MMPI), and Personality Assessment Inventory (PAI) are also lacking, only furthering the uncertainty of how clinicians would perceive a dimensional approach. There is a growing body of research regarding the clinician utility of dimensional personality pathology models, which is one component of the HiTOP system. Past research has determined that clinicians find dimensional models of personality pathology acceptable and often preferable to the categorical conceptualization of the DSM (Bornstein & Natoli, 2019; Glover et al., 2012; Hansen et al., 2019; Morey et al., 2014). For example, studies have found that clinicians consider dimensional personality traits to be more useful than categorical personality disorder diagnoses for clinical decision-making, treatment planning, comprehensively covering client difficulties, communication with the client, and generating global personality descriptions (Samuel & Widiger, 2006, 2011). One meta-analysis (Bornstein & Natoli, 2019) examined the clinical utility ratings of categorical and dimensional approaches to rating personality pathology across 11 studies, and the authors concluded that dimensional models were favored across the majority of indices of clinical utility. Nonetheless, it is not yet known whether these sentiments would extend beyond personality traits.
It is noteworthy that the clinical utility of the DSM has not been adequately surveyed (First et al., 2014). More than a decade ago, Mullins-Sweatt and Widiger (2009) detailed this concern, noting, “an entirely valid DSM could actually have relatively little clinical utility if it is not feasible for usage in general clinical practice” (p. 310). One study, accomplished for the APA’s DSM-5 field trials, concluded that the DSM-5 was feasible and clinically useful because 46% of clinicians identified the manual as “very or extremely easy” for assessment, and 46% reported that it was “useful or extremely useful” for routine clinical practice with adult clients (Mościcki et al., 2013). Another survey of psychiatrists found that the DSM was most useful for assigning diagnoses, teaching trainees or students, and communicating with other mental health providers, whereas it was least useful for selecting a treatment or assessing prognosis (First et al., 2018). In a separate sample, First and colleagues (2019) found that, at best, clinicians found the DSM to be “very or extremely useful” for administrative requirements (58% of respondents) or communicating diagnoses with other providers (55%) and, at worst, for selecting a treatment (23%) or determining a prognosis (18%). Beyond this handful of studies, quantitative information regarding clinician perception of the full DSM system is limited.
Given the lack of evidence, questions about the clinical application of HiTOP are certainly reasonable, and the clinician perception of this diagnostic approach is an important unexplored question. Indeed, in moments of differing opinions among researchers, perhaps it is most appropriate to consider the perspective of those most closely involved in patient care. The voices of clinicians are needed in the ongoing empirical shift away from the DSM, and we particularly need to assess the clinician perception of the utility, accessibility, and helpfulness of the full HiTOP model (Ruggero et al., 2019; Tyrer, 2018).
Accordingly, our goal in the present study was to assess the clinician perception of the HiTOP and DSM systems for the conceptualization of vignettes drawn from real clinical cases. Clinicians rated one of three clinical vignettes according to the three highest-order HiTOP levels: superspectra (i.e., the p factor), spectra (e.g., somatoform, internalizing, thought disorder), and subfactors (e.g., eating pathology, substance abuse, mania). These levels of the hierarchy were prioritized because these are the components of the system that differ most from the DSM. Given that there is limited distinction between the diagnoses of the DSM and the syndrome/disorder level of HiTOP at present, the fourth level of HiTOP was not assessed. The fifth level of HiTOP, symptom components and maladaptive traits (e.g., anxiousness, checking behaviors, avolition), was also not assessed because of similarities across systems. Additionally, signs, symptoms, and components were already presented in the vignettes, so to have clinicians rate these would be redundant. Further, this level includes more than 80 constructs, and it was not feasible for clinicians to rate them on a brief survey with minimal compensation.
In addition to rating the client on the three levels of HiTOP, clinicians rated them according to all possible DSM-5 diagnoses. Notably, whether the use of the DSM was typical for any given clinician, this approach largely dominates clinical training and billing for services, so all clinicians were likely familiar with the DSM approach and associated diagnoses. Following diagnostic ratings, the clinicians completed surveys assessing their subjective satisfaction with and overall perception of the HiTOP and DSM systems, as well as how the dimensional approach compared with their typical diagnostic categorization and their experience with the DSM.
Although the DSM carries the advantages of familiarity and existing infrastructure, prior research has evidenced a clinician preference for dimensions in conceptualizing personality pathology (e.g., Bornstein & Natoli, 2019; Samuel & Widiger, 2006, 2011). Accordingly, we hypothesized that clinicians would rate the HiTOP model equally or more favorably than the DSM on measures of clinical utility.
Method
The study was preregistered on OSF prior to data collection (https://doi.org/10.17605/OSF.IO/FNQ5D). There were no deviations from the preregistration, but we conducted additional exploratory analyses (reported below). The study was approved by Purdue University’s Institutional Review Board (Protocol No. IRB-2021-713). We report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study.
Participants
The study was advertised to clinicians via several electronic channels. This included emailing the clinical training directors of U.S. Veterans Affairs, U.S.-based training directors of the Association of Psychology Postdoctoral and Internship Centers (APPIC), members of the Indiana Psychological Association (IPA), and members of the National Association for Training Clinic Directors (NATCD). The investigators also shared the study via their professionally affiliated Twitter accounts and on the email LISTSERV and Facebook page for the Association for Behavioral and Cognitive Therapies (ABCT).
Two-hundred thirty-seven individuals who self-identified as licensed and actively practicing clinicians entered the study. Of these, 63 participants dropped out of the survey without completing any outcome items. Thirty Facebook bot responses were eliminated on the basis of identical and off-topic repeat responses in the free-response variable of the survey (e.g., 10 responses stating, “We can collect some questions about the mental illness of teenagers in the community”; six responses stating, “The psychological changes of some people in a specific environment”), all of which were responses to the Facebook group advertisement. Finally, as per our preregistration, we excluded respondents who completed the survey in under 4 min for a lack of effortful responding (n = 1). This resulted in a final sample size of 143 clinicians.
Procedure
After consenting to participate, respondents confirmed their status as actively practicing mental health providers who had completed a graduate degree in a mental-health-related field. If a respondent did not endorse these characteristics, the survey ended.
The study used a 3 (vignette; between subjects) × 2 (diagnostic system; within subjects) design. Qualifying clinicians read brief introductions to the two diagnostic systems (see Supplement A in the Supplemental Material available online). Next, each read one of three randomly assigned clinical vignettes that populated in a new window for continued reference throughout the study. The clinicians then made diagnostic ratings of the described client according to both the HiTOP and DSM-5 approaches; the order of systems was counterbalanced. Finally, the clinicians rated the clinical utility of both approaches. After study completion, respondents had the option to enter a raffle for one of thirty $20 gift cards to their choice of Amazon, Etsy, or Target.
Measures
The study’s primary outcome was the clinical utility of the target diagnostic approach as rated on the Clinical Utility Questionnaire (CUQ; First et al., 2004; Samuel & Widiger, 2006). Clinicians rated items on a 5-point Likert scale (1 = not at all useful, 5 = extremely useful). Clinicians completed this for both the HiTOP and DSM-5 systems. This measure includes six items that assess the utility of the target diagnostic system for (a) communicating information about the client with other mental health professionals, (b) ease of applying the system to the individual, (c) communicating with the clients about themselves, (d) comprehensively describing the individual’s psychopathology, (e) formulating effective interventions for the client, and (f) describing the individual’s global functioning. Composite scores for the overall clinical utility of each model were generated by summing the CUQ item scores. The total score was out of a possible 30 points, in which higher scores indicated greater utility.
Clinicians provided demographic information including their age, gender identity, and race or ethnicity. Additionally, a variety of professional characteristics of the clinicians were collected, such as theoretical orientation, typical methods of assessment, and years of clinical experience (Table 1). We also collected information regarding the clinicians’ familiarity with the DSM and HiTOP systems and whether they had a preference between the two approaches. Finally, at the end of the survey, respondents had the option to provide any thoughts on the topics of the study in a free-response box.
Professional Characteristics of Clinician Participants
Note: Participants had an average of 12.59 years of clinical experience (SD = 9.15) and spent an average of 47.08% working hours on direct client contact (SD = 25.36). CBT = cognitive behavioral therapy; ACT = acceptance and commitment therapy; DBT = dialectical behavioral therapy.
Participants could make multiple selections for these variables.
Procedure for applying diagnostic systems
The HiTOP consortium is currently developing a comprehensive, omnibus HiTOP assessment, but it is not expected to debut for some time (Ruggero et al., 2019; Simms et al., 2022). Even if such a measure were ready for use, it would likely be too long to reasonably expect clinicians to complete for the present study. Instead, the clinicians read brief descriptions of the various constructs and then rated whether, for the client depicted in the vignette, each construct was (a) not at all a problem, (b) somewhat of a problem, or (c) very much a problem. This response format was modeled after the scoring system for the gold standard of diagnosing DSM psychopathology, the Semi-Structured Clinical Interview for the DSM (SCID), in which raters score a clinical construct on a 3-point range: absent, subthreshold, and threshold. Definitions and ratings were prepared for three components of the HiTOP model, including superspectrum (general psychopathology/p factor), spectra (somatoform, internalizing, thought disorder, externalizing disinhibited, externalizing antagonistic, detachment), and subfactors (sexual problems, eating problems, fear, distress, mania, substance use, antisocial behavior; see Supplement B in the Supplemental Material).
The presentation of HiTOP constructs began at the top of the hierarchy, with the superspectrum. The survey used skip logic: Clinicians were asked only to score constructs that fell beneath the umbrella of difficulties they identified as “somewhat” or “very much” a problem on the previous level of the hierarchy. Patients could be rated as displaying elevated levels of multiple constructs per level. This followed the intended stepwise, cascading approach encouraged by the clinical workgroup of the HiTOP consortium. For example, if out of the five spectra, a clinician indicated only that the internalizing spectra was a problem for the client, then the clinician rated only the subfactors of the internalizing spectra: sexual problems, eating problems, fear, and distress. They were not asked to rate the mania, substance abuse, or antisocial behavior subfactors of the other spectra.
When the clinicians were asked to categorize the client according to the DSM-5, they were invited to use any reference materials (e.g., their DSM). The diagnostic options were organized to reflect the table of contents of the DSM-5, such that there were separate lists for each of the following diagnostic classes: neurodevelopmental disorders, schizophrenia and related disorders, bipolar and related disorders, depressive disorders, anxiety disorders, obsessive-compulsive and related disorders, trauma and stressor-related disorders, dissociative disorders, somatic symptom disorders, feeding and eating disorders, elimination disorders, sleep-wake disorders, sexual dysfunctions, gender dysphoria, impulse-control and conduct disorders, neurocognitive disorders, substance-related and addictive disorders, personality disorders, paraphilic disorders, and other mental disorders (see Supplement C in the Supplemental Material; APA, 2013). The clinicians could select as many diagnoses as they desired and did not have to make a selection for categories they perceived as irrelevant.
Vignette development
The clinical vignettes were designed to reflect the type of cases seen in practice and to include the types and amount of information that would be gleaned from a therapy intake. Clinical information was presented in terms of specific thoughts, behaviors, and problems that clients share during an intake, rather than in terms of traditional diagnostic criteria. Further, vignettes sculpted to be prototypical examples of DSM diagnoses were avoided because such prototypical, single-diagnosis cases are not reflective of the majority of outpatient clinical presentations (Zimmerman, 2016; Zimmerman et al., 2007). Instead, we aimed for diagnostically complex presentations drawn from real cases. This objective also aligned with the recommendation that future clinical utility research incorporate real clinical cases (Bornstein & Natoli, 2019).
Such vignettes of an appropriate length were not readily available in the literature. Instead, we derived vignettes from published clinical case studies to meet these goals and minimize investigator bias during vignette development. The first author examined the table of contents of every issue of the Clinical Case Studies journal over the past 5 years (i.e., 2016–2021). Publications that suggested some level of case complexity instead of a single diagnosis (e.g., titles referencing transdiagnostic treatment or diagnostic comorbidity) were screened for consideration as a vignette template. The three case studies that provided the most detail about psychosocial history and psychiatric symptoms were selected as the bases for the vignettes (Lui, 2017; Scheiderer et al., 2017; Smith et al., 2020). Multiple vignettes were developed to ensure that observed effects were not attributable to the characteristics of a certain case.
We consolidated each of these case studies into one-page vignettes (see Supplement D in the Supplemental Material) to streamline presentation and minimize the clinicians’ time commitment. The vignettes described the client’s gender, age, occupation, presenting problems, current psychiatric symptoms, interpersonal and occupational functioning, and symptom chronicity and course. To maximize experimental control, we held age and gender constant across vignettes. Previous data indicates that among recipients of mental health care in the United States, the majority of clients are female and between the ages of 25 and 39, or on average about 32 years of age (Center for Behavioral Health Statistics and Quality, Substance Abuse and Mental Health Services Administration, 2020). Thus, all vignettes depicted a 32-year-old woman. In describing psychiatric symptomatology, we generally avoided traditional diagnostic labels (e.g., “major depressive disorder”) and grouped criteria lists (e.g., “feelings of emptiness, affective instability, marked impulsivity, chronic suicidality”). We expanded certain details of symptomatology to mirror how a client would likely present the information (e.g., altering “sleep difficulties” to “waking frequently throughout the night and not being able to go back to sleep, sometimes waking for the day as early as 3 or 4am”).
Data analysis
Analyses were conducted in SPSS (Version 26; IBM Corp., 2019). For statistical inferences, α was set at .05, two-tailed. We adjusted the p-value criterion for exploratory analyses using false-discovery rate (FDR) correction (Benjamini & Hochberg, 1995) to control for Type I error. Effect sizes are reported for all main effects as η p 2s (η p 2 ≥ .01 indicates a small effect, η p 2 ≥ .06 indicates a medium effect, and η p 2 ≥ .14 indicates a large effect; Cohen, 1988). There were no missing items on the CUQ across all responses.
This study had a mixed design with both within- and between-subjects elements. The primary analysis was a 2 (diagnostic model: DSM vs. HiTOP) × 3 (vignette assignment: A, B, or C) mixed-model analysis of variance (ANOVA) with diagnostic model as the within-subjects variable, vignette assignment as the between-subjects variable, and clinical utility scores as the outcome. All pair-wise comparisons and simple effects were examined following significant omnibus tests. Descriptive information from the questions pertaining to clinicians’ demographic and professional characteristics are reported and, when applicable, correlated with the utility ratings for the DSM and HiTOP approaches.
Power analysis
Prior to data collection, G*Power (Faul et al., 2007) was used to conduct power analyses. We aimed to achieve 90% power to detect a small effect of Cohen’s f = 0.15 (f values between 0.10 and 0.24 are considered small) at the standard .05 α error probability. In our sample, this effect size was equivalent to an η p 2 of approximately .03. This target effect size was necessarily conservative given that the extant literature of clinician comparison of categorical and dimensional personality models has yielded anywhere from negligible to large effects (e.g., Glover et al., 2012; Hansen et al., 2019; Morey et al., 2014; Samuel & Widiger, 2006). To detect a difference between the HiTOP and DSM clinical utility ratings using a 2 (diagnostic system) × 3 (vignette assignment) mixed-model ANOVA with 90% power and no correlation between the two measures, G*Power suggested a sample size of 237.
A sample size of 237 clinicians was the best-case scenario. With a less conservative target of 80% power to achieve a small effect of Cohen’s f = 0.2, the target sample size was 102. Of note, this retained the specification that the repeated measures were not intercorrelated. If these values were intercorrelated even by .1, the total sample size specified by G*Power dropped to 93; a correlation of .3 dropped the sample to 72 and so on. In our sample, the repeated measures were correlated (r = –.14). A post hoc sensitivity analysis revealed that with this correlation of –.14, the present study had 80% power to detect f = 0.18, a small effect. This confirms that our study was adequately powered to detect small effects within the obtained sample.
Results
The order in which participants rated the diagnostic systems was counterbalanced, and there were no observed ordering effects for ratings of the DSM, t(141) = −0.075, p = .940, or HiTOP, t(141) = 0.15, p = .879. Professional characteristics of the sample of clinicians are listed in Table 1, and their demographic characteristics are detailed in Table 2. The sample was comprised of primarily female (62.7%) and White (86.7%) clinicians from the clinical psychology discipline (77.6%) with a mean age of 40.74 years (SD = 10.80). The most frequently endorsed theoretical orientations were second-wave cognitive behavioral therapy (CBT; 64.3%), third-wave CBT (53.2%), and interpersonal (29.4%). Most reported working with adult clients (83.2%), and the most common settings for clinical work were private practice (20.3%), Veterans Affairs (16.8%), and academic medical centers (15.4%). The sample was primarily comprised of clinicians with a PhD (51.4%) or doctorate of psychology (PsyD; 32.3%). On average, clinicians reported having 12.59 years (SD = 9.15) of clinical experience and spending 47.1% (SD = 25.36) of their working hours on direct client contact. Across participants, 87 reported having access to a DSM that they did not use during the study, 33 reported having access to a DSM that they used while completing the study, and 22 indicated that they did not have access to a DSM during the study.
Demographic Characteristics of Clinician Participants
Note: Participants had a mean age of 40.74 years (SD = 10.80).
Participants could make multiple selections regarding their race/ethnicity.
Results broadly evidenced higher clinical utility ratings of the HiTOP system compared with the DSM system. A 2 × 3 mixed-model ANOVA was conducted for predicting overall CUQ score, with a possible range of 6 to 30 points (Table 3). The main effect of diagnostic system revealed that HiTOP (M = 20.52, SD = 4.81) was rated more positively than the DSM system (M = 17.71, SD = 4.63) at a statistically significant level, F(1, 140) = 23.79, p < .001, η p 2 = .15. 2
Comparison of Clinical Utility Scores for DSM and HiTOP When Controlling for Vignette Assignment
Note: DSM = Diagnostic and Statistical Manual of Mental Disorders (Fifth Edition); HiTOP = Hierarchical Taxonomy of Psychopathology; CUQ = Clinical Utility Questionnaire.
We repeated the 2 × 3 ANOVA for predicting scores on each of the six individual CUQ items, as they assess incrementally informative components of utility (Table 3). These analyses were listed as potential exploratory analyses in the preregistration, so we adjusted p values using FDR correction. For each item, the minimum possible score was 1 point, and the maximum was 5 points. The HiTOP system was rated more favorably for five of six of the components of utility, with two effect sizes considered to be large, two medium, and one small. Specifically, HiTOP scored higher than the DSM for describing global functioning (mean difference = 0.86), F(1, 140) = 49.71, p < .001, η p 2 = .26; comprehensively describing psychopathology (mean difference = 0.73), F(1, 140) = 38.92, p < .001, η p 2 = .22; formulating effective interventions (mean difference = 0.54), F(1, 140) = 21.38, p < .001, η p 2 = .13; communicating clinical information to the client (mean difference = 0.50), F(1, 140) = 17.31, p < .001, η p 2 = 0.11; and ease of applying the system to the individual (mean difference = 0.29), F(1, 140) = 6.72, p = .011, η p 2 = .05. There was no statistically significant difference between the DSM and HiTOP systems for communicating information about the individual to other mental health providers (mean difference = −0.11), F(1, 140) = 0.73, p = .395, η p 2 = .01.
There were no statistically significant effects of vignette assignment for any outcome, including overall CUQ score, F(2, 140) = 0.55, p = .578, η p 2 = .01; communicating information about the client with other mental health professionals, F(2, 140) = 0.56, p = .553, η p 2 = .01; ease of applying the system to the individual, F(2, 140) = 0.24, p = .789, η p 2 = .00; communicating with the clients about themselves, F(2, 140) = 0.56, p = .564, η p 2 = .01; comprehensively describing the individual’s psychopathology, F(2, 140) = 0.12, p = .883, η p 2 = .00; formulating effective interventions for the client, F(2, 140) = 0.99, p = .374, η p 2 = .01; and describing global functioning, F(2, 140) = 0.51, p = .601, η p 2 = .01.
There were statistically significant interaction effects between diagnostic system and vignette assignment for four of these seven analyses. All interaction and simple effects pertaining to the analysis of individual CUQ items were adjusted using FDR correction. There was an interaction effect related to the overall CUQ score, F(2, 140) = 3.74, p = .026, η p 2 = .05; specifically, there was a statistically significant difference between diagnostic systems for Vignette A (mean difference = −5.13, p < .001) but not Vignette B (mean difference = −1.76, p = .081) or Vignette C (mean difference = −1.67, p = .101). There was also an interaction effect for ease of applying the system to the individual, F(2, 140) = 6.18, p = .003, η p 2 = .08, and simple-effects comparisons displayed a significant difference between diagnostic systems within Vignette A (mean difference = −0.89, p < .001) but not Vignette B (mean difference = 0.02, p = .919) or Vignette C (mean difference = −0.04, p = .837). Similarly, there was an interaction effect for formulating effective intervention outcomes, F(2, 140) = 5.91, p = .003, η p 2 = .08, and simple-effect comparisons displayed a significant difference between diagnostic systems within Vignette A (mean difference = −1.13, p < .001) but not Vignette B (mean difference = −0.25, p = .229) or Vignette C (mean difference = −0.27, p = .188). Regarding the interaction effect for the “comprehensively describing psychopathology” item, F(2, 140) = 3.59, p = .030, η p 2 = .05, there were still statistically significant effects between diagnostic systems for all vignettes: Vignette A (mean difference = −1.17, p < .001), Vignette B (mean difference = −0.43, p = .035), and Vignette C (mean difference = −0.60, p = .004). There were not statistically significant interaction effects for describing global functioning, F(2, 140) = 3.74, p = .026, η p 2 = .05; communicating to the individual, F(2, 140) = 1.83, p = .164, η p 2 = .03; or communicating with other providers, F(2, 140) = 0.90, p = .409, η p 2 = .01. Thus, omnibus effects of diagnostic system held across four of seven outcomes. Notably, the DSM was never favored for any omnibus or simple effects.
In accordance with the preregistered analytic plan, quantitative clinician factors including familiarity with the diagnostic systems, clinician age, years of clinical experience, and percentage of working hours spent with direct client contact were correlated with the CUQ outcomes (Tables S1a, S1b, and S1c in the Supplemental Material). Years of clinical experience was positively related to familiarity with the DSM (r = .27, p < .01) and negatively related with familiarity with HiTOP (r = −.30, p < .01). Years of clinical experience revealed several statistically significant negative relations with DSM outcomes: overall CUQ score (r = −.22, p < .01), communicating to the individual (r = −.20, p < .05), comprehensively describing psychopathology (r = −.18, p < .05), formulating effective interventions (r = −.22, p < .01), and describing global functioning (r = −.23, p < .01). There were also negative associations between clinician age and reported usefulness of the DSM for communicating to the patient (r = −.17, p < .05), as well as between clinician age and utility of the DSM for formulating effective interventions (r = −.17, p < .05). There was a positive association between familiarity with the DSM and reported ease of applying the DSM system (r = .18, p < .05). Only one statistically significant correlation between clinician factors and HiTOP-specific outcomes was observed; specifically, familiarity with HiTOP was positively associated with reported ease of HiTOP for communicating to other mental health providers about the individual (r = .18, p < .05).
Clinicians also rated how familiar they were with both diagnostic models, and 77 participants indicated that they were “not at all familiar” with HiTOP prior to participating in study. As an exploratory analysis to assess potential response bias toward clinicians familiar with HiTOP in our sample, we reran the 2 × 3 mixed-model ANOVAs with only those 77 participants who had no prior exposure to HiTOP. All p values were adjusted using FDR correction. The direction of effects held across all seven analyses, and the patterns of statistical significance held across six. HiTOP was still preferred for the outcomes of overall CUQ score (mean difference = 2.54, p = .001, η p 2 = .13), describing global functioning (mean difference = 0.97, p < .001, η p 2 = .33), comprehensively describing psychopathology (mean difference = 0.73, p < .001, η p 2 = .24), formulating interventions (mean difference = 0.48, p = .002, η p 2 = .12), and communication with the client (mean difference = 0.42, p = .017, η p 2 = .08). As with the initial analyses, there was no effect of diagnostic system for communicating with other providers (mean difference = −0.26, p = .094, η p 2 = .04). The only change was that HiTOP (M = 3.34, SD = 1.00) was no longer favored over the DSM (M = 3.14, SD = 0.90) for ease of applying the diagnostic system to the individual (p = .218, η p 2 = .02). As in the primary analyses described above, in this subsample, the DSM was not favored for any outcome.
When asked explicitly which model they would prefer for diagnosis in their clinical work, 73 clinicians (51.41%) indicated HiTOP, 54 (38.03%) indicated the DSM, and 15 (10.56%) indicated neither. Comments left by the clinicians in the free-response box at the end of the survey can be found in Supplement E in the Supplemental Material and are organized into comments about the DSM, about HiTOP, comparing both systems, and regarding the study itself. As an example, one clinician noted,
I love the idea of a dimensional approach to [diagnosis], but I don’t know how to use HiTop [sic]. I think having a specific diagnosis is often helpful for patients, as it helps them understand that they are not alone and that there is a name for what they are experiencing.
Another stated, “I appreciate these alternate diagnostic models and find them more useful than the DSM, whose only value for me is getting the proper code for insurance and flattening out wrinkly papers.”
Discussion
Since the debut of the HiTOP system, its clinical utility has been a crucial question. Prior to the present study, there were no data to speak to this question, and the present findings offer an encouraging perspective. Previous research has found that clinicians prefer HiTOP-friendly approaches to personality pathology over the DSM’s categorical personality disorder diagnoses (e.g., Bornstein & Natoli, 2019; Hansen et al., 2019; Samuel & Widiger, 2011), and the results of the present study reveal similar trends for a dimensional diagnostic system.
Among this sample of practicing clinicians, the DSM was not favored for any measure of clinical utility. Instead, findings coalesce with previous postulations regarding the potential benefits of HiTOP in clinical settings (e.g., Sauer-Zavala, 2022; Stanton et al., 2019). HiTOP was favored for its utility in communicating clinical information to the client, comprehensively describing psychopathology, and describing global functioning. This suggests that HiTOP is the better option for capturing the complexity of a client’s clinical presentation while offering an effective estimate of overall impairment, likely a reflection of the specificity of the lower hierarchy levels and the p factor at the highest hierarchy level. It also suggests that HiTOP may serve as a more accessible format for presenting diagnostic feedback to clients, an advantage for the crucial treatment factor of therapeutic alliance (Martin et al., 2000).
At the omnibus level, HiTOP was also favored for overall CUQ score, formulating effective interventions, and ease of applying the system to the individual, though simple-effects comparisons revealed that these findings were carried by Vignette A. This suggests that for some clients, clinicians perceive HiTOP as the better option for individualizing intervention services or as the easier system to use. Even so, such conclusions are revealing. One would expect that the DSM—with which our sample was largely trained and spent an average of 13 years using—would be less difficult to use. Further, the DSM system has many empirically supported treatment protocols based on its diagnoses (American Psychological Association Division 12, n.d.), and the same cannot yet be said for all components of the HiTOP system (Barlow et al., 2017; Hopwood et al., 2020). Despite what seems an auspicious advantage for the DSM in experience and treatment planning, clinicians again did not prefer it.
There was one omnibus outcome for which clinicians indicated no preference between HiTOP and the DSM: communicating information about the client to other mental health providers. It was somewhat unexpected that the DSM was not preferred for this metric of utility. Of the six clinical utility constructs assessed with the CUQ, the DSM arguably had its best chance to outperform HiTOP on this item because the DSM is the framework through which mental health care in the United States operates. Attempting to communicate with members of a DSM-entrenched system using a HiTOP lens could reasonably seem a difficult task.
This lack of preference for the DSM could prompt concern that the sample was driven by a response bias toward clinicians already familiar with HiTOP (and presumably favorable toward it). Rerunning the analyses including only the subgroup of 77 clinicians who were completely unfamiliar with HiTOP prior to the study was helpful in eliminating this concern. Indeed, the only observed change across the seven effects was that this subsample displayed no preference between systems for ease of applying the diagnostic system to the individual. HiTOP was preferred for the majority of utility constructs, and the DSM was preferred for none.
It seems feasible that the clinicians understood that to use a HiTOP lens for interprofessional communication would not mark a dramatic change in the way they communicate. Indeed, HiTOP includes the signs, traits, symptoms, and components that make up many of the diagnoses listed in DSM, and the two systems represent different ways of organizing these constructs. A clinician might already communicate at the level of signs and symptoms, for example describing a client as experiencing ruminative worry, insomnia, and difficulty concentrating. This description is not DSM specific. The lack of preference between HiTOP and the DSM for communicating with other providers might convey that clinicians do not rely on the DSM’s diagnostic labels to communicate with each other.
It is worth emphasizing the findings from the subgroup of HiTOP-naive clinicians beyond the lens of the professional communication item. These findings are quite striking because they indicate that HiTOP’s advantages were intuitive enough to be preferred over the DSM even among clinicians who were seeing it for the first time. This not only speaks optimistically of the promise of HiTOP but also counters concerns about its complexity (Haeffel et al., 2022).
Although our primary goal in the present study was to assess the perceived clinical utility of HiTOP, the implications regarding the perceived utility of the DSM are also notable. The DSM was not favored for any clinical utility outcome across all analyses. This is of concern given the influence and impact of the DSM, and these results should raise alarm to researchers advocating the continued implementation of the DSM system.
It is difficult to determine whether ratings favoring the utility of HiTOP were driven by an appreciation of HiTOP, a dislike of the DSM, or some combination of both. Trends across the limited free-response information collected in this study suggest a combination of the two forces. In the future, structured qualitative surveys could be well suited for arbitrating such comparisons.
Finally, despite effects favoring HiTOP across the CUQ, when clinicians were asked which diagnostic system they would prefer to use in clinical practice, only the slightest majority favored HiTOP: 51.41% said HiTOP, 38.03% said the DSM, and 10.56% said neither. This reflects the substantial progress that must be made for either diagnostic system to succeed in the long term.
Limitations
As described, Vignette A carried the main effects of HiTOP preference for a minority of outcomes: overall CUQ score, ease of applying the system, and formulating effective interventions. Clinicians assigned to Vignettes B or C did not indicate a preference between HiTOP and the DSM for these outcomes. Importantly, simple-effects comparisons across all outcomes never revealed a preference for the DSM; what varied was whether HiTOP was preferred or if there was no preference between HiTOP and the DSM.
We did not forge systematic differences between the vignettes on purpose, so it is difficult to know why Vignette A was associated with certain higher utility scores for HiTOP. This vignette required almost no elaborations or additions on behalf of the principal investigators (Balling and Samuel), as its associated case study (Scheiderer et al., 2017) provided ample detail about the symptoms the client was experiencing. Perhaps the revisions to Vignettes B and C were biased toward the DSM on the part of the principal investigators given their training backgrounds. Perhaps Vignette A came across as a less prototypical diagnostic case study. There is no doubt that prototypicality likely affects perceived utility of systems (e.g., Samuel & Widiger, 2011), but there remains a question of degree. The vignettes intentionally incorporated a degree of diagnostic complexity/comorbidity, as this is most representative of clients who present for outpatient treatment (Zimmerman, 2016; Zimmerman et al., 2007). Of course, not all clients will present with diagnostic ambiguity or comorbidity. Thus, research into meaningfully varied vignettes and associated effect differences would be highly informative.
As another limitation of the study, the approaches for applying the two diagnostic systems were imperfect proxies. For example, the HiTOP approach was likely more time and effort intensive than that of the DSM because it involved reading and processing introductory descriptions of each HiTOP construct, and this could have been of some detriment to the perception of HiTOP. Alternatively, the HiTOP approach telescoped the diagnostic process according to the clinician’s prior ratings, whereas the DSM approach presented all possible diagnostic constructs at once. Perhaps taking a screener approach to the DSM that guided clinicians toward likelier diagnoses (e.g., the SCID screener) would have resulted in higher DSM utility scores. Nevertheless, the diagnostic proxies accomplished the main goal of obligating the clinician to effortfully consider the DSM and HiTOP systems in relation to the client before making clinical utility ratings.
Future directions
This area of research would benefit from a broader measure of clinical utility to include domains such as assessing risk or developing therapeutic alliance. Interpretation guidelines for the CUQ would also be beneficial, as it is difficult to translate observed CUQ score differences into terms of clinical impact. Future research should explore cut points that might reflect the strength of a system’s clinical utility.
About three quarters of our sample consisted of mental health professionals from clinical psychology backgrounds. Future examinations of the clinical utility of HiTOP would benefit from obtaining ratings by other types of providers, particularly psychiatrists. It is reasonable to suspect that psychiatrists would differ in their preferences for the DSM over HiTOP, particularly because the DSM is a product of the primary governing body of psychiatry in the United States. It would also be of interest to assess utility beyond mental health clinicians to include providers such as primary care physicians, who often have a role in mental health treatment and associated referrals (Bornstein & Natoli, 2019). Further, this research must eventually extend beyond vignettes and into psychiatric clinics and hospitals, to be applied to the actual clients of surveyed clinicians.
The ultimate measure of clinical utility will be to assess whether a HiTOP approach improves treatment outcomes, and commentaries and criticisms of the system have made this clear (e.g., Tyrer, 2018; Zimmerman, 2021). Further exploration into the predictive validity of diagnostic models is of the utmost importance for future research, both for HiTOP and the DSM, as data is notably lacking in this area for both systems. It is worth emphasizing that the goal is to employ whatever system is most useful for improving the mental health of our clients. These are not merely academic questions. 3
Conclusion
Given the study’s limitations and in the absence of other data, our response to these results is not to suggest that HiTOP must urgently replace the DSM. However, for critics that assert HiTOP is too complicated, cumbersome, or unapplicable for clinicians, these data suggest this is currently unfounded. Theirs are certainly reasonable concerns, raised in the absence of data. Now that at least these data exist, we hope that a portion of these apprehensions are alleviated. From this vantage point, it appears that when it comes to HiTOP, clinicians are capable and interested.
A 2018 World Psychiatry commentary concluded, “if [HiTOP] can come forward with more clinical meat to add to their helping of science, things will certainly change” (Tyrer, 2018, p. 296). Ultimately, this single study does not serve as the definitive answer regarding the clinical utility of HiTOP or the DSM. However, the conclusion of the present study is that actively practicing clinicians displayed preference for HiTOP over the DSM in measures of clinical utility. We hope this study is an encouraging serving.
Supplemental Material
sj-docx-1-cpx-10.1177_21677026221138818 – Supplemental material for Clinician Perception of the Clinical Utility of the Hierarchical Taxonomy of Psychopathology (HiTOP) System
Supplemental material, sj-docx-1-cpx-10.1177_21677026221138818 for Clinician Perception of the Clinical Utility of the Hierarchical Taxonomy of Psychopathology (HiTOP) System by Caroline E. Balling, Susan C. South, Donald R. Lynam and Douglas B. Samuel in Clinical Psychological Science
Footnotes
Transparency
Action Editor: Aidan G. C. Wright
Editor: Jennifer L. Tackett
Author Contributions
Correction (March 2023):
Article updated to include the Open Data, Open Materials, and Preregistration badges and the Open Practices statement.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
