Abstract
The Alternative Model of Personality Disorders distinguishes between the severity of personality dysfunction (Criterion A) and individual differences in personality disorder expression (Criterion B). Several Criterion A measures exist, but few studies have compared these measures with each other. Moreover, debates about whether the constructs of Criteria A and B are redundant (i.e., weak incremental validity) should be framed around how different Criterion A measures perform relative to others. This study of 204 undergraduate students evaluated multiple measures of Criterion A. These measures were strongly correlated with Criterion B, but evidenced incremental validity (39% of outcomes, 5% average additional variance explained) with outcomes of psychopathology and interpersonal impairments, and less consistent incremental validity with suicidality, aggression, and mental health utilization. We discuss how these results inform the construct of Criterion A relative to Criterion B and evaluate strengths/weaknesses of Criterion A measures.
The Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5; American Psychiatric Association [APA], 2013) includes an alternative model for personality disorders (AMPD). This alternative model distinguishes between the severity of personality dysfunction (Criterion A) and the stylistic expressions in which personality disorder manifests (Criterion B). Criterion A is conceived as a single dimension reflecting self and interpersonal impairment, a description that accords with psychodynamic, interpersonal, personological, and attachment theories (Bender et al., 2011; Mulay et al., 2018; Pincus & Roche, 2019). Criterion B is conceived as a pathological trait model, usually represented through five pathological personality factors that link well with conceptualizations from other pathological trait models (Van den Broeck et al., 2014; Wright & Simms, 2014), normal personality trait models (Helle et al., 2017; Suzuki et al., 2015), and Research Domain Criteria (RDoC) constructs (Harkness et al., 2014) and also appear to capture the larger metastructure of psychopathology (Krueger, 2013; Wright & Simms, 2015). In the DSM-5, Criterion B is operationalized using the Personality Inventory for DSM-5 (PID-5; APA, 2013) or the brief form (PID-5-BF; APA, 2013).
The AMPD has received increasing research interest (Milinkovic & Tiliopoulos, 2020; Widiger et al., 2019; Zimmermann et al., 2019) and the World Health Organization accepted a similar model in the International Classification of Diseases (11th edition; ICD-11; World Health Organization, 2020). Yet, fundamental questions still remain in this new conceptualization of personality disorder. Specifically, how to best assess these dimensions of psychopathology and whether Criterion A is redundant with Criterion B. As such, this research examines several measures of Criterion A to determine their reliability, overlap, and validity, including whether any measure of Criterion A demonstrates sustained incremental validity over Criterion B in relation to important outcomes.
Assessment of Criterion A
Criterion A is operationalized in the DSM-5 with the Level of Personality Functioning Scale (LFPS; APA, 2013). It was originally designed to be a clinician-rated instrument and, for several years, there was not a self-report version directly operationalizing Criterion A, which led to research support of Criterion A lagging behind Criterion B (Roche et al., 2016). Now there are at least five self-report measures and two interview-rated measures of Criterion A (each reviewed below). Some of these measures have just been developed in recent years, and as such, there are few studies examining these measures, and even fewer comparing them. Yet this work is crucial in determining how choices of measurement may impact the validity of Criterion A reflected in the literature.
Incremental validity addresses the question of redundancy in the AMPD model as many have noted that there is conceptual and empirical overlap within constructs in Criteria A and B (e.g., Widiger et al., 2019). This is typically accomplished by adding Criterion B predictors in first, and then Criterion A predictor(s) in second, and vice versa, to examine additional variance explained in the second step. Across many studies, the trend is that Criterion A has less consistent evidence of incremental validity compared with Criterion B.
Measurement matters when assessing for overlap, as the choice of outcomes and assessment instruments for Criteria A and B factors into how redundant a set of constructs appear. Research that uses Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition (DSM-IV; APA, 1994) personality disorders as the key outcome finds less consistent evidence of incremental validity for Criterion A (Few et al., 2013; Hentschel & Pukrop, 2014; Hopwood et al., 2012), compared with outcomes that assess global impairment (e.g., Zimmermann et al., 2014).
For Criterion B, several studies have shown that using a normal range personality measure increases the likelihood of finding incremental validity of Criterion A, whereas using pathological traits decreases this likelihood (e.g., Berghuis et al., 2012; Hentschel & Pukrop, 2014; but see Sleep et al., 2020). Morey et al. (2020) demonstrated that pathological personality traits are saturated with global personality pathology that is not trait-specific, which may be why it is difficult to find incremental validity with measures of Criterion A. How to best address this overlap (e.g., only use Criterion B, replace Criterion B with normal range trait measures, or tolerate conceptual redundancy within a model) is an ongoing debate in the field.
For Criterion A, some measures may result in stronger criterion/incremental validity compared with other measures. When this occurs, it may reflect differences in the measures (e.g., better psychometrics) and/or it may reflect concerns about the clarity of the underlying construct. The choice of measurement tools can be based on a variety of factors (e.g., lab preferences, measure length, popularity of a measure, and translations available), which may lead to a situation where researchers fail to find incremental validity for Criterion A due to the measure that was used (Type II error). Alternatively, a research team may report incremental validity for the construct of Criterion A due to the measure used, when that finding may not replicate across most other Criterion A measures (Type I error). Studies designed with multiple measures of Criterion A are well positioned to examine these discrepancies and facilitate the synthesis of accumulating research on the validity of Criterion A as a construct. Next, we review the research to date on the validity and incremental validity of several Criterion A measures, organized by each measure, as well as the few studies that measure multiple Criterion A measures together.
LPFS–Self Report (LPFS-SR)
The LPFS-SR is an 80-item self-report measure that is strongly correlated with other global measures of personality impairment (Morey, 2017). The LPFS-SR is associated with interpersonal problems and personality disorders, and is also strongly correlated with normal range personality traits (rs ranging around .5) and pathological personality traits (rs = .6–.7; Hopwood et al., 2018; see also Hemmati et al., 2020). Higher LPFS-SR at baseline was significantly associated with greater daily experiences of negative affect, stress, and invalidation (Heiland et al., in press). Sleep and colleagues (2020) examined incremental validity for the LPFS-SR, finding that the LPFS-SR was able to increment normal range and pathological personality traits (PID-5) for the outcomes of personality disorder. The additional variance explained by Criterion A ranged from 3% to 13%, although the authors of this study viewed this additional variance as small and argued that it was evidence against meaningful incremental validity of Criterion A. In an ecological momentary assessment study, the LPFS-SR was associated with higher negative affect, lower positive affect, cold/submissive behavior, and perceptions of coldness in others, along with variability in most of these variables. After including the PID-5 dimensions, the LPFS-SR remained significant for lower dominant behavior and higher variability in dominant behavior, higher negative affect, and lower positive affect with higher variability in positive affect (Ringwald et al., 2021).
LPFS–Self-Report of Criterion A (LPFS-SRA)
The LPFS-SRA is a 12-item self-report measure that uses the verbatim language of the LPFS in the DSM-5 and asks participants to rate themselves. The LPFS-SRA has been related to interpersonal problems (Dowgwillo et al., 2018) and pathological personality traits (rs = .2–.5), along with identity diffusion, defensive styles, and attachment insecurity (Roche et al., 2018). In a 14-day diary study, the LPFS-SRA evidenced small (3%–5% additional variance) incremental validity over the PID-5 in explaining variance in daily ratings of self and interpersonal dysfunction (Roche et al., 2016). In a different 14-day diary study measuring Criterion A (LPFS) and Criterion B (PID-5-BF) daily, there was evidence of incremental validity (5%–11%) for Criterion A across daily outcomes of self-functioning, emotionality, thinking, awareness/regulation, and interpersonal domains (Roche, 2018). In particular, Criterion A was associated with all of these outcomes reflecting broad dysfunction, whereas the PID-5-BF domains were related to specific impairments (e.g., disinhibition and daily thinking difficulties, detachment/antagonism, and daily interpersonal problems), which reflects the goal of the AMPD to capture global impairment (Criterion A) and more specific manifestations of dysfunction (Criterion B).
LPFS–Brief Form 2 (LPFS-BF2)
The LPFS-BF2 is a 12-item self-report measure. The LPFS-BF2 was positively associated with a number of personality disorders, pathological personality styles (using the severity indices of personality problems), and global psychopathology symptoms (brief symptom inventory total score; Weekers et al., 2019). The LPFS-BF2 was also positively associated with measures of well-being, global symptom distress, and dysfunctional schema modes, and there was evidence that the LPFS-BF2 incremented the PID total score for well-being (.08 additional variance), symptom distress (.07 additional variance), and dysfunctional schemas (ranging from .03 to 0.10 additional variance when a given schema outcome was significantly incremented by Criterion A relative to Criterion B; Bach & Hutsebaut, 2018).
Self and Interpersonal Functioning Scale (SIFS)
The SIFS (Gamache et al., 2019) is a relatively new 24-item self-report measure. In this original study, the SIFS was related to low self-esteem, low satisfaction with life, lower personality organization, low empathy, pathological narcissism, anger, poorer mental and general health, and was correlated strongly (rs = .49–.81) with the PID-5-BF. In community (n = 280) and clinical (n = 194) samples, incremental validity (over the PID-5) ranged from 0.02 to 0.16 (M = 0.08), with larger increments found for low perspective-taking (0.16), low satisfaction with life (0.10), low self-esteem (0.09), and motor impulsivity (0.09), for the SIFS (Gamache & Savard, 2019).
DSM-5 Levels of Personality Functioning Questionnaire (DLOPFQ)/DLOP-SF
The DLOPFQ is a 132-item measure (Huprich et al., 2018), with a shorter 24-item form (DLOP-SF; Siefert et al., 2020). The DLOPFQ was significantly related to insecure attachment, dependency, and lower well-being, correlated around 0.6 with pathological personality traits, but found incremental validity (additional variance explained ranging from 0.03 to 0.14) compared with the PID-5 in its association with attachment, dependency, and well-being (Huprich et al., 2018). The shorter form was associated with interpersonal problems, lower well-being, and insecure attachment style (Siefert et al., 2020).
Informant Measures
The original informant/clinician-rated measure is the LPFS, which has been used by clinicians with higher scores relating to greater PD symptoms, lower social functioning, risk composite, and level of care rating (Morey et al., 2013). Lay raters have also been able to use this measure with adequate inter-rater reliabilities (average rater intraclass correlation coefficient [ICC] = .7–.9), and it similarly relates to PD diagnosis, Axis I diagnoses, and psychodynamic and interpersonal variables (Roche et al., 2018; Zimmermann et al., 2014). Regarding incremental validity, Few and colleagues (2013) examined clinician-rated measures of the LPFS and pathological traits, finding that Criterion A did not increment Criterion B in associations with personality disorders.
Interviews have also been designed to specifically address the Criterion A construct. The semi-structured interview for personality functioning scale of DSM-5 (STiP-5.1; Hutsebaut et al., 2014) is a 45-minute interview, with adequate average inter-rater reliability (.81–.92), and is related to personality disorders, pathological personality traits, global psychopathology symptoms as assessed by the brief symptom inventory total score, and severity of personality problems (Berghuis et al., 2013; Hutsebaut et al., 2017). Zettl and colleagues (2019) found positive associations with personality disorder symptoms and that scores differentiated whether a patient was diagnosed with a PD or not. The STiP was correlated with pathological personality traits (PID-5) in around the 0.2 range except for disinhibition (r = .45; Weekers et al., 2021).
The Structured Interview for Personality Disorders Module I (SCID-AMPD) is a 45- to 90-minute interview (Bender et al., 2018), with good inter-rater reliability (average inter-rater reliability = .93) and associations with number of suicide attempts, psychiatric hospitalizations, and severity of diagnoses (Kampe et al., 2018). The interview effectively distinguishes between patients diagnosed with a PD or not (Buer Christensen et al., 2020). We are unaware of any research examining the incremental validity of these Criterion A interviews over and above Criterion B traits to predict outcomes.
Comparing Criterion A Measures
McCabe et al. (2021) noted that no previous research had compared self-report measures of Criterion A, and then they examined subscales of the LPFS-BF2, DLOPFQ, and LPFS-SR, finding these subscales shared substantial overlap (rs > .5) and were also strongly associated with the five-factor model. The Criterion A subscales were more strongly related within a measure than between different measures purportedly measuring the same construct (e.g., empathy), which they evaluated as poor discriminate validity at the subscale level. A few studies have examined cross-method comparisons of Criterion A. Specifically, the STiP was positively related to the LPFS-BF2 (r = 0.49; Weekers et al., 2021) and the LPFS-SRA was positively related to the LPFS–observer-rated (LPFS-OR) measure (r = .27; Roche et al., 2018). This study also published subscale-level data, and consistent with McCabe and colleagues (2021), subscales were not aligning well across method, evidencing poor discriminant validity.
A comparative content analysis for self-reported Criterion A measures (e.g., LPFS-SR, DLOP, LPFS-BF, and SIFS) found that graduate student and PhD-level raters viewed the content of these measures as very similar (associations higher than .8, many higher than .9), especially for the content measuring identity and intimacy difficulties (Waugh et al., 2021).
These raters were also asked to rate each Criterion A measure for content reflected in the Criterion B domains, thus comparing content overlap of Criterion A and Criterion B. They found that many Criterion A instruments appeared to contain more overlapping content for negative affectivity and detachment, relative to the other Criterion B domains. The authors reported the total amount of Criterion A versus Criterion B construct representation (higher scores indicating less Criterion B overlap). They reported effect sizes for LPFS-SR (d = 0.94), LPFS-BF2 (d = 0.84), SIFS (d = 0.63), and DLOP (d = 1.15). These results may suggest that, based on content overlap, the domains of Negative Affect (NA) and Detachment (DET) will be strongly associated with Criterion A measures, and that the DLOP may be more distinct, and the SIFS more overlapping, with the PID-5-BF. These findings are also largely consistent with Zimmermann and colleagues (2020) who found that several self-report Criterion A measures (e.g., LPFS-SR, LPFS-BF) and the average of PID-5-BF domains submitted to factor analysis form a single severity factor, indicating potential redundancy between Criterion A and Criterion B.
In summary, very few measures have compared Criterion A instruments, and even fewer have compared how Criterion A measures perform relative to each other in predicting important outcomes within the same sample.
Summary
Several self-report and interview measures are available to assess for Criterion A. These measures are often evaluated by examining reliability, and demonstrate criterion validity by examining associations with global impairment (such as a total score from a broad measure of psychopathology) and distinguishing level of severity or belonging to a diagnostic group or severity indicator (e.g., history of hospitalizations, number of suicide attempts).
Incremental validity is also examined to determine how distinct Criteria A and B are as constructs. The majority of this literature uses mono-method self-report frameworks, and inconsistently finds small but significant (less than 0.10) additional variance explained by Criterion A measures. Some measures have performed better (DLOPFQ) and using outcomes related to global impairments, or temporally dynamic data, tend to increase incremental validity coefficients.
The Present Research
This research compares five self-report measures of Criterion A to examine their reliability, criterion validity, overlap, and incremental validity with a Criterion B measure. More specifically, we hypothesized that all Criterion A measures will be sufficiently reliable, and that all Criterion A measures will be positively associated with each other. We chose outcome variables that were related to mental health utilization (e.g., history of psychiatric and psychotherapy utilization) as a proxy for levels of severity used in other studies. We also included a measure of interpersonal problems to capture interpersonal dysfunction. We employed a general symptom inventory appropriate for the college-level sample, meant to represent general dysfunction, and also included so-called “critical items” used by that measure to mark suicidality and aggression.
We assessed Criterion B through the PID-5-BF and hypothesized substantial (rs = .5–.6) associations with Criterion A measures, as has been found in previous research. We further hypothesize that measures of Criterion A will increment Criterion B for the aforementioned outcomes. We did not have a priori hypotheses about which Criterion A measures would be more strongly associated with outcomes, but we report these findings to help elucidate the similarities and differences among these measures of the same purported construct.
We also included a new interview-based measure of Criterion A and similarly evaluate its reliability, criterion validity, association with self-reported Criterion A measures, overlap with Criterion B, and incremental validity. We present these results separately because all of the other variables in the study are self-report and more easily comparable. For instance, one would expect a lower association with outcome measures due to unshared method variance and, if the interview-based Criterion A measure does not increment, it may be due to a variety of reasons that are difficult to disentangle (e.g., the measure is weak, the construct is redundant, and/or the self-reported outcomes may not have been significant with any interview-rated scale).
Method
Participants and Procedure
We collected data for this project across two semesters. In the spring of 2019, participants completed self-report surveys online in exchange for course credit (n = 108). In the fall of 2019, participants (n = 96) were interviewed by undergraduate research assistants (RAs) using the brief personality interview (see the appendix) and then were asked to complete self-report surveys on a lab computer. The interviews were audiotaped so that multiple RAs could later code the interview using the LPFS (APA, 2013). The research was institutional review board (IRB)-approved and followed appropriate ethical guidelines.
The total combined sample was 204 participants. Participant’s age ranged from 18 to 24 years (M = 18.90, SD = 1.47) and predominately identified as heterosexual (89%), non-Hispanic (93%), single (43%), and female (69%). The sample largely identified as Caucasian (74%), with smaller portions identifying as Asian (13%), African American (9%), Other (4%), American Indian or Alaskan Native (2%), and Native Hawaiian or Pacific Islander (<1%). The brief personality interview was only collected in the fall of 2019, and three participant interviews were not coded due to an oversight, resulting in a sample size of 93 for that variable. The SIFS and DLOP-SF were also only collected in fall 2019 (the authors became aware of these measures after spring 2019), and due to some missing data, the sample sizes were 91 (DLOP-SF) and 92 (SIFS).
Measures
Self-Report Criterion A Measures
The five self-report surveys of Criterion A used for this study all assess for personality dysfunction and all can be calculated with an average score representing the single dimension of Criterion A. The LPFS self-report (LPFS-SR; Morey, 2017) is an 80-item self-report questionnaire employing a 4-point Likert-type scale ranging from 1 (totally false) to 4 (very true). The LPFS-SR is scored using unique weights for items that represent more severe personality dysfunction content. The LPFS-SRA (APA, 2013; Roche et al., 2016) is a 12-item self-report questionnaire created by taking the LPFS verbatim from the DSM-5 and asking participants to use the coding scheme about themselves. As such, a 5-point scale ranging from 0 to 4 is used (with unique anchor points for each number across each item). The LPFS-BF2 (Bach & Hutsebaut, 2018) is a 12-item self-report measure employing a 4-point Likert-type scale ranging from 0 (very false or often false) to 3 (very true or often true). The SIFS (Gamache et al., 2019) is a 24-item self-report questionnaire employing a 5-point Likert-type scale ranging from 0 (not at all) to 4 (totally). The DSM-5 Level of Personality Functioning Questionnaire short-form (DLOP-SF; Siefert et al., 2020) is a 24-item self-report questionnaire employing a 6-point Likert-type scale from 1 (strongly disagree) to 6 (strongly agree).
LPFS–Interview-Rated
The Brief Personality Interview was developed specifically for this study to target difficulties with identity, self-direction, empathy, and intimacy (see the appendix). This roughly 20-minute interview was coded for Criterion A personality dysfunction using the LPFS (APA, 2013). The LPFS is a 12-item coding instrument, with a 5-point scale ranging from 0 (healthy) to 4 (impaired). This instrument captures elements of personality dysfunction described in the DSM-5. Once inter-rater reliability was established, we calculated the average score across the raters for each of the 12 LPFS items. Then, we averaged these items to comprise the four subscales of identity (ID), self-direction (SD), empathy (EPY), and intimacy (INT) along with the average score of all subscales (LPFSavg). These scores are denoted as LPFS-IR. Descriptive and reliability information for all Criterion A measures is listed in Table 7, and reviewed in the “Results” section.
Criterion B
The PID-5-BF (APA, 2013) is a 25-item instrument that was derived from its parent measure, the 220-item PID-5 (Krueger et al., 2012). Items are rated on a 4-point scale, ranging from 0 (very false) to 3 (very true), and is scored to represent the five pathological trait domains of Negative Affect (NA, α = .74), Detachment (DET, α = .64), Antagonism (ANT, α = .74), Disinhibition (DIS, α = .69), and Psychoticism (PSY, α = .79).
Mental Health Utilization
Participants were asked items from the 2016 National Survey on Drug Use and Health survey. Items assessed for history of inpatient treatment (Have you ever stayed overnight or longer in a hospital or other facility to receive treatment or counseling for any problem you were having with your emotions, nerves, alcohol/drug addiction, or mental health?), outpatient treatment (Have you ever received any outpatient treatment or counseling for any problem you were having with your emotions, nerves, alcohol/drug addiction, or mental health at any of the places listed below?), and psychiatric medication history (Have you ever taken any prescription medication that was prescribed for you to treat a mental or emotional condition?). The rate of any mental health utilization for this sample was 35.3%, which included inpatient treatment (6.4%), psychiatric medication (17.6%), and outpatient treatment (30.4%). The percentage of those endorsing inpatient treatment without endorsing outpatient treatment was 44%, while the percentage of those endorsing psychiatric medication without endorsing outpatient treatment was 20%.
Interpersonal Problems
Interpersonal problems were measured using the Inventory of Interpersonal Problems–Short Circumplex (IIP-SC; Soldz et al., 1995). The IIP-SC measures interpersonal problems that an individual does too much, or has difficulty doing (e.g., I do . . . too much, It is hard for me to). Items from the IIP-SC span themes of (low/high) agentic and (low/high) communal interpersonal problems. The 32-item measure is rated using a 5-point Likert-type scale ranging from 0 (not at all) to 4 (extremely). We used the total (elevation) score, calculated by averaging the subscales of the IIP-SC once they were z scored using college student sample norms (Hopwood et al., 2008; α = .84, M = −0.08, SD = 0.69, skewness = 0.51, kurtosis = −0.11).
Psychological Distress
The Counseling Center Assessment of Psychological Symptoms (CCAPS) 34 was released in September 2009 and updated in 2012 (Locke et al., 2012). It is a 34-item instrument on a 5-point Likert-type scale, ranging from 0 (not at all like me) to 4 (extremely like me). Participants endorse symptoms experienced over the past 2 weeks, which organizes into the subscales of Depression, Generalized Anxiety, Social Anxiety, Academic Distress, Eating Concerns, Hostility, and Alcohol Use. For this study, we use the distress index, which includes items from each subscale to capture a more general aspect of psychological distress (α = .93, M = 1.19, SD = 0.81, skewness = 0.75, and kurtosis = 0.14). We also examined three individual “critical items” that the CCAPS-34 identifies as important clinical markers. This includes suicidal ideation (I have thoughts of ending my life, M = 0.32, SD = 0.59, skewness = 1.53, and kurtosis = 0.80), aggressive ideation 1 (I have thoughts of hurting others, M = 0.16, SD = 0.43, skewness = 2.54, and kurtosis = 5.29), and aggressive non-control (I’m afraid I may lose control and act violently, M = 0.25, SD = 0.53, skewness = 1.88, and kurtosis = 2.24). These three items were square root transformed to reach acceptable levels of skewness (<2) and kurtosis (<7; West et al., 1995), and descriptive information was reported after those transformations.
Data Analysis
Convergent and Criterion Validity
For convergent validity, we calculated the correlations between Criterion A measures. For criterion validity, we calculated the correlations between each Criterion A measure and the outcome variables. To examine whether certain Criterion A measures were more strongly associated with a given outcome, we compared correlation magnitude strengths using Steiger’s Z test for dependent correlations (Steiger, 1980). This method tests the difference between two dependent correlations that share one variable in common. We used online software to calculate Steiger’s Z (Lee & Preacher, 2013). Steiger’s Z is sensitive to sample size, so we used the lowest n for each comparison to be more conservative. For most comparisons, this was n = 198; however, any comparison using either the SIFS or DLOP-SF was computed with n = 88 due to these measures only being collected in fall 2019, resulting in a smaller sample size.
Associations and Incremental Validity With Criterion B Measures
We calculated correlations between each Criterion A measure and the Criterion B domains, comparing magnitude strengths using Steiger’s Z. We also entered the five PID-5-BF domains into a regression to examine the amount of variance the Criterion B dimensions were able to explain for each Criterion A measure.
For incremental validity, we first report a series of multiple regressions, where each outcome was regressed on the PID-5-BF domains, to report the relationship between Criterion B and the outcomes. Then, we used a series of hierarchical linear regressions for each outcome model. We first entered all five PID-5-BF domains into Step 1 and then added each Criterion A measure separately into the regression to examine the amount of additional variance explained by a Criterion A measure once Criterion B domains were accounted for. We then reversed this process by adding in each Criterion A measure separately into Step 1, followed by all five PID-5-BF domains in Step 2. We then evaluated the relative additional variance explained in Step 2 of Criterion A versus Criterion B.
We replicated these incremental validity analyses while using correlations corrected for unreliability (see Supplemental Table 1). We also replicated these incremental validity analyses using the Predicted Residual Sum of Squares (PRESS) effect size estimates that are free from over-fitting models with more predictors (see Supplemental Table 2).
Analysis of Interview-Based Criterion A Measure
In the spring of 2020, eight undergraduate RAs were given the deidentified audio files for each participant interview and were asked to rate the interviews using the LPFS (APA, 2013). Prior to this, RAs were given introductory information about the LPFS and a journal reading on the LPFS was completed. Due to the impact of coronavirus disease 2019 (COVID-19), several RAs were unable to complete all of the ratings for the 96 participants. Specifically, five RAs rated at least 92 of participant interviews. The remaining three RAs completed 65, 43, and 42 interviews, respectively (essentially stopping once the university ended in-person learning). Reliability scores cannot be calculated with incomplete data, and the correlation between the LPFS average score of eight raters versus five raters was 0.97. Given this, we retained the ratings for the five RAs who completed most of the ratings, and we use these five RAs ratings moving forward. Inter-rater reliability was evaluated using a two-way, random ICC calculations, to quantify the reliability of a single rater (ICC [2,1]) and reliability across the mean of the five raters (ICC [2,5]).
Next, we evaluated the LPFS-IR in terms of convergent validity (correlations with other Criterion A measures), criterion validity (correlations with the outcomes), overlap with Criterion B domains (correlations), and incremental validity (regression).
Results
Descriptive Information and Intercorrelations Among Criterion A Measures
Descriptive information for all Criterion A self-report measures is presented in Table 1. Across measures, the mean level for Criterion A is low compared with the range of the scales, consistent with prior research using Criterion A measures in a sample of college students (Roche et al., 2018). The reliability across measures is adequate for the average Criterion A score. Reliability can increase with a larger number of items per scale, and the Criterion A measures vary from 12 to 80 items. Therefore, we applied the Spearman–Brown correction and set the number of items at 80, resulting in all scales having reliabilities above .94. Associations among Criterion A measures are presented in Table 2. Criterion A measures were strongly correlated (mean r = .68, range = .54–.79).
Descriptive Information for Criterion A Measures.
Note. LPFS-SR theoretical range is based on unique weighting for individual items. The LPFS-BF2 has a two-factor structure, so reliability for the subscales is reported above. LPFS-SRA = Level of Personality Functioning Scale–Self-Report of Criterion A; LPFS-SR = Level of Personality Functioning Scale–Self-Report; LPFS-BF2 = LPFS–Brief Form 2; SIFS = Self and Interpersonal Functioning Scale; DLOP-SF = DSM-5 Levels of Personality–Short Form; SD = self-direction; ID = identity; EPY = empathy; INT = intimacy.
Interrelationships Among Criterion A Measures.
Note. LPFS-SRA = Level of Personality Functioning Scale–Self-Report of Criterion A; LPFS-SR = Level of Personality Functioning Scale–Self-Report; LPFS-BF2 = LPFS–Brief Form 2; SIFS = Self and Interpersonal Functioning Scale; DLOP-SF = DSM-5 Levels of Personality–Short Form.
p < .05.
Criterion A Measures and Outcomes
The correlation between Criterion A measures and outcomes are presented in Table 3. Consistent with expectations, all Criterion A measures were positively associated with history of outpatient therapy, history of medication, psychological distress, interpersonal problems, suicidal ideation, and aggressive non-control, and most were also associated with aggressive ideation. No Criterion A measures were associated with history of inpatient treatment.
Associations Among Criterion A Measures and Outcome Variables.
Note. Superscripts denote significant differences among Criterion A measures for a given outcome variable, calculated using Steiger’s Z test for dependent correlations. Point biserial correlation values are reported for the inpatient, outpatient, and medication variables, which are dichotomous. LPFS-SRA = Level of Personality Functioning Scale–Self-Report of Criterion A; LPFS-SR = Level of Personality Functioning Scale–Self-Report; LPFS-BF2 = Level of Personality Functioning Scale–Brief Form 2; SIFS = Self and Interpersonal Functioning Scale; DLOP-SF = DSM-5 Levels of Personality–Short Form; CCAPS = Counseling Center Assessment of Psychological Symptoms.
p < .05.
Specific pair-wise correlation magnitude comparisons are noted in the superscripts of Table 3. For outpatient therapy, the SIFS and DLOP-SF were stronger in magnitude compared with the LPFS-BF2. For interpersonal problems, the LPFS-SR was of a stronger magnitude compared with all others. For psychological distress, the LPFS-SR, LPFS-BF2, and SIFS were stronger in magnitude compared with the LPFS-SRA. For suicidal ideation, the LPFS-SR, LPFS-BF2, and SIFS were stronger in magnitude compared with the DLOP-SF. For aggressive ideation, the LPFS-SRA, LPFS-SR, and LPFS-BF2 were stronger in magnitude compared with the DLOP-SF. For aggressive non-control, the LPFS-BF2 was stronger in magnitude compared with the LPFS-SRA, and the LPFS-BF2, LPFS-SR, and SIFS were stronger in magnitude compared with the DLOP-SF.
Evaluating Overlap Among Criterion A and Criterion B of the AMPD
The correlation between Criterion A measures and the Criterion B measure are presented in Table 4. The average association between Criterion A and Criterion B was 0.51. Correcting for attenuation resulted in a higher average association (0.65), with all associations for detachment in the .8–.9 range and the SIFS associated with negative affectivity at .85, and detachment at .96. We will next highlight significant magnitude differences.
Associations Among Criterion A Measures and Criterion B Measure.
Note. The r values are reported in the correlation section and standardized betas are reported in the regression section. Superscripts denote significant differences among Criterion A measures for a given PID-5-BF variable, calculated using Steiger’s Z test for dependent correlations. Regression analyses included Criterion A measures (columns) that were regressed onto the five PID-5-BF scales (rows). LPFS-SRA = Level of Personality Functioning Scale–Self-Report of Criterion A; LPFS-SR = Level of Personality Functioning Scale–Self-Report; LPFS-BF2 = Level of Personality Functioning Scale–Brief Form 2; SIFS = Self and Interpersonal Functioning Scale; DLOP-SF = DSM-5 Levels of Personality–Short Form; NA = Negative Affectivity; ET = Detachment; ANT = Antagonism; DIS = Disinhibition; PSY= Psychoticism.
p < .05.
For NA, the SIFS had a stronger magnitude compared with the LPFS-SRA, LPFS-SR, and LPFS-BF2. For ANT, the LPFS-SRA and LPFS-SR had a stronger magnitude compared with the LPFS-BF2 and DLOP-SF. For DIS, all other measures were stronger in magnitude compared with the DLOP-SF, and, for Psychoticism, the LPFS-SR was stronger in magnitude compared with all other measures.
Table 4 also included multiple regressions to determine the amount of variance shared among Criterion A and Criterion B measures. The self-report measures of Criterion A had substantial overlap with Criterion B (range = 0.50–0.64). The consistent associations were for NA and DET. After this, specific measures appeared to be more related to certain PID-5-BF domains (e.g., LPFS-SRA and ANT, LPFS-SR and PSY, and SIFS and DIS).
Criterion B and Outcome Variables
Before examining incremental validity, we first demonstrate that Criterion B domains are related to the outcomes (see Table 5). Specifically, there were strong associations with the outcomes of interpersonal problems and psychological distress, and moderate associations with the other variables (except history of inpatient treatment).
Multiple Regression Associations for PID-5-BF Variables and Outcome Variables.
Note. Coefficients are reported as odds ratios for binary outcomes and standardized beta coefficients for dimensional outcomes. Adj. r2 = Nagelkerke’s R2 for mental health utilization outcomes that were modeled using a binary logistic model. PID-5-BF = Personality Inventory for DSM-5–Brief Form; NA = Negative Affectivity; ET = Detachment; ANT = Antagonism; DIS = Disinhibition; PSY = Psychoticism; CCAPS = Counseling Center Assessment of Psychological Symptoms.
p < .05.
Incremental Validity of Criterion A Measures
For each outcome, the PID-5-BF was entered in Step 1, and then each Criterion A measure was entered into Step 2 in separate models to examine whether a Criterion A measure explains additional variance in the outcomes. Steps 1 and 2 were then reversed. These models are presented in Table 6. In all cases, when a Criterion A measure was significant, it was in the expected positive direction.
Incremental Validity of Criterion A and Criterion B for Outcomes.
Note. Criterion B is entered into a step using all five PID-5-BF scales together. Criterion A scales entered into a step separately for each outcome model. Binary logistic regression used for inpatient, outpatient, and medication variables, which are dichotomous. For these variables, we report Nagelkerke’s R2 for the Step 1 and Step 2 variance. For all other variables, we report the adjusted R2. Inp = ever had inpatient therapy; Out = ever had outpatient therapy; Med = ever had psychiatric medication; Interp prob = interpersonal problems; CCAPS = Counseling Center Assessment of Psychological Symptoms; PID-5-BF = Personality Inventory for DSM-5–Brief Form; LPFS-BF2 = LPFS–Brief Form 2; Sui ide = suicidal ideation critical item; Agg ide = aggressive ideation critical item; Agg non-control = aggressive non-control; LPFS-SRA = Level of Personality Functioning Scale–Self-Report of Criterion A; LPFS-SR = Level of Personality Functioning Scale–Self-Report; SIFS = Self and Interpersonal Functioning Scale; DLOP-SF = DSM-5 Levels of Personality–Short Form.
p < .05.
We reported the adjusted r2 for the total models (Steps 1 and 2) to document the amount of variance explained by Criteria A and B together. 2 We focus first on the additional variance explained by Criterion A (second set of rows in Table 6). Looking across the outcomes (columns), consistent support for incremental validity was found for psychological distress and interpersonal problems, with less consistent support for suicidal ideation. There was infrequent support for incremental associations of mental health utilization (DLOP-SF was the only measure that incremented Criterion B) and aggressive non-control (LPFS-BF2 was the only measure that incremented Criterion B).
In total, there were 40 possible incremental validity findings (8 outcomes × 5 measures). Criterion A incremented Criterion B 15 times (39% of findings). When significant, the average additional variance explained was 5.47 (the average additional variance from all 40 models, including nonsignificant findings, was 2.83). In contrast (looking at the third set of rows in Table 6), Criterion B incremented Criterion A 29 times (73% of findings). When significant, the average additional variance explained was 14.07 (the average from all 40 models was 11.83). It was rare to find incremental validity only for Criterion A (2 times), while finding incremental validity only for Criterion B was more common (17), and when both Criteria A and B evidenced incremental validity, in every case Criterion B had a stronger (11) or equal (1) magnitude compared with Criterion B. Overall, Criterion B evidenced incremental validity about twice as often (15 vs. 29), and the average magnitude of a significant finding was about 2.5 times as large for Criterion B (5.47 vs. 14.07).
We recalculated incremental validity in two other ways. First, we reran all analyses using the correlation matrix corrected for attenuation due to the unreliability of measures (see Supplemental Table 1). This was done because Criterion B measures had weaker reliability compared with Criterion A measures, which may have presented more favorable incremental validity score for Criterion A. Several models did not converge with Criterion A entered second due to multicollinearity with DET, and when PID-5-BF variables were entered second, DET could not fit in the model, so was excluded. This perhaps led to an underestimation of incremental validity for those outcomes as it evaluated the incremental validity of the PID-5-BF without DET, and DET was significantly associated with several outcomes (see Table 5). Standard errors of these estimates are not reliable, precluding significance testing.
We also recalculated incremental validity using the PRESS statistic, which derives R2 values that are free from over-fitting models with more predictors. This was done because Criterion B measures had a higher number of scales compared with Criterion A, which may have presented a more favorable incremental validity score for Criterion B due to model overfitting. Full details for calculating the PRESS R2 statistic are found in Supplemental Table 2, and results were only reported for dimensional outcomes as it was unclear how to implement this statistic using dichotomous outcomes.
Incremental validity for Criterion A measures on average was 2.83% (including nonsignificant findings) using the adjusted R2 approach (reported already). Incremental validity values were lower in analyses that accounted for unreliability (1.79%) or used the PRESS R2 approach (2.16%). Incremental validity for Criterion B measures on average was 11.83% (including nonsignificant findings). Incremental validity values were higher in analyses that accounted for unreliability (14.98%), but lower in analyses that used the PRESS R2 approach (6.64%).
Inter-Rater Reliability of Criterion A Interview
The inter-rater reliability results for the LPFS-IR are presented in Table 7. Considering the average LPFS score, reliability of individual raters was moderate (.52), while the reliability averaged across five raters was good (.84). These magnitudes are comparable to Roche and colleagues (2018) who also found inter-rater reliability around .4–.5 for individual undergraduate raters and around .8 for the average across five undergraduate raters. Examining the reliability of subscales across raters, the identity subscale had the weakest reliability (.65), while the other subscales demonstrated adequate to good reliability (.75–.90).
Inter-Rater Reliability for LPFS–Interview-Rated (LPFS-IR) Measure.
Note. Average rater based on five raters. ICC = intraclass correlation coefficient; LPFS-IR = Level of Personality Functioning Scale–Interview-Rated; SD = self-direction; Skew = skewness; Kurt = kurtosis; alpha = Cronbach’s alpha (internal consistency of items comprising the scales); ID = identity; EPY = empathy; INT = intimacy.
The LPFS-IR evidenced small to medium correlations with other measures of Criterion A, such as the LPFS-SRA (r = .39, p <.001), LPFS-SR (r = .22, p = .037), LPFS-BF2 (r = .31, p = .003), SIFS (r = .38, p < .001), and DLOP-SF (r = .28, p = .009). The LPFS-SRA correlation was slightly larger compared with prior research (.27) using a different stimulus for LPFS coding (Roche et al., 2018). There was also small overlap with the Criterion B dimensions of NA (r = .29, p = .005), DET (r = .25, p = .018), ANT (r = .25, p = .018), DIS (r = .11, p = .307), and PSY (r = .27, p = 009). This magnitude is consistent with prior research using a different interview to correlate with the PID (Weekers et al., 2021). The LPFS-IR was significantly associated with history of outpatient treatment (r = .35, p < .001), interpersonal problems (r = .27, p = .013), and CCAPS distress (r = .30, p = .004), but not significantly associated with history of inpatient or medication, or the critical items of the CCAPS.
For incremental validity, the LPFS-IR only incremented the PID-5-BF traits in the outcome of history of outpatient therapy. Specifically, when entered second, the LPFS-IR explained an additional 7% of variance (compared with the PID-5-BF entered second, which explained 19% additional variance). However, when entered together, the LPFS-IR was the strongest association (p = .02) and none of the PID-5-BF dimensions remained significant at p values of < .05.
Discussion
This research examined how the construct/measures of Criterion A were related to each other and key outcomes indicative of psychopathology. We first focus on the broader picture of Criterion A as a construct, and then compare Criterion A measures.
The Construct of Criterion A
First, measures of Criterion A appear to be reliable, and intercorrelations suggest that most self-report measures capture a similar construct. Criterion A measures are significantly and consistently related to markers of mental health utilization, interpersonal problems, psychological distress, and more specific clinical concerns, such as suicidal ideation and markers of aggression.
Criterion A also shares significant variance with Criterion B (regression-based variance explained, ranging from 0.50 to 0.64 across self-report measures, many zero-order correlations in excess of 0.5), which increases when correcting for the low reliability of Criterion B measurement (correction for attenuation). Despite this, most Criterion A measures significantly incremented interpersonal problems and psychological distress, and over half incremented suicidal ideation. Thus, regardless of the measure chosen to capture the Criterion A construct, it seems that small but significant incremental validity is found for these key outcomes. Less consistent incremental validity was found for treatment utilization and aggressive non-control, which may be due to the strengths of a given Criterion A measure and should be replicated in future studies to increase confidence in the link between those outcomes and the construct of Criterion A.
These findings are in contrast to other studies that tend to find Criterion A does not increment when the outcome is defined narrowly as personality disorder diagnostic criteria. We agree with Morey (2019) that research on incremental validity should move beyond traditional PD constructs as outcomes. Interestingly, Sleep and colleagues (2020) also found incremental validity in Criterion A at around 6% in their study and concluded that this level of incremental validity was low. There does not seem to be a clear consensus in the field for what percentage is necessary to be meaningful and likely depends in part on the outcome assessed. Consistent with prior research, Criterion B is more frequently and more powerfully incrementing Criterion A than the reverse.
The PID-5-BF domains of NA and DET were the strongest and most consistent associations with the Criterion A measures (see Table 4), and these were also the domains most consistently associated with the outcomes (see Table 5). This is in line with other research that found NA is broadly associated with many negative outcomes and does not have the discriminative nuance that the other pathological trait domains possess (Roche, 2018). Indeed, Saulsman and Page (2004) also noted neuroticism to be at the core of many personality disorders. Thus, it may be that the construct of negative affectivity is appropriately conceptualized as a core PD impairment (Criterion A), even if it is also employed as a stylistic difference in how personality dysfunction is expressed (Criterion B). Detachment was even more strongly associated with Criterion A (adjusted for reliability, the associations were higher than .8) and shares a clear conceptual overlap with intimacy impairment contained within the Criterion A construct.
Conceptual overlap across PD criteria is not necessarily problematic and, in fact, is quite common in other assessment literatures such as cognitive assessment. For instance, on the WISC-V, the visual spatial index (VSI) and fluid reasoning index (FRI) both rely upon visual-spatial processing, have similar looking tasks to derive index scores, and are correlated. Yet assessment psychologists readily recognize the nuances in these scores and can make meaningful distinctions when discrepancies arise. Similarly, negative emotion content is found in both Criteria A and B, and it may be more fruitful to consider how these similarities/discrepancies in scores fit together, rather than viewing overlap as a problem to be eliminated.
Diagnostically, Criteria A and B function in different roles (Criterion A captures severity and commonality; criteria B captures individual differences in expression). But they also connect to different theoretical frameworks, as Criterion A emphasizes personological and psychodynamic themes, whereas Criterion B emphasizes multivariate and trait themes (Mulay et al., 2018). Thus, the decision for retaining Criteria A and B should also account for this added clinical utility in having the AMPD represent many different theoretical lenses with which to see a client.
Comparing Criterion A Measures
All Criterion A measures (self-report) were significantly correlated with one another and had similar (within .15) magnitudes of association with the outcomes. One exception is the DLOP-SF, which had weaker associations with the CCAPS critical items, compared with most other Criterion A measures.
The overlap between Criterion A and Criterion B was substantial, with NA and DET being the main associations. This may be partially explained by content overlap as a comparative content analysis study (Waugh et al., 2021) found that content overlap was strongest for NA and DET in all of the Criterion A measures (except for the LFPS-SRA, which was not included in their study). Also consistent with this research, the measure with the most content overlap with Criterion B was noted to be the SIFS, and the least was noted to be the DLOP-SF and indeed, in this study, the strongest regression variance explained was for the SIFS (64%) and the weakest was for the DLOP-SF (49%), suggesting the association is partially explained by shared content overlap.
Some Criterion A measures have weaker associations with certain domains. These associations matter because a study using one Criterion A measure may find a significant association, whereas using a different Criterion A measure may not find that association. For instance, the LPFS-SR had stronger associations with psychoticism, which may make it better positioned to find associations with thought disturbance. The SIFS had stronger associations with negative affectivity, which may make it better positioned to find associations with emotional distress, and the DLOP-SF had weaker associations with disinhibition, which may make it harder to find associations with externalizing outcomes. In each of these cases, researchers would not want to find a result (or miss a result) due to measurement-specific differences in how the construct of Criterion A is assessed. This inconsistency may reflect differences in measures, or it may reflect a lack of criterion validity for the construct of Criterion A and the associated outcomes. Studies employing multiple Criterion A measures may be well positioned to disentangle these different possibilities. We hope our research can be useful to synthesize such results into a better understanding of the nomological net of the Criterion A construct.
Similarly, the incremental validity findings are more robust if they are demonstrated across several measures. This was found for interpersonal problems, psychological distress, and somewhat for suicidal ideation. But there were also measure-specific effects (e.g., treatment utilization incremented by DLOP-SF, aggressive non-control incremented by LPFS-BF2). Researchers should therefore be cautious in claiming that Criterion A fails to increment Criterion B, as we demonstrated in this study that the measure of Criterion A can influence whether a particular outcome will be incremented. However, some outcomes are incremented more consistently than others and, when the incrementation appears measurement specific, we suggest caution in its interpretation and invite the need to replicate the findings across a new sample.
We cannot conclude that one Criterion A measure is superior to the others based on the totality of this research. However, one can consider trade-offs across different measures. The LPFS-SRA is taken verbatim from the DSM-5, making it easy to compare it with interview ratings using the LPFS. However, in this study, it demonstrated a range restriction compared with the other measures, and the items themselves have different scale point descriptions on each item, likely increasing completion time compared with other measures. The LPFS-SR is a comprehensive view of Criterion A, but is the longest measure (80 items) and has a complex scoring routine. The LPFS-BF2 is the briefest measure; however, it uses a two-factor subscale structure (unlike all others with a four-factor structure), so if subscale-level comparisons are desired this measure may be less precise. The SIFS and DLOP-SF are both 24-item measures, with a four-factor structure found in other measures, and represent the middle road in terms of completion length for the surveys.
Interview-Rated Measure of Criterion A
The LPFS-IR was designed to be a free and briefer interview measure (20 minutes vs. 45–90 minutes) to assess for Criterion A personality dysfunction. The LPFS-IR appears to be reliable across five raters (ICC average rater score 0.84), and the items comprising the average score also appear to have adequate reliability (α = .79). However, inter-rater reliability for a single rater is only moderate (ICC single rater = 0.52). Taken together, it suggests the LPFS-IR is sufficiently reliable to use when averaging across five raters, but may not be reliable enough when using a single rater (such as in clinical practice with a single clinician rating a client).
Of note, the reliability (internal consistency) for the LPFS-IR subscale of self-direction is poor (.38), while the ICC across raters was adequate (.75). This suggests that coders can reliably agree on elevations of self-direction problems, but the three items comprising this scale do not hang together well. However, the LPFS-SRA uses the exact same items in a self-report format, and the internal consistency reliability for that subscale is .65. This may further suggest that the poor reliability (internal consistency) from the LPFS-IR may not be the items themselves, but an inadequate set of interview questions to tap the construct of self-direction difficulties in a college sample. These comparisons highlight the usefulness of multi-method data to better understand strengths/limitations in how a construct is measured.
The LPFS-IR was associated with self-reported measures of Criterion A, establishing convergent validity. The weaker correlation magnitudes (especially compared with other Criterion A measures) is likely explained by unshared method variance with the self-reported outcome variables. The LPFS-IR was associated with some Criterion outcomes (outpatient therapy, interpersonal problems, and psychological distress), but was not related to the others, providing less consistent validity support for this measure. The LPFS-IR only incremented the PID-5-BF in explaining additional variance in the history of outpatient treatment variable. This is an important outcome, yet future research will need to establish a broader base of incremental validity, and this may be facilitated by considering outcomes beyond self-report.
Strengths and Limitations
This study contained some strengths including a multi-method assessment of Criterion A and a comprehensive examination of Criterion A self-report measures within the same study. This study also has notable limitations, which can be organized by participants, measures, and data analysis.
Participant Limitations
The participant sample size was small, although this study focused mainly on convergent validity and therefore a smaller n would be appropriate in expecting moderate to large effect sizes, and the sample size is comparable to other studies evaluating interview-based methods for Criterion A. The sample was also drawn from a university setting, which, while not precluding psychopathology, the range restriction in Criterion A measures suggest it was a relatively healthy sample. While structural relationships between variables do tend to generalize in clinical and nonclinical samples (Bach et al., 2018; O’Connor, 2002), future research should replicate these findings in a clinical sample. In addition, a more diverse sample (e.g., race/ethnicity, gender, and age range) would also aid in evaluating the generalizability of these findings.
Measure Limitations
While five measures of Criterion A are the most that anyone has compared to date, several other measures do exist or have been published since this research was collected. Future research should continue to integrate these measures within a single study to understand their overlap and relative performance. We chose in this study to only use the average Criterion A scores, rather than to use subscales. We chose this because extant research suggests the subscales evidence poor discriminant validity (McCabe et al., 2021), the construct of Criterion A was designed to be a single dimension (Morey, 2019), and factor analysis results suggest that many Criterion A measures have a reasonable one-factor structure (Bliton et al., in press; Hummelen et al., 2021). Nevertheless, examining subscale-level data may be fruitful in determining which aspects of the Criterion A construct are contributing to an association with a particular outcome. This level of detail was beyond the scope of this article. The level of impairment from the LPFS-IR was low, which in some ways is expected in a relatively healthy college sample. However, it may also signal an underestimated level of impairment by RA coders (e.g., Zimmermann et al., 2014), or inadequacies of the stimuli (brief personality interview) or coding tool (LPFS) to capture psychopathology.
Our outcome measures were exclusively self-report, and may underrepresent psychopathology that could be assessed through other sources (e.g., informant reports, life history data). It may be particularly important to take a multi-method approach toward outcomes in evaluating method-based differences of AMPD instruments. More generally, the criterion validity of Criterion A measures is still being established, and a more thorough assessment of key severity markers is needed.
Future research should also consider collecting longitudinal data to evaluate these measures since to date only two measures (LPFS-SRA, LPFS-SR) have been evaluated using this type of data, and research found more nuanced patterns for Criteria A and B when considering longitudinal outcomes (Ringwald et al., 2021; Roche, 2018). Indeed, many of the Criterion A constructs themselves reflect temporally dynamic processes (e.g., inconsistent sense of self, vacillating goals, emotional lability, and if-then contingencies for personality impairment worsening when perceiving rejection or difference of opinions) that would be more directly captured in a temporally dynamic data collection approach.
Data Analysis Limitations
With a larger sample, more advanced data analyses could have been performed, including structural equation modeling (SEM), to determine whether the one-factor structure presumed in the Criterion A measures was reasonable. Similarly, with SEM, we could have modeled Criterion A as a latent construct, and then examined path models toward the outcomes (or perhaps latent outcomes), which would have arguably been more parsimonious than running separate models. This was not possible, given the small sample size, and we also feel it was more useful to treat Criterion A measures separately to facilitate comparison among them.
Despite these limitations, the present research provides a comprehensive comparison of Criterion A measures, finds some evidence for modest incremental validity with Criterion B, and introduces a new interview to capture Criterion A efficiently. We hope this moves the conversation forward on how to conceptualize and assess personality dysfunction.
Supplemental Material
sj-docx-1-asm-10.1177_10731911211059763 – Supplemental material for Comparing Measures of Criterion A to Better Understand Incremental Validity in the Alternative Model of Personality Disorders
Supplemental material, sj-docx-1-asm-10.1177_10731911211059763 for Comparing Measures of Criterion A to Better Understand Incremental Validity in the Alternative Model of Personality Disorders by Michael J. Roche and Sarah Jaweed in Assessment
Footnotes
Appendix
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
