Abstract
The assessment of intimate partner violence (IPV) by mental health, medical, and criminal justice practitioners occurs routinely. The validity of the assessment instrument they use impacts practitioners’ ability to judge ongoing risk, establish the type of IPV occurring, protect potential victims, and intervene effectively. Yet, there is no known compendium of existing assessment measures. The purpose of this article is threefold: (1) to present a systematic review of measures used to identify or predict IPV, (2) to determine which of these measures have psychometric evidence to support their use, and (3) to determine whether any existing measure is capable of differentiating between situational couple violence and intimate terrorism. A systematic search was conducted using PsycINFO, PsycARTICLES, PubMed, and MEDLINE. Studies on the reliability or validity of specific measures of IPV were included, regardless of format, length, discipline, or type of IPV assessed. A total of 222 studies, on the psychometric properties of 87 unique measures, met our a priori criteria and were included in the review. We described the reliability and validity of the 87 measures. We rated the measures based on the Consensus-based Standards for the Selection of Health Measurement Instruments–revised criteria and other established validity criteria, which allowed us to generate a list of recommended measures. We also discussed measures designed to differentiate IPV types. We conclude by describing the strengths and weaknesses of existing measures and by suggesting new avenues for researchers to enhance the assessment of IPV.
Keywords
When working with distressed couples, assessing for and identifying intimate partner violence (IPV) is essential because of the potential for serious injury or even death in situations of severe partner abuse. According to estimates by the National Intimate Partner and Sexual Violence Survey, within the United States, 22.3% of women and 14% of men experienced severe physical violence by an intimate partner in their lifetimes (Breiding, 2014). According to the FBI’s Supplementary Homicide Reports, which synthesizes data from all homicide cases in the United States between 1980 and 2008, 41.5% of all women murdered were murdered by current or past intimate partners (Cooper & Smith, 2011). In other words, the stakes are high in accurately determining the extent and lethality of IPV and making immediate and effective decisions on how to proceed because of the potential for serious harm and death.
The first step toward successful prevention of future episodes of IPV is often detection by medical personnel, mental health practitioners, or law enforcement officials (Ramsay et al., 2002). Several studies have demonstrated the benefits of using multimodal assessment techniques to detect IPV (e.g., Webster & Holt, 2004). Patients who endorse behaviors and symptoms in self-report measures sometimes do not endorse the same behaviors and symptoms in interviews and vice versa (O’Leary et al., 1992). This demonstrates the need for well-validated measures of IPV.
The first objective of this review is to identify existing measures of IPV and all studies presenting evidence for their validity, and the second objective is to describe each measure’s reliability, validity, and rating based on the COnsensus-based Standards for the selection of health Measurement INstruments (COSMIN)–revised criteria. To achieve these two objectives, we aspired to present a list of validated assessment instruments.
Evidence of an IPV Typology
An important and often overlooked consideration is that not all IPV is the same (Bates, 2016). In 1995, Michael P. Johnson proposed a typology that included two distinct types of IPV: intimate terrorism (also called patriarchal terrorism, coercive controlling violence) and situational couple violence (previously called common couple violence). He defined intimate terrorism as “a form of terroristic control of wives by their husbands that involves the systematic use of not only violence, but economic subordination, threats, isolation, and other control tactics” (Johnson, 1995, p. 284). Conversely, situational couple violence is not motivated by control but is a dynamic in which “conflict occasionally gets ‘out of hand,’ leading usually to ‘minor’ forms of violence, and more rarely escalating into serious, sometimes even life-threatening, forms of violence” (Johnson, 1995, p. 285).
Several studies have validated Johnson’s types (Graham-Kevan & Archer, 2003; Johnson, 1999, 2009). However, few researchers have attempted to create and validate assessment measures that take Johnson’s types into account. Having such a measure would be instrumental in assisting clinician’s decision-making processes when working with couples experiencing IPV. The most important implication of the distinction between couples who have experienced intimate terrorism versus situational couple violence is that there are differences in the most appropriate response to the disclosure of IPV in these two types of couples; therefore, the ability to differentiate correctly between these two types of IPV is vitally important for treatment decisions. For this reason, we include an objective of locating any measures that seek to differentiate between types of IPV in our review. For the sake of clarity, moving forward in this article, we will refer to Johnson’s two types of IPV as “types,” and we will refer to differing modes of abuse such as physical, sexual, or psychological as “forms” of IPV.
Prior Reviews of IPV Risk Assessments
Previous large-scale systematic reviews have focused on IPV risk assessment measures used in the criminal justice field or on screening measures used in the medical field (Arkins et al., 2016; Nichols et al., 2013; Rabin et al., 2009). The previous reviews each provided information on 21, 19, and 10 measures, respectively. All three were limited in scope to one specific discipline, leading to the exclusion of measures not directly designed as medical or criminal justice screeners.
Rabin and colleagues (2009) completed a systematic review of 21 IPV screening measures used in medical settings. Their review excluded measures of frequency or severity. They concluded that even the most commonly used measures are understudied, that few measures have been validated for male victims, and that utility is important in addition to psychometric properties.
Nichols and colleagues (2013) performed a systematic review of 19 measures that predict the risk of IPV perpetration. They concluded that there are many risk assessments available and most provide above chance predictions of who will engage in IPV. They also urged users of these measures to match their chosen measure to the purpose they are seeking to fulfill (e.g., predict violence vs. decide on treatment needs).
Arkins and colleagues (2016) focused their systematic review on screening measures used in a mental health context. They also excluded measures of severity and frequency in addition to measures for perpetrators. They concluded that about 30% of the 10 measures they identified had strong psychometric properties and that further research is needed on cultural and gender sensitivity.
While dozens of measures for assessing IPV were identified in these three reviews, none of them included information on measures designed to differentiate types of IPV. Thus, the third objective of this review is to identify any studies attempting to create a measure that can differentiate types of violence or provide evidence that an existing measure can be used to identify types of IPV.
Summary of Objectives
To summarize, we had three objectives: (1) to develop a comprehensive list of measures assessing the future risk of recidivism of IPV, current presence of IPV, or current severity of IPV; (2) to describe each measure’s reliability, validity, and rating based on the COSMIN-revised criteria, which would allow us to provide practitioners a list of validated assessment instruments; and (3) to identify and evaluate any assessment instruments that purport to identify types of IPV.
Method
To locate IPV assessment instruments (Objective 1), we conducted a systematic search in PsycINFO, PsycARTICLES, PubMed, and MEDLINE during September 2018 including all available sources published before the date of the search, using the following Boolean operators: (“intimate partner violence” or “domestic abuse” or “domestic violence” or “IPV” or “patriarchal terrorism” or “common couple violence” or “interpersonal violence” or “couple violence” or “intimate partner terrorism”) AND (“assessment” or “questionnaire” or “measure” or “interview” or “scale.”).
The search produced 15,778 records (i.e., articles, books, chapters, dissertations). Following the initial search, 55 records were located through Google Scholar searches and the reference sections of records identified in the first search. Inclusion and exclusion criteria were determined a priori. We excluded records not available in English and records without editorial- or peer-review processes, such as dissertations. The remaining records were examined first in a title screen and then an abstract screen to determine relevance for this review. Of the original records, 13,495 were available in English and had editorial- or peer-review processes. After duplicates were removed, 10,301 records remained. The titles and abstracts of these records were screened to determine relevance to this review (whether the records presented reliability or validity studies on specific IPV assessment measures). The full text of 249 of these records was read to determine eligibility, and a resulting 211 records were included in the review. The results of the search are presented in a flow diagram according to Preferred Reporting Items for Systematic Reviews and Meta-Analyses standards (Figure 1).

Preferred reporting items for systematic reviews and meta-analyses flow diagram.
After compiling the list of records, we read the records to identify the names of specific IPV assessment measures. In our review, we included measures of intimate partner abuse of any form (physical, psychological, economic, etc.); assessments in any format (e.g., self-report measures, interviews); and both well-known, widely used measures and lesser-known, regional measures. We excluded measures of the child, elder, or family abuse; measures focused on short-term interactions or casual relationships rather than committed relationships; measures of violence without a specific focus on IPV; and measures of IPV treatment outcomes, willingness to end the relationship, or other such outcome measures.
Once our list of assessments was complete, the name of each measure on the list and the words “validation” or “validity” or “development” were searched one at a time in the previously used databases to attempt to locate further validation studies. An additional 11 studies were located in these searches, leading to a final total of 222 records in the review. For each measure, the Online Appendix describes the measure’s name, author(s), format (e.g., self-report, interview), and any unique uses or specific populations for which the measure is intended. The Online Appendix also includes information on any records presenting data on the reliability or validity of that instrument.
To achieve Objective 2, we rated each measure according to the COSMIN updated criteria for strong measurement properties (Mokkink et al., 2018). The COSMIN updated system is a system of examining studies of measurement properties to ensure methodological quality and strong evidence in support of a measure (Prinsen et al., 2018; for more information on the development of the COSMIN system, see Mokkink et al., 2017; Terwee et al., 2018). The COSMIN system provides ratings of sufficient (+), indeterminable (?), or insufficient (−) across a number of categories of reliability and validity. Guidelines provided for cutoff values that are considered sufficient for various types of reliability and validity, including structural validity, internal consistency, reliability, measurement error, hypothesis testing for construct validity, cross-cultural validity/measurement invariance, criterion validity, and responsiveness. Indeterminable ratings indicate not enough information was provided to determine a sufficiency rating. Table 1 is a table from the COSMIN Manual summarizing these criteria (reproduced with permission of the authors; Mokkink et al., 2018, p. 28). For each study, we made a rating for every type of reliability and validity assessed. An ideal measure would receive sufficient ratings across studies. In contrast, inconsistent ratings across studies reflect poorly on the measure if there is not a reasonable explanation for why these inconsistencies occurred. The COSMIN system has been used in several previous systematic reviews of measures (e.g., Statham et al., 2019; Wei et al., 2016).
Updated Criteria for Good Measurement Properties.
Source. Table reprinted, with permission, from the COSMIN User Manual (Mokkink et al., 2018, p. 28).
Note. The criteria are based on, for example, Terwee et al. (2018) and Prinsen et al. (2018). AUC = area under the curve; CFA = confirmatory factor analysis; CFI = comparative fit index; CTT = classical test theory; DIF = differential item functioning; ICC = intraclass correlation coefficient; IRT = item response theory; LoA = limits of agreement; MIC = minimal important change; RMSEA = root mean square error of approximation; SEM = standard error of measurement; SDC = smallest detectable change; SRMR = standardized root mean residuals; TLI = Tucker–Lewis index.
a “+” = sufficient, “−” = insufficient, and “?” = indeterminate. b To rate the quality of the summary score, the factor structures should be equal across studies. c Unidimensionality refers to a factor analysis per subscale, while structural validity refers to a factor analysis of a (multidimensional) patient-reported outcome measure. d As defined by grading the evidence according to the Grading of Recommendations Assessment, Development and Evaluation (GRADE) approach. e This evidence may come from different studies. f The criteria “Cronbach α < .95” was deleted, as this is relevant in the development phase of a PROM and not when evaluating an existing Patient‐Reported Outcome Measures (PROM). g The results of all studies should be taken together, and it should then be decided if 75% of the results are in accordance with the hypotheses.
In order to make specific recommendations on measures to use, we used a series of thresholds that are described in Figure 2. Our first standard was that a measure must have been tested in more than one validation study. Second, both the reliability and the validity of the measure must have been studied. Third, there must have been at least two validation studies that used different methods from one another. These validation methods could have been of the same broad category but unique in the specific evidence provided. For example, there could have been multiple studies of criterion validity but with different specific criterion measures. Fourth, researchers unaffiliated with the developers of the measure must have completed one of the validation studies. In the fifth threshold, we incorporated the COSMIN ratings. The measure must have at least half positive COSMIN ratings, demonstrating that evidence against using a measure is not stronger than evidence for it. Finally, the measure must have been studied with multiple samples. We have allowed that these samples be of the same type of source, such as all hospital samples, if the participants are from different specific locations, such as different specific hospitals. The measures that met these requirements are presented as “recommended” in the Results section. We view these criteria as minimal and encourage further evaluation of the measures when selecting an instrument for a specific purpose.

Thresholds for determining which measures of intimate partner violence to recommend.
Results
We found 87 distinct measures of IPV (when alternate forms and lengths were counted separately). Information on each measure including its length, delivery format, purpose/intended population of use, and the reliability and validity data from studies on the measure is included in the Online Appendix.
Recommended Measures of IPV
Based on the aforementioned criteria, we recommend the 18 (of 87) measures listed in Table 2. As of now, these are the measures that have been studied multiple times, including at least once by an unaffiliated researcher, and demonstrated validity for the intended purposes using multiple methods. The measures generally fit into the following five broad categories: (1) brief screening instruments, (2) measures of severity or that identify high-risk cases, (3) measures of criminal recidivism, (4) measures of constructs and attitudes related to violence without overtly asking about violent behaviors, and (5) measures of specific forms of abuse (e.g., physical, sexual, economic).
Critical Findings: Recommended Measures of Intimate Partner Violence With Brief Descriptions.
Note. See Online Appendix for more details on each of these measures and their validity studies.
a The designation as having been studied in multiple distinct cultural groups requires that a validity study has been conducted that specifically sought to assess whether the measure is valid in a cultural group distinct from the original validation sample. This designation was given if the measure was tested in distinct racial, ethnic, or sexual minority groups.
Brief screening instruments
The Abuse Assessment Screen (AAS); the Hurt, Insult, Threaten, and Scream (HITS); and the Woman Abuse Screening Tool (WAST) were the three brief screening measures that met our criteria. The AAS was designed specifically to screen pregnant women because they are a high-risk population (McFarlane et al., 1992). This measure’s validity has been assessed in seven distinct studies, plus one study validating an alternate form of the measure that is specifically intended for women with disabilities. Reliability was also assessed in four of the seven studies. The HITS was designed to be as brief as possible and contains only four items (Sherin et al., 1998). It was originally designed to be used by family care physicians. All seven validity studies on this measure received sufficient COSMIN ratings. Considering the brevity of the measure, it is somewhat surprising and quite promising that when reliability was also assessed (in three of the seven studies), these ratings were also sufficient (but not for the Spanish translation). Finally, the WAST is also meant for screening by family care physicians, and both the long-form and short-form met our criteria for recommendation (Brown et al., 1996). There are nine validity studies on the long-form and four on the short form. The WAST has been tested with men and women and in six different countries. Six studies also assessed reliability, all finding sufficient results.
Measures of severity/high-risk cases
Two recommended measures, the Danger Assessment (DA) and the Index of Spouse Abuse (ISA), focus on violence severity. The DA focuses on the severity in that it is intended to identify the most severe cases (J. C. Campbell, 1986). It was designed to predict femicide or near femicide. The DA has been validated in 11 studies. It also has a short form (DA-5) and has been used as the basis for additional measures (e.g., lethality screen). Two of the 11 studies also assessed reliability. The ISA aims to capture the full spectrum of severity and produces severity scores for both physical and nonphysical abuse (Hudson & McIntosh, 1981). The ISA’s reliability and validity have been assessed in eight studies.
Measures of criminal recidivism
One actuarial measure designed to predict criminal recidivism, the Domestic Violence Risk Appraisal Guide (DVRAG), is recommended. The DVRAG combined the Ontario Domestic Assault Risk Assessment (ODARA) with content from the Psychopathy Checklist—Revised (PCL-R) to create a risk assessment for criminal justice purposes (Hare & Neumann, 2006; Hilton et al., 2008; Hilton et al., 2004). This measure’s reliability and validity have been assessed in two studies. However, there is a large overlap between this measure and the ODARA, so studies on the ODARA are also helpful to understanding the utility of the DVRAG.
Measures of constructs and attitudes surrounding violence
This somewhat diffuse category of recommended instuments captures measures that do not ask about overt abuse behaviors. Instead, these measures ask about correlates of abuse, ideas and attitudes around abuse, or abuse as a broader concept. The Intimate Justice Scale (IJS) is a measure of “ethical dynamics” and broader patterns in relationships (Jory, 2004). To avoid demand characteristics, it has relatively low face validity for a measure of abuse. The IJS is also used to predict the severity of IPV. The IJS has been validated in two studies, one of which also assessed reliability. Similarly, the Propensity for Abusiveness Scale (PAS) has intentionally low face validity (D. G. Dutton, 1995). This measure is intended to be administered to perpetrators rather than victims of violence, an unusual characteristic among measures, and the only measure for use with perpetrators (not victims or both members of a couple) on our recommended list. The PAS has been validated in five studies. Of the five, four also assessed reliability. The PAS measures propensity toward abusiveness based on history, personality traits, and mental health symptoms.
The Intimate Partner Violence Attitude Scale (IPVAS) and Wife Abuse Inventory (WAI) are both risk assessments designed to identify people at risk for abuse. The IPVAS identifies people’s attitudes toward IPV to attempt to identify those at risk of perpetrating or being a victim of violence in the future (B. A. Smith et al., 2005). Its reliability and validity have been studied in three samples. The WAI is meant to predict those at risk of becoming victims for prevention purposes (Lewis, 1985). Two studies on this measure’s reliability and validity have been completed, receiving all sufficient COSMIN ratings.
Finally, the Women Experience with Battering (WEB) Scale is a measure of women’s experiences in their abusive relationships (P. H. Smith et al., 1995). Items ask about psychological vulnerabilities and emotional reactions to the abusive relationship to understand not the specific behaviors that are occurring but how the victim is impacted. The WEB Scale’s validity has been assessed in three studies, one of which also assessed reliability.
Measures of forms of violence
Finally, six recommended measures focused on specific forms of abuse (e.g., physical, sexual, economic). The Abuse Behavior Inventory was developed from feminist theory and focuses on intimate terrorism (Shepard & Campbell, 1992). Three studies on its reliability and validity were located. Similarly, the Controlling Behaviors Scale (CBS) focuses only on behaviors aimed at controlling one’s partners (Graham-Kevan & Archer, 2003). Its reliability and validity have been assessed in four studies, three of which showed very promising and sufficient results. These measures do not include any items intended to capture situational couple violence.
Three recommended measures focused on nonphysical abuse. The Multidimensional Measure of Emotional Abuse (MMEA) is a measure of emotional abuse, broken down into multiple categories of emotional abuse (e.g., intimidation and denigration; Murphy & Hoover, 1999). There are three reliability and validity studies on the MMEA. The Psychological Maltreatment of Women Index is a measure of psychological maltreatment of women (Tolman, 1989). It is among the older measures on the list and has reliability and validity evidence from four studies. Finally, the Scale of Economic Abuse (SEA) is a measure specifically designed to capture economic abuse (Adams et al., 2008). The SEA has reliability and validity evidence from four studies.
Lastly, the NorVold Domestic Abuse Questionnaire measures four forms of abuse: emotional, physical, sexual, and health care abuse. This measure is intended for research purposes. It includes items about abuse by medical professionals in addition to IPV items. It has two reliability and validity studies both receiving sufficient COSMIN ratings.
Discussion
In identifying 87 measures of IPV varying in length, intended purpose, and format, we attempted to describe each measure’s reliability, validity, and rating based on the COSMIN-revised criteria (for details of the 87 measures, see the Online Appendix). Based on this review, we recommend the measures in Table 2 as well validated. The variety of available and recommended measures makes it more likely that researchers and practitioners can adopt a measure designed for their purpose and population. Among the measures we located were measures designed specifically for unique forms of abuse (e.g., emotional, economic); a measure that presented pictorial items; measures as brief as two items and as long as 92; and measures for specific populations such as military members, women with disabilities, lesbians, and pregnant women. The distinctiveness of the measures should lead those considering assessing IPV to attend to the purposes for which these measures were developed. A few aspects of our review merit further discussion and research.
Selecting Criterion Measures: In Search of the Gold Standard
While there is certainly a great amount of commendable validity research on IPV assessment, more needs to be done. Many of the measures in this study were validated by being compared with one another. Further support for IPV measures can be built through diversifying the criterion measures used.
The Conflict Tactics Scale (CTS)
The most commonly used “gold standard” measure was the CTS ( we do not distinguish among versions for the purposes of this discussion). The CTS has many validity studies in support of its use, and many of these studies come from large and diverse samples. At the same time, the CTS has faced criticism for not asking about certain forms of abuse, hierarchically ranking items by severity where a true hierarchy may not exist, not assessing for the motivation behind violent acts, and framing violence only in the course of an episode of conflict when not all violence occurs in this context (DeKeseredy & Schwartz, 1998). The CTS is so widely used that providing a full analysis of its strengths, weaknesses, and applications is beyond the scope of this article (for reviews of this measure, see Archer, 1999; Schafer, 1996). In our COSMIN ratings, the CTS received a mix of sufficient and insufficient ratings across different studies and types of validation. While the measure is one of the most widely studied and remains a useful resource, using other criterion measures in addition to the CTS could help demonstrate a measure’s utility across contexts.
Criminal recidivism
In addition to the CTS, one of the most common outcomes/comparison measures used across studies, particularly in the area of criminal justice, was criminal recidivism. Recidivism provides important information for criminal justice officials, but care must be taken not to equate recidivism with the severity of violence or with reoffending. Some studies suggest that recidivism may be mediated by jail time (Williams & Stansfield, 2017). If someone is in prison, they cannot reoffend because of the physical separation from their partner. This should be factored into studies using recidivism as an outcome. Additionally, not all IPV perpetrators are arrested or incarcerated, and criminal recidivism captures only those who are. Recidivism measures also often do not capture nonphysical violence. Finally, some research suggests that arrest reduces subsequent reoffending even after the perpetrator is released (Maxwell et al., 2002).
In general, the lack of a clear “gold standard” for criterion validity impedes efforts to validate IPV measures. In their classic article on assessing construct validity, D. T. Campbell and Fiske (1959) describe how conducting a multitrait–multimethod matrix enables a thorough description of the construct validity in the absence of an easily identifiable criterion. We found few studies met their—albeit high—bar. Therefore, we urge a “buyer beware” approach to selecting an IPV measure by paying close attention to the criterion and outcome measures used and ensuring that the measures chosen are used only for the purposes for which they have been validated.
No Measures Validly Categorize Types of IPV (Objective 3)
Returning to Johnson’s (2009) typology of IPV, two of the 87 measures were intended to differentiate the types of IPV; however, in both measures, the authors interpreted Johnson’s two types as two ends of a severity continuum. The Continuum of Conflict and Control Relationship Scale (CCC-RS) was created using exploratory factor analysis and intended to distinguish the two types of couple violence (Carlson et al., 2017). Despite the intention and methodology, both the “controlling violence” factor and the “relational conflict” factor include items describing controlling behaviors and conflict escalation behaviors. Our evaluation of this measure led us to conclude that the factors align more with IPV severity than with the stated intent of distinguishing the two types of IPV in Johnson’s typology. We reached this conclusion based on the fact that severity of violence seems to be the main distinction and based on both factors including items describing behavior that should—in theory—be distinguishing characteristics of Johnson’s types of IPV.
The IJS (Jory, 2004) is a measure of IPV intended to categorize levels of severity. It is distinctive among measures of IPV in that it includes items about relationship dynamics rather than specific behaviors and, as such, it appears to be less face valid in an attempt to reduce minimization or denial of violence. Friend et al. (2011) maintain that the IJS can distinguish between couples experiencing situational couple violence and intimate terrorism; however, to establish the validity of the IJS for this purpose, they compared groups classified by the IJS to groups classified by the severity of violence endorsed on the CTS. This method of validation has two problems: CTS scores have not been shown to differentiate IPV types and, as with the CCC-RS, severity was conflated with type. While severity often coincides with the type of violence, this is not always the case, so it is important that severity and type not be equated (Kelly & Johnson, 2008). Therefore, our conclusion is that neither the CCC-RS nor the IJS can validly categorize respondents as experiencing intimate terrorism versus situational couple violence. While it could be argued that measures of severity would suffice, the ability to identify intimate terrorism before it escalates in severity could have a profound prophylactic effect.
Several measures are based on a theoretical grounding of coercive control, for example, the Intimate Partner Violence Control Scale (IPVCS; Bledsoe & Sar, 2011), the Checklist of Controlling Behaviors (CCB; Lehmann et al., 2012), the Coercive Control Measure (CCM; M. A. Dutton et al., 2005), and the CBS (Graham-Kevan & Archer, 2003). However, these measures do not incorporate situational couple violence. Additionally, further validity evidence is needed to demonstrate their utility. Our search turned out only one validity study each on the IPVCS, CCM, and CCM.
In the end, the goal of quickly and easily differentiating IPV types appears illusory despite the importance of making decisions about the safest, most effective treatment for couples. Court-mandated interventions, such as the widely used Duluth curriculum, consist largely of educating men about their beliefs about women and IPV from a power and control standpoint (Armenti & Babcock, 2016). The treatment is intended to be provided in men-only groups. While these programs are effective for some patients, overall, the effect sizes have been small to null/nonexistent (Babcock et al., 2004; Feder & Dugan, 2002). Some evidence even suggests that men perpetrating low levels of violence may increase the severity of their violence following group treatment after learning from the more violent men in the group (Armenti & Babcock, 2016).
The programs that target beliefs about controlling women are designed around the assumption that men are coercive controlling offenders, when this type of violence is actually far less common (Johnson, 2001). Some situationally violent men find it insulting that educational programs such as this treat them as if they are intentionally seeking to control their partners (Raab, 2000). This suggests that treatment may not be “one-size-fits-all,” and that separating the types of IPV perpetrators into separate groups that target their specific motivations toward violence may be beneficial.
Another treatment method sometimes used in cases of IPV is conjoint couples’ therapy. Some researchers have found no differences in effectiveness between couple therapy and single-sex group therapy (Brannen & Rubin, 1996; O’Leary et al., 1999); however, other researchers have found that couple therapy may be more effective than single-sex groups (Mills et al., 2019; Simpson et al., 2008; Stith et al., 2004). There is also evidence that targeting relationship problems and skills rather than beliefs may be effective for reducing abuse (Wray et al., 2013). Kelly and Johnson (2008) describe how situationally violent couples may benefit from custody mediation whereas victims of intimate terrorism may be coerced into not asserting their needs in a mediation setting. A few recent studies have tried to target situationally violent couples for couple treatment and have found results demonstrating that for these couples, conjoint treatment seems to be safe and effective (Armenti & Babcock, 2016).
These studies, and others, make clear that the ability to categorize IPV as either situational couple violence versus intimate terrorism can be important to couples’ safety and ability to tolerate different treatment approaches. Having a well-validated measure of Johnson’s two types could be instrumental in aiding treatment decisions.
Limitations of Our Review
Although the COSMIN rating system is a strong, empirically based rating system, it does have limitations. The quality of a measure is difficult to condense down to a simple plus or minus and attempting to do so will naturally result in losing some of the complexities involved in describing the validity of the measures we reviewed. For example, criterion validity is scored based on whether the measure correlates with the gold standard, but there are inherent difficulties in determining which criterion to use when the options for a criterion measure are less than or greater than one. Any effort to reduce complex evaluations such as validity to categories such as “+,” “?,” or “− ” is necessarily imperfect. For example, Cronbach’s αs of .70 or greater are necessary but not sufficient for earning a “+” yet looking at the symbols alone will result in missing the distinction between measures with Cronbach’s αs of .70 and .95 or may lead to discounting measures that were very close to reaching the cutoff. Because of this, readers are strongly encouraged to read the primary source article about a measure if interested in using it.
It is also possible that there are excellent measures or validity studies missing from our review. For example, this review was limited to studies available in English, and there may be measures in languages other than English that would have received sufficient ratings across studies. The exclusion of any measure that fits our inclusion criteria is unintentional. Although we attempted to create as exhaustive a list as possible, the literature is always expanding.
Recommendations in our review are based on standards of validity evidence and not creativity. As we noted before, there was an impressive variety in our review, and we urge readers to review all of the measures in the Online Appendix not just those that are recommended in Table 2. Validity studies are designed for a particular set of inferences to be drawn from that measure (Cronbach & Meehl, 1955). It is plausible that the best measure for a particular purpose is not included on our recommended list but is still well suited to specific situations and populations (e.g., IPV Assessment Icon Form for populations with low literacy).
Finally, this review focused on reliability and validity in determining recommended measures. We did not complete an in-depth analysis of the wording of the items in the recommended measures. Therefore, further evaluation may be needed to determine whether these measures use inclusive, respectful language and are applicable across cultures. This is an important consideration because IPV is a sensitive topic, and many individuals may be hesitant to report on their experiences. Ideally, researchers will soon examine the specific item content of items in the measures and whether they are cross-culturally valid.
Although not a direct limitation of this study, overwhelmingly, the records included in this review did not include diverse populations or assessment of cross-cultural validity. Of the 18 recommended measures, 11 were evaluated in multiple distinct cultural groups including at least one minority group (see Table 2 for a list). However, one study in one minority group is not enough, and further translations and cross-cultural validity studies are needed for all measures.
Strengths of Our Review
Despite limitations, we compiled—in one article—the current options for English-language measures assessing IPV across multiple disciplines. With over 200 studies evaluated as part of this review, this article catalogs a wider, more cross-disciplinary range of IPV assessment instruments than in previous reviews. Measures were compiled that were designed to be used in various unique situations and with diverse populations. These measures came from several different disciplines. It is our hope that in compiling resources in this way, professionals across fields (e.g., mental health workers, criminal justice professionals, nurses, researchers, and others) can learn from each other’s progress. Using this study’s findings as a starting point, researchers can use knowledge of what has and has not been included in previous assessments to design new, innovative studies and create measures that address aspects of IPV that are underaddressed in existing measures, such as situational couple violence versus intimate terrorism. In addition, further replication and validation of measures across settings and samples will go a long way toward understanding IPV, both in research and practice.
This review also systematically compiled a list of recommended measures with the strongest evidence in support of their use for measuring IPV. This list can help clinicians and researchers narrow down the decision of which measure to use from among the many options available. Information is also provided on the populations and situations with which these measures have been studied to allow for ease of deciding which measures are best for which purposes.
Conclusion
To summarize our review, over 80 distinct measures of IPV exist. Of the many IPV measures, 18 show strong evidence of their utility for detecting violence, predicting recidivism, and even predicting homicide. Many measures exist to evaluate IPV in specific populations in specific contexts. However, despite claims to the contrary, none can validly differentiate which couples are experiencing situational couple violence versus intimate terrorism. Future research on IPV assessment is needed, with a particular need for further validation studies on several promising yet understudied measures and for the design of measures capable of differentiating IPV types.
Implications for Practice, Policy, and Research
This review has relevant implications for practitioners, policy makers, and researchers. Implications are summarized in Table 3. Practitioners are urged to use validated measures, such as the ones recommended in this review. They are encouraged to choose which measures to use carefully, considering the purpose of their assessment. Finally, criminal justice workers are encouraged to incorporate validated measures into their decision-making processes. Policy makers who create mandated intervention programs are urged to consider what measures are used in evaluating fit of individuals to interventions. Police departments that have policies requiring the use actuarial tools in decision making should require only instruments which have strong validation evidence. The same is true for policies passed about universal screening in health care settings. This review points to several areas of research that need further exploration. We recommend that researchers conduct further validation studies on several understudied measures, using strong and varied research methods. Further research is also needed to construct measures that incorporate types of IPV. If these advances in practice, policy, and research are made, the safety of victims and the quality of care that couple experiencing IPV will hopefully continue to improve.
Supplemental Material
Supplemental Material, sj-pdf-1-tva-10.1177_15248380211013413 - Evaluating Measures of Intimate Partner Violence Using Consensus-Based Standards of Validity
Supplemental Material, sj-pdf-1-tva-10.1177_15248380211013413 for Evaluating Measures of Intimate Partner Violence Using Consensus-Based Standards of Validity by Erin F. Alexander, Bethany L. Backes and Matthew D. Johnson in Trauma, Violence, & Abuse
Footnotes
Authors’ Note
Christina Balderrama-Durbin, Mark Lenzenweger, and Clara Mildenberger provided guidance and feedback on this project. Their assistance is gratefully acknowledged.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
The supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
