Abstract
Empathy is a core component of social cognition that can be indexed via behavioral, informant-report, or self-report methods of assessment. However, concerns have been raised regarding the lack of convergence between these assessment approaches for cognitive empathy. Here, we provided the first comparison of all three measurement approaches for cognitive and affective empathy in a large adult sample (N = 371) aged 18 to 101 years. We found that poor convergence was more of a problem for cognitive empathy than affective empathy. While none of the cognitive empathy measures correlated with each other, for affective empathy, self-report was significantly associated with both behavioral and informant-report assessments. However, for both cognitive and affective empathy, there was evidence for poor discriminant validity within the measures. Out of the three assessment approaches, only the informant-report measures were consistently associated with indices of social functioning. Importantly, age did not moderate any of the tested relationships, indicating that both the strengths and the limitations of these different types of assessment do not appear to vary as a function of age. These findings highlight the variation that exists among empathy measures and are discussed in relation to their practical implications for the assessment of empathy.
Empathy is a core component of social cognition that plays a fundamental role in social relationships (Henry et al., 2016). Although the precise definition of empathy varies considerably across the literature (Hall & Schwartz, 2019), most agree that it is a multifaceted construct that includes both cognitive and affective components (Davis, 1983; Decety & Cowell, 2014; de Waal & Preston, 2017). While cognitive empathy refers to the ability to understand the perspectives or feelings of others, affective empathy refers to one’s emotional response to someone else’s situation. Although both are important for social well-being at all stages of the human lifespan, they appear to be differentially impacted by normal adult aging (Bailey et al., 2008; Khanjani et al., 2015; Sun et al., 2018). Specifically, while most prior work has shown that cognitive empathy declines in older adulthood (Beadle & de la Vega, 2019; Henry et al., 2013), some studies have shown that affective empathy remains stable (Bailey et al., 2008) or possibly even improves (Grainger et al., in press; Sun et al., 2018; Sze et al., 2012; Ze et al., 2014). However, some notable inconsistencies have been observed in both the literature, with some studies identifying age-related stability or improvement in cognitive empathy (Chen et al., 2014; Katzorreck & Kunzmann, 2018; Ze et al., 2014) and others age-related decline in affective empathy (Chen et al., 2014).
One possibility is that these prior inconsistencies simply reflect differences in measurement approach across individual studies, as three distinct approaches have been used: (a) behavioral tasks, (b) self-report, and (c) informant-report. Behavioral tasks index “state” empathy by assessing real-time empathic responses to stimuli that is often emotionally evocative (such as pictures or videos depicting an individual in pain or distress). 1 Participants are asked to state what they believe a central protagonist is thinking or feeling (to index their cognitive empathy) and/or describe their own subjective emotional response to that protagonist (to index their affective empathy). By contrast, self-report and informant-report tasks assess “trait” empathy by asking hypothetical questions about how a person believes they would typically respond in specific scenarios. Self-report assessments are completed by the individual and therefore require intact self-awareness and insight. Informant-reports are completed by someone who knows the participant well (i.e., partner, close friend, or family member) and while they have the advantage of providing an independent evaluation of an individuals’ empathic capacity, their validity is contingent on the availability of an informant who is: (a) willing to provide an accurate report and (b) has sufficient cognitive empathy of their own to accurately detect, understand and describe how that individual thinks and feels.
Self-report assessments are the most efficient of these three approaches to administer, and consequently, it is not surprising that they are the most commonly used in empathy research (Hall & Schwartz, 2019). However, self-report empathy measures are often treated as proxies for actual empathic ability, and this has prompted researchers to examine the extent to which the data generated by self-report empathy tasks is a true reflection of an individual’s actual empathic abilities. Most notably, Murphy and Lilienfeld (2019) used meta-analytic methodology to examine the degree to which self-reported cognitive empathy ratings are related to performance on behavioral cognitive empathy tasks. Aggregating data across 85 studies, they found that self-reported cognitive empathy scores accounted for <1% of the total variance in behavioral task performance. Stosic et al. (2022) examined the degree to which self-reported empathy predicts behaviorally indexed social cognitive function. Consistent with Murphy and Lilienfeld (2019), they found that self-reported cognitive empathy was unrelated to performance on nine different standardized social cognitive tasks. However, findings from Israelashvili et al. (2019) were more promising. They aggregated data from six independent datasets and found a small to moderate correlation (r = .20) between self-report and behavioral cognitive empathy and concluded that a person’s beliefs about their own cognitive empathy ability somewhat reflects their actual ability. However, although Israelashvili et al. found some agreement between the two assessment types, the level of agreement was only modest. Taken together, these findings suggest that self-report cognitive empathy scales may not be suitable proxies for actual cognitive empathy ability. One possible explanation for the poor convergence between self-report and behavioral empathy tasks is that people are not very accurate in self-reporting their own empathic capacity. Indeed, Vazire and Carlson (2010) argue that self-perceptions of personality are often distorted by “blind spots” in one’s true knowledge about themselves. These blind spots can occur because of inaccurate knowledge about oneself or by motivated cognitive processes (e.g., engaging in thoughts or beliefs to maintain or enhance feelings of self-worth). Informant-reports are more objective than self-reports and may in fact provide a more accurate evaluation of an individual’s traits (Vazire & Carlson, 2010; Wright et al., 2021). Indeed, many traits and abilities, including some aspects of empathic behavior (i.e., showing care or concern), are highly visible and are therefore easily observed and evaluated by other people.
Yet self- and informant-reports are affected by different favorability biases that can influence the way in which people evaluate themselves and others. While self-report is associated with a bias to emphasize traits associated with one’s self-efficacy and competence, informant-reports are associated with a bias to emphasize the target’s cooperative traits (Ashton & Lee, 2010). Given that many different factors have the power to differentially influence evaluations provided by self- and informant-report measures, it might be expected that evaluations by these two assessment types would be similar but not identical. While there is very little work to date that has examined how self- and informant-reports are related in the context of empathy, the general finding in the broader person perception literature is that self- and informant-reports correlate modestly (see Connelly & Ones, 2010), suggesting some but not perfect agreement between these two types of evaluations. As mentioned earlier, very little is known about the degree to which informant-reports are related to self-report and behavioral measures of empathy. Cliffordson (2001) was the first to examine the level of agreeance between self- and informant-report empathy ratings. Using a sample of parents and their adolescent children, they found modest agreement between self- and informant-report ratings on the Interpersonal Reactivity Inventory (IRI, Davis, 1983). Since Cliffordson’s study—which was published over two decades ago—only a single study has directly examined the agreeance between self- and informant-report measures of empathy in healthy adults. Using a sample of adults in the nursing profession, Roth and Altmann (2021) found a moderate correlation between these two assessment types when the informant was an intimate partner (r = .35), suggesting good convergence between self- and informant-report empathy. They also examined the relationship between informant-report and behavioral cognitive empathy performance. These two assessment types were not correlated, suggesting that an informant may be no better, and could potentially be even worse at estimating an individual’s actual cognitive empathic capacity relative to the individual themselves. However, it is important to note that Roth and Altmann (2021) used a global measure of empathy for the self- and informant-report measures, which measured aspects of both cognitive and affective empathy. It is therefore unclear whether the strong convergence identified between the self- and the informant-report was evident for both or just one of these empathy components. No study to date has examined convergence between self- and informant-report measures separately for the two empathy subcomponents (i.e., cognitive and affective); therefore, this study will be the first to address this gap.
With respect to affective empathy, a small handful of studies have examined convergence between behavioral and self-report measures with several showing relatively good convergence (Van der Graaff et al., 2016; Westbury & Neumann, 2008). For instance, Westbury and Neumann (2008) found scores on the Balanced Emotional Empathy Scale correlated strongly with subjective empathy ratings to emotional film clips. Van der Graaff et al. (2016) also collected emotional ratings in response to film stimuli and found modest correlations with self-reported empathic concern in response to negative films. These findings suggest that self-reported trait empathy may be a reliable proxy for behavioral affective empathy. However, Foell et al. (2018) validated the English version of the Multifaceted Empathy Test (MET) that involves viewing images of people in different emotional situations. For affective empathy, they identified a significant correlation between the empathic concern subscale of the IRI (IRI-EC) and ratings of affective empathy in response to positive images (r = .22); however, no significant association emerged for negative images. They also examined convergence between the cognitive component of the MET and the perspective-taking scale of the IRI (IRI-PT) and found no correlation, aligning with prior work showing poor convergence for cognitive empathy (e.g., Stosic et al., 2022). No study to date has investigated the degree to which informant-reports are associated with behaviorally indexed affective empathy. However, it might be expected that, given that affective empathy is a highly visible trait (i.e., it often involves feeling emotions and displaying care or concern toward another person), it may be better judged by an informant than the self. This study will provide the first direct test of this possibility. Another important question that remains unanswered in this literature is the degree to which these different measurement approaches are associated with important functional outcomes. Given that strong empathic abilities are necessary to interact effectively with others, and clinical groups that present with social function impairment often show deficits in empathy (Henry et al., 2016), empathic capacity should be associated with a range of important social function outcomes. Indeed, prior studies have identified links between empathy and numerous social indices, including sociality, personality, loneliness, and relationship quality (e.g., Bailey et al., 2008; Beadle et al., 2012; Fulz & Bernieri, 2022; Grühn et al., 2008). However, most of these studies used a self-report measure to index empathy, and more importantly, none examined whether associations between empathy and broader social functioning differ as a function of the measurement approach used to index empathy.
The Present Study
Measurement type plays an important role in the evaluation of personality traits, and this appears to also be the case for empathy. However, no study to date has directly compared the three available assessment approaches in relation to both cognitive and affective empathy. We aimed to address this gap by examining how evaluations of empathy are influenced by measurement type. First, using a large, adult lifespan sample, we aimed to cross-validate prior work that has identified relatively weak associations between self-report and behavioral cognitive empathy assessments, and to also extend this work by concurrently examining these associations in the context of affective empathy. Second, we aimed to establish whether self-report and informant-report empathy assessments show reasonable convergent validity for cognitive and affective empathy. Third, we aimed to examine whether informant-report is accurate in predicting an individual’s behavioral task performance and whether this differs as a function of empathy type (i.e., cognitive v affective). Fourth, we aimed to assess whether any observed relationships between empathy and broader social function differ according to the measurement approach used to index empathy. Finally, we aimed to test whether any of the observed relationships between the empathy assessments differ as a function of participant age. This latter question was of interest considering broader literature showing that cognitive and affective empathy might be differentially impacted by normal adult aging (see Beadle & de la Vega, 2019).
Hypotheses
Based on prior research (Israelashvili et al., 2019; Melchers et al., 2015; Murphy & Lilienfeld, 2019; Roth & Altmann, 2021; Stosic et al., 2022), we predicted that self-report would be poorly correlated with behavioral performance for cognitive empathy. For affective empathy, self-report and behavioral measures were expected to be correlated based on prior work that has shown convergence between these two assessment types (Van der Graaff et al., 2016; Westbury & Neumann, 2008). For informant-reports, we expected a weak correlation with behavioral performance for cognitive empathy given its low trait visibility. In contrast, for affective empathy, we expected informant-report to be significantly correlated with the behavioral performance given its high trait visibility. For all analyses involving the behavioral affective empathy task, we expected any observed relationships to be strongest for negative as opposed to positive emotions. This is because feelings of care or concern (which is primarily indexed via the IRI) are typically experienced when observing others in unfortunate situations (as opposed to happy uplifting situations). Prior work in the broader person perception literature has shown that self- and informant-report measures correlate modestly; therefore, we expected to find moderate correlations between these two assessment types for both components of empathy. We expected empathy to be correlated with broader social functioning outcomes, but it was unclear whether the strength of these associations would differ according to measurement type. The investigation of age effects was exploratory; therefore, we made no a priori predictions regarding the moderating role of participant age.
Method
We report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study. This study was part of a larger research program that included other assessments but only the measures relevant to this study are reported here.
Participants
We conducted a priori power analyses using G*Power (Faul et al., 2007) to determine the minimum sample sizes required to detect (a) a moderate correlation and (b) a moderate-sized interaction between age and measurement type in our regression model. These revealed that a minimum of 134 participants were required to detect a moderate correlation (r = .30), and a minimum of 119 participants was required to detect a moderating effect of age, with 95% power. However, we exceeded this sample size and ended up including a total of 371 participants in this study. Participants ranged in age from 18 to 101 years (M = 49.10, SD = 24.82), and 65% of the sample was female. The younger and middle-aged adults (i.e., <60 years) were recruited from the local Brisbane community, and all older adults (60+ years) were recruited from the Sydney Memory and Aging Study (MAS; Sachdev et al., 2010), The Sydney Centenarian Study (SCS; Sachdev et al., 2013), and the local Sydney community. Community participants were recruited via numerous sources including (a) advertisements on social media, university websites, and interest groups, (b) recruitment flyers at local community events, and (c) word of mouth from other participants. Participants were excluded if they had a current neurological or psychiatric illness, or if they could not speak English fluently. All participants aged 60 years of age and older were screened for abnormal cognitive decline using the Mini-Mental State Exam (MMSE, Folstein et al., 1975), and all scored above the recommended cut-off of 26, resulting in no exclusions for this reason.
Measures
Behavioral Assessment
The Multifaceted Empathy Test (MET, Foell et al., 2018) was used to provide a behavioral assessment of cognitive and affective empathy. The MET includes images of people displaying a range of different emotional states in naturalistic contexts. An important strength of the MET is that provides a concurrent assessment of both cognitive (MET-Cog) and affective (MET-Aff) empathy. To index cognitive empathy, participants were required to select the label that best described how the person in the image was thinking or feeling from a list of four possible emotion response options. Responses were scored as correct or incorrect, and the total score (out of a possible 40) was converted into a percentage to reflect the overall accuracy in the task. For affective empathy, participants were asked to state how much they empathized with the individual in each photograph, on a scale of 1 (not at all) to 9 (very much). Because the word “empathize” can have a number of different meanings, we explained to participants (verbally and via standardized instructions on the computer screen) that “empathize” in this study refers to the degree to which seeing someone expressing a particular emotion makes them feel that emotion themselves (e.g., the extent to which they feel sad when seeing someone else feeling sad). The trials were organized into eight blocks of 10 such that participants viewed a block of cognitive empathy trials followed by a block of affective empathy trials until all blocks were completed. All trials were presented in a randomized sequence within each block. The MET is comprised of 80 trials in total (40 for each empathy subtype), and there was no time limit to complete the task. We calculated an average empathy rating score for all affective empathy trials in the MET with higher scores indicating higher overall empathy.
Self-Report Assessment
The IRI (Davis, 1983) is the most widely used self-report empathy task to date (Hall & Schwartz, 2019) and therefore was used to index self-reported empathy in this study. Although the IRI includes four 7-item subscales, only the two relevant to this study were used. The perspective-taking subscale (IRI-PT self) was used to index cognitive empathy and includes questions such as “I believe there are two sides to every question and try to look at them both.” The empathic concern subscale (IRI-EC self) was used to index affective empathy and includes questions such as “I would describe myself as a pretty soft-hearted person.” All responses were made on a 5-point scale that ranged from 0 “Does not describe me well” to 4 “Describes me very well,” and responses were averaged to create a mean score for each empathy component. Higher scores indicated greater empathic capacity. The factor structure for this questionnaire is reported in Supplementary Materials (S1).
Informant-Report Assessment
The IRI (Davis, 1983) was also used to index informant-reported empathy. Consistent with the self-report measure, we only used the perspective-taking and empathic concern subscales. The informant version of the IRI was originally adapted from the self-report version by Cliffordson (2001) for use with parent populations but has since been adapted and used in adult populations (see e.g., Chander et al., 2021; Dermody et al., 2016; Rankin et al., 2005). The questions and responses are identical to the self-report version except that they are worded in the third instead of the first person. For example, “The participant believes there are two sides to every question and tries to look at them both” (IRI-PT informant), and “I would describe the participant as a pretty soft-hearted person” (IRI-EC informant). Participants were given an envelope with a letter to the informant explaining the reason for their involvement in the study, along with the questionnaire and a reply-paid envelope, and were told to ask someone who knows them well (i.e., a close friend or family member) to complete the questionnaire on their behalf and post it back to the research team. Consistent with the self-report assessment, total scores were created by averaging the ratings for the 7 items in each subscale. Higher scores indicated greater empathic capacity. The factor structure for this questionnaire is reported in Supplementary Materials (S1).
Social Function Measures
Social Engagement
The Lubben Social Network Scale (Lubben, 1988) is a self-report questionnaire that indexes social engagement with family and friends. It includes 12 questions in total that enquire about an individual’s relationships with their family and friends. All responses were made on a 6-point scale, and two separate total scores were calculated for friends and family. Higher scores indicated greater social engagement.
Social Behavior
The Peer-Report Social Functioning Scale (Henry et al., 2009) is an informant-rated measure that provides an assessment of one’s social behavior. It includes three subscales that enquire about the degree to which an individual engages in behavior that is (a) socially appropriate (e.g., How often does the participant speak positively about others?) (b) socially inappropriate (How often does the participant embarrasses people unintentionally?) and (c) prejudicial (How often does the participant comment on someone else’s race in a negative way?). All responses were made on a 4-point scale (never, rarely, occasionally, and frequently), with some questions reverse scored. Higher scores indicate higher levels of socially appropriate, inappropriate, and prejudicial behaviors.
Procedure
After providing written informed consent, older adults recruited from the community completed the MMSE. All older adults recruited from ongoing longitudinal aging studies (i.e., MAS and SCS) had recently completed the MMSE as part of a separate study, and all scored above the recommended cutoffs for cognitive decline; therefore, the MMSE was not administered again for these participants. Next, all participants completed assessments of mood and provided basic demographic information. Participants were randomly assigned to complete one of four counterbalanced orders of a large testing battery that included a series of well-validated cognitive and social cognitive tasks. This testing battery took between 3 and 4 hours to complete, and all participants were offered brief breaks throughout the testing session. Upon completion of all tasks, participants were provided with the informant assessments along with instructions explaining how these should be completed. All participants were compensated with $60 at the end of the testing session. This study was approved by the University of Queensland Human Research Ethics Committee (Approval number: 2017000770).
Analytic Approach
Multitrait-Multimethod Analysis
We investigated the assessment of two theoretically distinct constructs—cognitive empathy and affective empathy—via three measurement methods—behavioral assessment, self-report ratings, and informant-report ratings. We therefore measured six variables (see Figure 1 for a schematic representation).

Schematic Representation of How the Measures Collected in This Study Relate to the Three Methods of Assessment and the Two Theoretically Distinct Empathy Components
To assess the relationships between these variables we constructed a multitrait-multimethod matrix (MTMM; Campbell & Fiske, 1959) in R using the Hmisc package (Harrell, 2021). This matrix contains three distinct sets of correlations. The first is homotrait–heteromethod values which reflect the correlations between indices of the same trait collected using different methods (e.g., behavioral cognitive empathy and self-report cognitive empathy). The second is heterotrait–homomethod values which are the correlations between indices of different traits collected using the same method (e.g., behavioral cognitive empathy and behavioral affective empathy). The third is heterotrait–heteromethod values which are the correlations between indexes of different traits using different methods (e.g., behavioral cognitive empathy and self-report affective empathy).
Three key psychometric assessments are discernible via the assessment of the pattern of relationships across these sets of values (see Campbell & Fiske, 1959). To extract the required outcomes (averages of each set of values), we used the MTMM function within the R package Multicon (Sherman, 2015). The first is convergent validity, a measure of the capacity for multiple measures of the same thing to converge on a single value. Convergent validity is present when variables that aim to measure the same trait have a strong correlation (i.e., significant homotrait–heteromethod values). The second is discriminant validity that refers to the capacity for a measure to adequately discriminate between the theoretically distinct traits it aims to measure. Discriminant validity is present among the measurement types when the heterotrait–heteromethod values are not high, especially relative to the homotrait–heteromethod values. The third is method variance referring to variation in variables attributable to measurement type. If method variance is not present, the heterotrait–homomethod values are not different from the heterotrait–heteromethod values. If measurement variance is present, the heterotrait–homomethod values will be higher than the heterotrait–heteromethod values.
Pearson Correlations
Prior work has shown that emotional valence can influence the degree of convergence between affective empathy measures (Foell et al., 2018). Therefore, we separated scores on the MET-Affective according to valence and then conducted Pearson correlations between scores on the MET-Affective and the self- and informant-report measures and the MET-Cognitive, separately for positive and negative trials.
We were also interested in assessing the degree to which the three empathy measures were related to broader social function measures. We conducted Pearson correlations and adjusted for multiple comparisons using Bonferroni correction.
Moderated Regression
We used moderated regression to test whether any of the relationships between the different empathy measures differed as a function of age. This was achieved using PROCESS, Model 1 (Hayes, 2018). Age was used as a continuous moderator in all analyses because trichotomizing this variable (i.e., into three separate age-groups) would lead to substantial reductions in statistical power (see e.g., MacCullum et al., 2002; Maxwell & Delaney, 1993). Predictor variables and age were mean-centered for all analyses.
Results
Descriptive Statistics
The total sample mean scores and standard deviations for each of the three assessments are reported in Table 1.
Descriptive Statistics for Performance on Each of the Three Assessment Types, Presented Separately for Cognitive and Affective Empathy.
Note. Across all tasks, higher scores indicate greater empathy. Some participants did not complete all three assessments which resulted in different samples sizes across the three measurement types. Ten very old adults (i.e., 90 years+) completed an abbreviated testing battery that did not include the behavioral assessment (i.e, MET) and two older adults did not complete the MET-Cog. There was some attrition in the informant-report assessments, which resulted in reduced sample sizes for this assessment type. IRI-PT (self) = self-report perspective taking subscale of the IRI; IRI-PT (informant) = informant-report perspective-taking subscale of the IRI; MET-Cog = cognitive empathy component of the MET; IRI-EC (self) = self-report empathic concern subscale of the IRI; IRI-EC (informant) = informant-report subscale of the IRI; MET-Aff = affective empathy component of the MET; M = mean, SD = standard deviation.
MTMM Analysis
Table 2 presents the MTMM matrix.
MTMM Matrix Comprising the Pearson Correlations Among the Six Assessed Empathy Variables.
Note. The homotrait–heteromethod values are presented in bold. The heterotrait–heteromethod values are italicized. The remaining correlations are the heterotrait–homomethod values. MTMM = multitrait-multimethod.
p < .001, **p < .01, *p < .05.
Convergent Validity
The average correlation for the homotrait–heteromethod values was small (r = .13). Only two of the six homotrait-heteromethod values were significantly different from zero. These both involved affective empathy and included (a) the correlation between self- and informant-report affective empathy and (b) self-report and behavioral affective empathy. These two significant correlations were small-to-moderate in magnitude (both rs = .21).
Discriminant Validity
The average correlation for the heterotrait–heteromethod values was small (r = .10). It was also smaller than the average correlation for the homotrait–heteromethod values. Again, only two out of the six heterotrait–heteromethod values were significant. These included (a) the correlation between self-report affective empathy and behavioral cognitive empathy and (b) the correlation between self-report cognitive empathy and behavioral affective empathy. The former correlation was small in magnitude (r = .11) and the latter was small-to-moderate in magnitude (r = .23).
Method Variance
The average correlation for the heterotrait–homomethod values was moderate (r = .34). Two of the three heterotrait–homomethod values were significant. These included the correlations between (a) self-report cognitive and affective empathy and (b) informant-report cognitive and affective empathy. Both were moderate in magnitude (rs = .46 and .59, respectively).
Correlations With Behavioral Affective Empathy Separately by Valence
Pearson correlations between the behavioral affective empathy task separately for positive and negative trials, and (a) behavioral cognitive empathy, (b) self-report cognitive and affective empathy, and (c) informant-report cognitive and affective empathy, are reported in Table 3. These analyses revealed that the correlation with self-reported affective empathy was significant for negatively valenced but not positively valenced trials. We tested whether the strength of these correlations differed significantly using Fischer’s z test. This revealed that the correlation between behavioral and self-reported affective empathy was significantly stronger for negatively valenced compared with positive valenced trials (z = 2.35, p < .01). The behavioral affective empathy task was also significantly correlated with self-reported perspective taking for both positive and negative trials. However, the correlation was significantly stronger for negatively valenced trials (z = 1.67, p = .047).
Pearson Correlations Between Behavioral Affective Empathy and (a) Behavioral Cognitive Empathy, (b) Self-Report Empathy, and (c) Informant-Report Empathy, Presented Separately by Emotional Valence.
Note. MET-Aff = affective empathy component of the MET; MET-Cog = cognitive empathy component of the MET; IRI-PT (self) = self-report perspective-taking subscale of the IRI; IRI-EC (self) = self-report empathic concern subscale of the IRI; IRI-PT (informant) = informant-report perspective-taking subscale of the IRI; IRI-EC (informant) = informant-report subscale of the IRI.
p < .01, *** p < .001.
Correlations Between the Empathy and the Social Function Measures
Pearson correlations were conducted between the three empathy measures and the social function measures, separately for cognitive and affective empathy. As can be seen in Table 4, although a few weak correlations were identified with some of the social function measures and behavioral and self-reported empathy, most correlations involving these two assessment types were not significant. However, for the informant-rated assessment of empathy, significant correlations—ranging from small to large in magnitude—emerged with all five social function measures.
Pearson Correlations Between Measures of Empathy and Social Functioning.
Note. MET-Cog = cognitive empathy component of the MET; MET-Aff = affective empathy component of the MET; IRI-PT (self) = self-report perspective-taking subscale of the IRI; IRI-PT (informant) = informant-report perspective-taking subscale of the IRI.
p <.05. ** p <.01. *** p <.001.
Moderated Regression Analyses
Cognitive Empathy
We examined whether self-reported cognitive empathy scores were associated with behavioral cognitive empathy scores and whether age moderated this relationship. The overall model was significant, R2 = .02, F(3, 354) = 2.68, p = .047. IRI-PT (self-report) was not associated with MET-Cog scores, b = .34, p = .195, nor was age, b = −.01, p = .118. The interaction was not significant, b = .02, p = .089, indicating no moderating effect of age. Next, we examined whether self-reported cognitive empathy scores were associated with informant-reported cognitive empathy scores, and again whether age was a moderator in this relationship. The overall model was significant, R2 = .12, F(3, 198) = 9.37, p < .001. IRI-PT (self-report) was significantly associated with IRI-PT (informant-report) scores, b = .20, p =.027, as was age, b = .01, p < .001. However, the interaction was not significant, b = 0.00, p = .201, suggesting no moderating effect of age. Finally, we examined whether informant-reported cognitive empathy scores were associated with behavioral task performance (i.e, MET-Cog). The overall model was not significant, R2 = .02, F(3, 191) = 1.46, p = .227. IRI-PT (informant-report) was not associated with MET-Cog scores, b = .31, p = .249, nor was age, b = −.02, p = .056. The interaction was not significant, b = .01, p = .576, again suggesting no moderating effect of age.
Affective Empathy
The same analytic approach outlined above was used for affective empathy. We first examined whether self-reported affective empathy was associated with behavioral task performance, and whether this relationship was moderated by participant age. The overall model was significant, R2 = .14, F(3, 357) = 18.62, p <. 001, and both IRI-EC (self-report), b = .57, p < .001, and age, b = .02, p < .001 were significantly associated with MET-Aff scores. However, the two-way interaction was not significant, b = 0.00, p = .826, suggesting no moderating effect of age. Next, we examined whether self-reported affective empathy was associated with informant-reported affective empathy, and whether this relationship was moderated by participant age. The overall model was significant, R2 = .12, F(3, 198) = 8.64, p < .001. IRI-EC (self-report) was significantly associated with IRI-EC (informant-report), b = .26, p < .001, as was age, b = .01, p <.001. The interaction was not significant, b = −0.00, p = .513, indicating no moderating effect of age. The final set of analyses focused on testing whether informant-report affective empathy was associated with behavioral task performance (i.e., MET-Aff), and again whether age moderated this relationship. The results revealed that the overall model was significant, R2 = .07, F(3, 193) = 4.58, p = .004. IRI-EC (informant-report) was not associated with MET-Aff performance, b = .10, p = .521, but age was, b = .01, p <.001. However, the interaction was not significant, b = −.00, p = .673, again indicating no moderating effect of age.
Discussion
Empathy plays a fundamental role in social well-being. It is therefore not surprising that extensive literature has now emerged focused on trying to better understand empathic function at all stages of the human lifespan. We used three widely used measures of empathy, with the goal of providing the first direct test of whether these three measurement approaches provide consistent evaluations of cognitive and affective empathy across the adult lifespan.
In line with our prediction, there was no significant correlation between the self-report and behavioral measures of cognitive empathy. This finding aligns with a growing literature identifying poor convergence between these two forms of assessment for cognitive empathy (Foell et al., 2018; Melchers et al., 2015; Murphy & Lilienfeld, 2019; Roth & Altmann, 2021; Stosic et al., 2022). It has been argued that the IRI-PT measures one’s motivation to adopt someone else’s perspective and may in fact be more like a measure of affective empathy than cognitive empathy (Murphy et al., 2020; Murphy & Lilienfeld, 2019). Consistent with this idea, we found that the IRI-PT was significantly correlated with the behavioral affective empathy task, despite not being related to the behavioral cognitive empathy task. Together, these findings provide further evidence that the IRI-PT is not the best indicator of cognitive empathy ability and adds to growing concerns regarding the use of the IRI-PT as a proxy for cognitive empathy ability (Murphy & Lilienfeld, 2019; Stosic et al., 2022). For affective empathy, the self-report and behavioral empathy measures were significantly correlated, and this was in line with our prediction. Although this correlation was only small-to-moderate in magnitude, it aligns with prior work that has shown convergence between self-report and behavioral affective empathy measures (Van der Graaff et al., 2016; Westbury & Neumann, 2008). We examined whether this relationship differed as a function of valence in the behavioral measure. In line with our prediction, we found the correlation was indeed strongest for responses to negatively valenced trials. As explained earlier, the IRI-EC taps into feelings of care and concern, which often arise in response to seeing someone else experiencing an unfortunate situation. It therefore makes sense that the IRI-EC was more strongly associated with responses to negatively valenced stimuli than positive stimuli (although see Foell et al., 2018). Together, these findings demonstrate that the disconnect identified between self-report and behavioral measures may be unique to the cognitive component of empathy.
One possible explanation for the greater convergence between self-report and behavioral measures for affective empathy may be that these two assessment approaches have shared measurement error. Indeed, the MET and IRI have similar response formats (i.e., self-report) for affective empathy. With such assessments, social desirability biases can also emerge (see Paulhus, 1986), and prior work has shown that people who wish to be perceived by others in a favorable way are more likely to provide inflated ratings of their empathic capacity (Sassenrath, 2019). It could be argued that the affective empathy component of the MET is not a true behavioral task because it requires a person to self-report how certain stimuli make them feel rather than an objectively index how they are feeling. However, this is a broad limitation of many of the currently available lab-based affective empathy tasks—they involve making self-ratings about emotional information (e.g., Beadle & de la Vega, 2019; Kanske et al., 2015; Kuypers, 2017). However, it is noteworthy that prior work has established convergence between physiological emotional responses (i.e., facial muscle responding) and self-reported empathy (Drimalla et al., 2019; Westbury & Neumann, 2008), suggesting that similarities in response types may not be the only reason for the higher convergence between affective empathy measures. Future work should include more objective ratings of affective empathic responding, such as tasks that measure actual empathic behavior in real-life situations, to provide a more nuanced understanding of the relationship between self-report and behavioral affective measures. Despite the evidence for convergence between the self-report and behavioral measures of affective empathy, there was also evidence of poor discriminant validity within the three assessment types. This is because, as noted earlier, a small-to-moderate association was also identified between the behavioral affective empathy task and the self-reported cognitive empathy task, while a smaller but still significant relationship also emerged between the behavioral cognitive empathy task and the self-reported affective empathy task. The self-report and informant-report measures of cognitive and affective empathy were also correlated, and these correlations were moderate-to-large in magnitude. Because there was no evidence of poor discriminant validity between the behavioral and informant-report tasks (for either empathy component), the most parsimonious explanation of this validity problem is that it is being driven by the self-report measure (i.e., the IRI). As noted, concerns regarding the validity of the IRI have been identified previously (Murphy et al., 2020; Siu & Shek, 2005), and the current findings provide further evidence that the self-report variant of the IRI, in particular, may not be sensitive at indexing distinct subcomponents of empathy.
The second aim of this study was to establish convergence between self-report and informant-report assessments of empathy. In line with findings from the broader person perception literature, we predicted moderate correlations between these two assessment types for both components of empathy. In contrast with this prediction, we found no significant correlation between self- and informant-report for cognitive empathy. However, for affective empathy, these two assessment types were significantly correlated (r = .21) providing some evidence for convergent validity. As noted earlier, trait visibility is an important consideration for self- and informant-report measures of personality (Vazire, 2010). Indeed, traits with high visibility are often best judged by an independent observer, whereas traits with low visibility are often better judged by the self (Vazire, 2010). Given that cognitive empathy is a relatively internal process (i.e., as it involves thinking about what someone else is feeling), it might be more difficult for an informant to judge, and this might explain the poorer agreement between the self and the informant measures for cognitive relative to affective empathy. However, this was the first study to examine self-other agreement separately for the two empathy subcomponents, therefore we encourage future studies to extend on this work by replicating these findings with other validated measures of empathy.
The third aim of this study was to assess whether informant-reports are accurate in predicting an individual’s behavioral task performance. We expected informant-reports to be associated with behavioral performance for affective but not cognitive empathy due to differences in trait visibility. However, we found no relationship between these two assessment types for either empathy type, suggesting informant-reports may not be appropriate for predicting behavioral performance. This was particularly surprising for affective empathy, given it is arguably a highly visible trait that should be more easily observed by informants. A closer inspection of the instructions provided to participants suggests that the behavioral and informant (and self-report) affective empathy measures may have tapped into different aspects of affective empathy. This is because the IRI enquires about one’s propensity to experience care and concern toward another person, whereas the affective empathy component of the MET assesses the degree to which one shares the emotional state of another person (i.e., affective resonance). These two measures are therefore not conceptually equivalent. Nevertheless, because a central tenet of prominent models of empathy is that emotion sharing is a precursor to higher level feelings of care and concern (de Waal & Preston, 2017), the MET and IRI should theoretically be indirectly conceptually related, and indeed, the self-report IRI and MET were correlated for affective empathy. As noted earlier too, the MET and IRI both require self-reporting and therefore have shared measurement error, and this might explain why we found better convergence for the self- vs informant-report and behavioral affective empathy measures.
The fourth aim of this study was to assess the degree to which empathy is related to broader social functioning and whether measurement type is important in determining this relationship. The results showed that, of the three measurement approaches, the informant-rated assessments were most consistently and strongly associated with levels of social engagement and social behavior, with the strongest correlations observed between socially appropriate behavior and informant-rated cognitive and affective empathy. As noted earlier, self- and informant-reports are differentially affected by certain favorability biases in responding such that self-report is associated with a bias to emphasize competence, and informant-report is associated with a bias to emphasize cooperative traits (Ashton & Lee, 2010). Given that, like the informant-report IRI, the Peer Report Social Functioning Scale was also completed by an informant, similar response biases and consequently shared measurement error may have contributed to the high correlations between these particular assessments. Together, these findings are interesting and potentially very important, as they suggest that an independent observer may in fact provide the most accurate reflection of an individual’s empathic capacity.
Finally, we aimed to test whether any of the observed relationships were moderated by participant age. The results showed that age did not moderate any relationships between any of the measures, which indicates that inconsistencies in measurement across empathy tasks are not a problem exclusive to a particular age-group but instead an issue across the lifespan. As noted earlier, the findings in the empathy and aging literature are relatively mixed, and the results from this study indicate that these inconsistencies could indeed be due to the differences in approaches used to measure empathy.
Limitations
Although this study was the first to compare three different empathy measurement approaches in a lifespan sample, there are some limitations that need to be acknowledged. First, we conducted a confirmatory factor analysis (CFA) to establish convergent and discriminant validity (see Supplementary Material 2) and found very little evidence to support a two-factor model of empathy. Therefore, although the measures included in this study were theoretically designed to tap into two distinct empathy subcomponents, they failed to do so, which indicates a problem with discriminate validity within these measures. Indeed, the strongest correlations between measures were observed when the same method was used to measure different traits (e.g., informant-report cognitive and affective empathy), which suggests that variance due to measurement type is largely contributing to the poor validity observed between these measures. However, the CFA also revealed that a single-factor model of empathy did not provide a strong fit for the data either, which again points to concerns regarding the validity of these measures. The key point to take away from these results is therefore that three empathy measures used in this study appear to have been tapping into different aspects of empathic processing, and therefore, careful consideration should be made when selecting an empathy measure.
The smaller sample available for the comparisons that included the informant-rated measure also needs to be acknowledged as a limitation. However, missing data is common when using informant-report measures and the level of attrition in our study is similar to previous work (e.g., Roth & Altmann, 2021). Another limitation of our study design was the use of a single informant as opposed to multiple informants to index empathy. Prior work has shown that the level of agreement between self and informant ratings increases with more informants used (see Connelly & Ones, 2010; Roth & Altmann, 2021). Because the use of a single informant in our study may have underestimated the strength of any relationships involving the informant measure, it would be valuable for future work to cross-validate these findings with multiple informants to establish whether the strength of any of the observed relationships changes. However, for some people (and in particular older adults), access to multiple informants may be challenging; therefore, the use of a single informant in this study might be regarded as a strength because it allowed us to collect responses from a more socially diverse sample, including those with smaller social networks and those with rich social networks. Another limitation of this study was that we were unable to test the moderating role of informant type. Participants were asked to ensure that only someone who knows them well completed the informant questionnaire on their behalf, but we did not collect any details regarding the relationship between the participant and informant. Given that the degree of closeness between an individual and informant is important for agreement between self- and informant-reports in empathy and personality traits more broadly (Ashton & Lee, 2010; Connelly & Ones, 2010; Lee & Ashton, 2017), an interesting avenue for future research is to test whether relationship type and/or reported closeness moderates the degree of agreeance between empathy measures. Finally, a single measure was used to index behavioral empathy rather than a battery of behavioral tasks. Although the measure used (the MET) is widely regarded as a validated measure of this type of empathy (see e.g., Henry et al., 2016), the static nature of the MET’s stimuli clearly limits its ecological validity and means that it may not evoke the same type or strength of empathic response as a real-world social encounter. Indeed, prior work has revealed considerable variation in tasks that assess interpersonal accuracy (Schlegel et al., 2017), and therefore, it remains to be established whether the results observed with the MET in this study replicate when a different empathy measure is used that includes different task features (e.g., multimodal stimuli). As others have suggested (see Murphy & Lilienfeld, 2019), a battery of behavioral empathy measures that can be aggregated to form latent variables would provide the most accurate assessment of one’s empathic capacity, and we would encourage future studies to consider this approach rather than relying on single empathy measures.
Conclusion
To conclude, this study replicates prior work but also meaningfully extends the current understanding of the role of measurement type in the assessment of empathy. The results showed that affective empathy measures, in general, show better convergence than cognitive empathy measures and that this is true across the entire adult lifespan. However, despite evidence for convergence within the affective empathy measures, the results also showed that there is a lack of discriminant validity for both empathy components. In particular, the IRI showed the poorest discriminant validity, aligning with prior work suggesting that the EC and PT subscales may not be adequate for measuring distinct empathy subcomponents (Murphy et al., 2020; Siu & Shek, 2005). Importantly though, out of all three assessment approaches, informant-reports were most consistently correlated with actual social functioning, suggesting that this approach may be particularly valuable, especially in clinical settings where poor social function has been linked to many important prognostic outcomes. However, the key conclusion to emerge from this study is that careful consideration must be taken when choosing an empathy measure for research or clinical purposes. This is because the three currently available approaches for measuring empathy appear to provide quite distinct insights into empathic capacity, and this is particularly the case for cognitive empathy. Given the large variability among empathy measures, we would encourage researchers and clinicians to avoid relying on a single approach for measuring empathy, and in any situation where empathy assessments are being used to inform decisions that directly impact upon the individual being assessed (such as vocational or clinical environments), a multi-method assessment approach may be warranted.
Supplemental Material
sj-docx-1-asm-10.1177_10731911221127902 – Supplemental material for Measuring Empathy Across the Adult Lifespan: A Comparison of Three Assessment Types
Supplemental material, sj-docx-1-asm-10.1177_10731911221127902 for Measuring Empathy Across the Adult Lifespan: A Comparison of Three Assessment Types by Sarah A. Grainger, Kate T. McKay, Julia C. Riches, Russell J. Chander, Rhiagh Cleary, Karen A. Mather, Nicole A. Kochan, Perminder S. Sachdev and Julie D. Henry in Assessment
Supplemental Material
sj-docx-2-asm-10.1177_10731911221127902 – Supplemental material for Measuring Empathy Across the Adult Lifespan: A Comparison of Three Assessment Types
Supplemental material, sj-docx-2-asm-10.1177_10731911221127902 for Measuring Empathy Across the Adult Lifespan: A Comparison of Three Assessment Types by Sarah A. Grainger, Kate T. McKay, Julia C. Riches, Russell J. Chander, Rhiagh Cleary, Karen A. Mather, Nicole A. Kochan, Perminder S. Sachdev and Julie D. Henry in Assessment
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by an Australian Research Council Discovery Project Grant (grant no. DP1701011239). S.A.G. is the recipient of an Australian Research Council Discovery Early Career Researcher Award (grant no. DE220100561). J.D.H. was supported by an Australian Research Council Future Fellowship (grant no. FT170100096).
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
