Abstract
The Rosenberg Self-Esteem Scale is the most frequently used measure of self-esteem in the social sciences. These items are often administered with a different number of response options, but it is unclear how the number of response options impacts the psychometric properties of this measure. Across three experiments (Ns = 739, 2,358, and 1,461), we evaluated how different response options of the Rosenberg influenced (a) coefficient alpha estimates, (b) distributions of scores, and (c) associations with criterion-related variables. Observed coefficient alpha estimates were lowest for a 2-point format compared with response formats with more options. However, supplemental analyses using ordinal alpha pointed to similar estimates across conditions. Using four or more response options better approximated a normal distribution for observed summary scores. We found no consistent evidence that criterion-related correlations increased with more response options. Collectively, these results suggest that the Rosenberg should be administered with at least four response options and we favor a 5-point Likert-type response format.
Keywords
Global self-esteem is a widely studied construct in the social sciences and is positively associated with adjustment in a wide range of life domains including close relationships, psychological health, school, and work (see Orth & Robins, 2022 for a review). The 10-item Rosenberg Self-Esteem Scale (RSES; Rosenberg, 1989) is the most frequently used measure of self-esteem in the literature (Blascovich & Tomaka, 1991; Donnellan et al., 2015; Pegler et al., 2019; Zeigler-Hill, 2013). This measure was originally developed using a Guttman scale (see Chapter 2 and Appendix D in Rosenberg, 1989), but the items are now widely used with Likert-type response scales (e.g., Blascovich & Tomaka, 1991; Donnellan et al., 2015) such as 4-point (e.g., Heatherton & Wyland, 2003; Schmitt & Allik, 2005), 5-point (e.g., Gray-Little et al., 1997), and even 11-point (e.g., Leung, 2011) versions. Given the relevance of self-esteem for a wide range of domains in the social and behavioral sciences and the considerable variability in response options used with the RSES (Pegler et al., 2019), it is important to evaluate how the number of response options influences the psychometric properties of this popular measure.
Several recent studies have evaluated how the number of response options influences the psychometric properties of personality and attitude measures (Adelson & McCoach, 2010; Cox et al., 2017; Dawes, 2008; Finn et al., 2015; Jacoby & Matell, 1971; Leung, 2011; Matell & Jacoby, 1971; Müssig et al., 2021; Simms et al., 2019). This literature suggests that coefficient alpha estimates of internal consistency tend to increase as the number of response options increases from two to about five or six response options (Simms et al., 2019). For example, Finn et al. (2015) found small increases in alpha estimates when expanding from the traditional True/False format for MMPI-2-RF items to an extended four option scale. Simms et al. (2019) found increases in alpha from 2-point (e.g., Agree versus Disagree) to 6-point (e.g., Strongly Agree to Strongly Disagree) response versions of a big five personality trait measure. There was little improvement beyond six response options. A similar pattern was reported by Müssig et al. (2021) for measures of the big five personality traits and emotion regulation.
Researchers might expect that increasing alpha estimates will generate stronger correlations between scale composites and other variables because of assumed reductions in measurement error. After all, measurement error attenuates correlations with criterion variables (see, for example, Schmidt & Hunter, 1996). Counter to this expectation, existing studies tend to find the correlations with criterion-related variables are similar regardless of the number of response options for a given item pool (e.g., Cox et al., 2017; Finn et al., 2015; Simms et al., 2019). Finn et al. (2015) referenced increasing systematic error with more response options as a possible explanation of these findings. Systemic error variance (see, for example, Clifton, 2020; Streiner, 2003) captures consistent ways of responding across all items that are unrelated to an individual’s standing on the latent variable. Accordingly, there is no actual reduction in measurement error with more response options under this systematic error explanation.
In sum, there are reasons to suspect that increasing response options might increase coefficient alpha estimates for the Rosenberg items but have little, if any, influence on criterion-related validity estimates. This generalization seems consistent with the small existing response option literature surrounding the RSES. Pegler et al. (2019) reported a small correlation between estimates of coefficient alpha and the number of response options (r = .12) in their survey of the self-esteem literature in social/personality psychology. Leung (2011) investigated the psychometric properties of a Chinese translation of the RSES across 4-, 5-, 6-, and 11-point response alternatives using a sample of secondary students from Macau. Internal consistency estimates were similar across the 4 versions (min = .76 for the 6-point version; max. = .81 for the 11-point version). No clear pattern of differential criterion-related correlations emerged across the four versions. However, Leung (2011) recommended an 11-point version because summary scores were more normally distributed with 11 response options compared with alternatives. Observed scores were more negatively skewed for the versions with fewer response options.
To be sure, the shape of the distribution of scores as well as measures of central tendency are relevant when considering the impact of response options. Observed scale scores from measures using few response options might produce skewed distributions that provide a truncated perspective on the actual distribution of that variable (see also Dawes, 2008). Thus, there could be reasons to prefer a greater number of response options to the extent that administering the RSES with few response options per item might distort the distribution of scores. However, the 11-point response scale used by Leung (2011) seems unwieldy, especially if researchers wish to label all response options (see Simms et al., 2019). Indeed, Müssig et al. (2021) found that participants rated an 11-point format as having the lowest ease of use compared with 2-, 5-, and 7- response formats for personality measures. Although the 11-point format is rare in the self-esteem literature (Pegler et al., 2019), it is important to test whether this format is superior to versions with fewer response options.
Present Studies
The present studies were designed to evaluate the impact of the number of response options on the psychometric properties of the Rosenberg Self-Esteem scale. Specifically, in Studies 1 and 2, we compared versions with 2-, 3-, 4-, 5-, and 7-point scales (a 6-point condition was added to Study 2); and in Study 3, we compared versions with 2-, 5-, and 11-point scales, respectively. Study 1 was an exploratory effort using college students whereas Study 2 was a preregistered confirmatory study using a more representative sample collected by a market research firm. Study 3 was not preregistered but evaluated the 11-point format against the 2- and 5-point versions using a college student sample. In all three studies, participants completed measures of criterion-related constructs and were then assigned by the survey software to complete the Rosenberg with different response options. Data and R code for each study are available on the OSF website (https://osf.io/ae5gt/?view_only=e9c92f075b6b44478d30110c2f13eb08). The following questions were investigated.
Question 1: Does Increasing the Number of Response Options Increase Estimates of Internal Consistency?
Based on previous work, we expected to observe that estimates of alpha (and thus the average inter-item correlation for the 10 items) would increase as the number of scale points increased, but with an expectation that increases would not be pronounced once the number of scale points reached 5 or 6 (Simms et al., 2019).
Question 2: Does Increasing the Number of Response Options Decrease Estimates of Central Tendency and Reduce Skewness?
Given results from Leung (2011) and the observation that self-esteem scores are negatively skewed (e.g., Baumeister et al., 2003), we expected that self-esteem scores would more closely approximate normal distributions with increasing scale points.
Question 3: Does Increasing the Number of Response Options Impact Criterion-Related Correlations?
Given results from the personality trait literature (e.g., Cox et al., 2017; Finn et al., 2015; Müssig et al., 2021; Simms et al., 2019) and results reported by Leung (2011), we did not expect to see systematic differences in criterion related correlations for the RSES when considering formats with more response options. Measures of the big five were administered in each study and we expected results in line with previous studies. For example, Donnellan and colleagues (2016) reported correlations between the big five and global self-esteem measured by the Rosenberg. Robins et al. (2001) evaluated associations between a single-item measure of self-esteem and the big five domains. Donnellan et al. (2016) reported correlations between the RSES and the big five of .35, .32, .41, −.56, and .15 for extraversion, agreeableness, conscientiousness, neuroticism, and openness, respectively, based on a sample of over 1,100 college students. Robins, Tracy and colleagues (2001) found correlations of .38, .13, .24, −.50, and .17 for these same domains in a sample of over 300,000 internet participants. Life satisfaction was also used as a criterion in each study. Again, we expected results in line with previous findings. Lucas and colleagues (1996) reported concurrent correlations between .54 and .65 for life satisfaction and self-esteem; Donnellan et al. (2016) reported a correlation of .45 from the large college student sample. Measures of the HEXACO domains (Ashton & Lee, 2007) and psychopathic personality traits were collected in Study 1 and Study 3. Ruchensky and Donnellan (2017) reported correlations between the RSES and the HEXACO domains of .03, −.16, .63, .15, .28, −.08, and .15 for honesty/humility, emotionality, extraversion, agreeableness, conscientiousness, openness, and altruism, respectively; as well as correlations between the Rosenberg and psychopathic traits associated with the triarchic model (e.g., Patrick et al., 2009; Patrick & Drislane, 2015) of .54, −.17, and –.43 for boldness, meanness, and disinhibition, respectively.
Study 1 Method
Participants
Participants were college students who passed two data quality check items (a self-reported honesty item and a serious responding item modeled on Aust et al., 2013) as part of an online survey. Data from up to 739 participants were analyzed for this report out of the 829 responses with unique ID numbers from the Department of Psychology research participant pool. Of the 829 responses, 785 participants responded affirmatively to the honesty question and 743 indicated that they took part seriously. A total of 741 participants passed both quality checks but two of those individuals were missing all self-esteem responses and were excluded. The sample was 72.3% women and 26.9% men, with the remaining 0.8% identifying as non-binary, self-describing, or preferring not to disclose. Participants ranged from 18 to 27 years old (M = 19.59, SD = 1.47). The self-identified racial-ethnic breakdown was 74.3% White, 11.0% Asian, 6.8% Black, 3.8% multi-racial, with the remaining 4.1% identifying as a different race, indicating a preference not to disclose, or skipping the question.
Measures
Big Five Domains
The 60-item Big Five Inventory 2 (BFI-2; Soto & John, 2017) measured extraversion (alpha = .87, M = 3.35, SD = 0.74), agreeableness (alpha = .78, M = 3.81, SD = 0.57), conscientiousness (alpha = .84, M = 3.59, SD = 0.66), negative emotionality (alpha = .89, M = 2.99, SD = 0.80), and open-mindedness (alpha = .84, M = 3.62, SD = 0.67) using 12 items each. The BFI-2 was administered with a 5-point response scale.
Life Satisfaction
The five-item Satisfaction with Life Scale (SWLS; Diener et al., 1985; alpha = .88, M = 4.79, SD = 1.27) measured this attribute with a 7-point response scale.
HEXACO Domains
The HEXACO-100 (HEXACO; Lee & Ashton, 2018) measured honesty-humility (alpha = .78, M = 3.26, SD = 0.51), emotionality (alpha = .84, M = 3.47, SD = 0.60), extraversion (alpha = .86, M = 3.25, SD = 0.60), agreeableness (alpha = .82, M = 3.04, SD = 0.53), conscientiousness (alpha = .83, M = 3.55, SD = 0.54), and openness (alpha = .80, M = 3.19, SD = 0.56). Each domain scale has 16 items and responses were made on a 5-point scale. This inventory also includes a four-item altruism scale (alpha = .55, M = 3.85, SD = 0.64).
Psychopathic Traits
The Triarchic Psychopathy Measure (TriPM; Patrick, 2010) measured disinhibition (alpha = .87, M = 1.78, SD = 0.43), meanness (alpha = .90, M = 1.66, SD = 0.47), and boldness (alpha = .79, M = 2.57, SD = 0.40). Boldness and meanness were each assessed using 19 items and disinhibition was assessed using 20 items. Items were rated on a 4-point scale (1 = False, 2 = Mostly false, 3 = Mostly true, 4 = True).
Procedure
Participants completed the survey online in exchange for extra credit or course credit. The survey was created and collected using the Qualtrics platform. Personality measures were completed in a fixed order (HEXACO, TriPM, SWLS, and the BFI-2). Participants were then assigned to complete the Rosenberg (1989) items in one of five conditions: 2-, 3-, 4-, 5-, or 7-point response options. The exact wording for each set of response options is in the online supplement (Table S1). After completing the self-esteem items, a survey quality item was presented (“How would you rate the quality of this survey”) with 6 response options (Very poor; Poor; Fair; Good; Very good; Excellent). Demographic questions were then presented. The last two items were a yes/no “I have answered all questions honestly” item and the Aust et al. (2013) seriousness check item. The exact wording of the seriousness check item is the online supplement with response options—“I have taken part seriously” or “I have just clicked through, please throw my data away.” Again, participants had to complete the self-esteem items, answer yes to the honesty question, and indicated that they took part seriously to be included in Study 1.
Study 1 Results
Does Increasing the Number of Response Options Increase Estimates of Internal Consistency?
Estimates of coefficient alpha were calculated within each condition and reported in Table 1 along with the average inter-item correlation. The test of the equality of alpha coefficients was evaluated using the R package cocron (Diedenhofen & Musch, 2016). The null hypothesis was rejected (chi-square = 27.87, df = 4, p < .001). A follow-up test dropping the 2-point condition yielded a marginal result (p = .0536). Thus, it appeared that internal consistency estimates were lower for the 2-point response version compared with the other response conditions. Liu and Weng (2009) developed an effect size metric for the differences between two alphas that was used to quantify the difference between the highest alpha (.92 for the 5-point condition) and lowest alpha (.79 for the 2-point condition). The estimate was .65 in a Cohen’s d-metric. The difference between the 3-point condition (alpha = .87) and the 5-point condition (alpha = .92) was moderate (d = 0.33). In sum, the 2-point condition had a comparatively worse alpha compared with the other conditions, but all estimates were near values that are typically deemed acceptable for research purposes (see Lance et al., 2006 for a thoughtful consideration of cut-offs associated with reliability coefficients).
Coefficient Alpha Estimates for the Rosenberg Self-Esteem Scale by Response Option Condition for Each Study.
Note. Sample size is the minimum number of item responses per condition. CI = Confidence Interval; AIC = Average Inter-item Correlation.
Does Increasing the Number of Response Options Decrease Estimates of Central Tendency and Reduce Skewness?
Estimates of central tendency in the Percentage of Maximum Possible metric (POMP; Cohen et al., 1999) are reported in Table 2. The POMP metric offers a way to put responses across all conditions on a common metric to facilitate comparisons. An individual’s POMP score is calculated using the following equation: [(observed scale score – minimum possible scale score) / (maximum possible scale score – minimum possible scale score)] * 100. An ANOVA comparing mean POMP metrics for each condition yielded a statistically significant result (F = 18.83, df = 4, 734, p < .001). Visual inspection of means suggested relatively high scores for the 2- and 3-point conditions, similar means for the 4- and 5-point conditions and then an increase with the 7-point condition. A series of post hoc comparisons using a Bonferroni correction suggested no differences between the 2- and 3-point conditions (p = 1.0) but differences between the 2-point condition and each of the other conditions (largest p = .0039). There was no evidence that the 3-point condition was different from the 7-point condition (p = .21), whereas it was different from the 4- and 5-point conditions (ps < .001). The 4- and 5-point conditions were not significantly different (p = 1.0), but the 4-point condition was different from the 7-point condition (p = .0203). The 5-point condition was different from the 7-point condition (p = .0104).
Descriptive Statistics for the Rosenberg Self-Esteem Scale by Response Option Condition for Each Study.
Note. All scores are reported in the Percentage of Maximum Possible score metric.
The inspection of the percentage of the sample reporting the highest possible score (i.e., a POMP score of 100) suggested that ceiling effects were more likely with the 2- and 3-point conditions compared with greater response options (24% of the 2-point sample had a maximum possible score compared with fewer than 4% of the 4-, 5-, and 7-point conditions, respectively). Thus, the number options appears to influence the shape of the distribution of scores.
Does Increasing the Number of Response Options Impact Criterion-Related Correlations?
Correlations between self-esteem and the big five domains and life satisfaction for each response format are reported in Table 3. These correlations were generally consistent with previous studies pointing to the strongest effect sizes for neuroticism and extraversion. The same pattern was evident for the associations between self-esteem and life-satisfaction. We tested the similarity of correlations using a meta-analytic approach by computing the Q statistic for testing homogeneity of effect sizes (Cochran, 1954) for each set of correlations with a different criterion variable. The Q statistic was statistically significant for extraversion (p = .0045), neuroticism (p = .0051), and life satisfaction (p = .0190). A visual inspection suggested that the size of criterion-related correlations for the Rosenberg administered with the 4-point scale was more distinct from the others. We repeated the calculation of the Q statistic dropping this condition and all significant results were eliminated (smallest p = .0772 for extraversion). Given the unexpected result for the 4-point version, we reserved judgment to see if this pattern was replicable.
Correlations Between Self-Esteem and the Big Five Domains and Life-Satisfaction by Response Option Condition for Each Study.
Note. E = Extraversion; A = Agreeableness; C = Conscientiousness; N = Negative Emotionality; O = Open-mindedness; LS = Life Satisfaction.
To further quantify differences between correlations, Cohen’s q (Cohen, 1988) was computed between the largest and smallest correlation with each criterion variable. This effect size measure is the absolute value of the r-to-z transformed correlations. These effect sizes varied considerably (extraversion: 0.45; agreeableness: 0.15; conscientiousness: 0.24; negative emotionality: 0.37; open-mindedness: 0.25; life satisfaction: 0.36). However, no clear pattern was evident in the table such that the highest correlations were consistently observed when self-esteem was measured with more response options. In other words, we did not find evidence that administering the RSES with more scale points consistently improved criterion-related validity coefficients.
We conducted the same tests of the homogeneity of correlations for the HEXACO domains (and altruism; Table 4). The Q statistic was only statistically significant for emotionality (p = .0282). Cohen’s q between the largest and smallest correlation for each domain (and altruism) were consistent (honesty/humility: 0.29; emotionality: 0.30; extraversion: 0.26; agreeableness: 0.30; conscientiousness: 0.30; openness: 0.23; altruism: 0.26). But, again, there was no clear pattern to the differences indicating that formats with more response options showed increased criterion-related validities. Finally, we conducted the same tests for the TriPM domains (Table 5). The only significant Q statistic was for boldness (p = .0092). Cohen’s q between the largest and smallest correlation for each domain were 0.18 for disinhibition, 0.21 for meanness, and 0.40 for boldness. Thus, the same general pattern held for the HEXACO domains and the triarchic domains—results were consistent with previous studies and there was no evidence that correlations are stronger when the Rosenberg was administered with more response options.
Correlations Between Self-Esteem and the HEXACO Domains by Response Option Condition for Studies 1 and 3.
Note. H = Honesty/Humility; E = Emotionality; X = Extraversion; A = Agreeableness; C = Conscientiousness; O = Openness; Alt = Altruism.
Correlations Between Self-Esteem and the Triarchic Domains by Response Option Condition for Studies 1 and 3.
Exploratory Analysis: Did Reports of Survey Quality Differ by Condition?
There was an indication that means for the quality of the survey item differed across response conditions using ANOVA (F = 2.82, df = 4, 732, p = .0244). The smallest mean value was 4.24 (SD = 0.92) for the 4-point condition whereas the largest mean value was 4.54 (SD = 0.92) for the 7-point condition. This translated to a d-metric effect size of 0.33. The analysis was not preregistered and defied a simple interpretation given the overall patterns of means (2-point: 4.30; 3-point = 4.50; 5-point = 4.38; Pooled SD = 0.92).
Study 1 Discussion
Results from Study 1 suggest that observed alpha coefficients were lowest for the 2-point condition and relatively similar for the 3- to 7-point conditions. This is partially consistent with previous research in which internal consistency was found to increase as response options increased from two with a plateau around five or six (Simms et al., 2019). Second, and consistent with Leung (2011), we found that increasing the number of response options produced closer approximations of normal score distributions especially when considering scores in the 4-, 5-, and 7-point conditions. Third, although we observed some differences in associations between criterion-related variables among response option conditions, no consistent pattern favoring response formats with more options emerged. Exploratory analyses point to potential differences in perceived survey quality across conditions but again the pattern defied an easy explanation. To test the robustness of all these results, we conducted a second, preregistered study using a larger, more diverse sample and which added a 6-point condition. Indeed, Simms et al. (2019) noted the importance of testing response option differences in samples comprised of participants other than college students. They also advocated for 6-point formats based on a response option study with big five trait domains.
Study 2 Method
Participants
Participants were members of online panels maintained by Qualtrics. A consumer/general population sample of adults ages 18 or older with a mix of age and income levels and an equal balance of women and men was requested. The target sample size was 2,400 based on the heuristic of recruiting approximately 400 participants per condition. The design decision was preregistered (https://osf.io/ae5gt/?view_only=e9c92f075b6b44478d30110c2f13eb08). The sample included in this report completed self-esteem items and passed quality control checks: a correct response to a directed response item embedded within the BFI-2 (select “disagree a little”), an affirmative response to item asking if they responded honestly, and a positive response to serious responding item (Aust et al., 2013). The first quality check was used by the market research firm to identify 2,403 quality responders and the second and third quality control checks were preregistered and imposed by the researchers (similar to Study 1). Of the 2,403 responders, 2,389 participants responded affirmatively to the honesty question and 2,371 indicated that they took part seriously. A total of 2,358 participants passed the honesty and seriousness check and were included in this report. The sample was 48.5% women and 50.3% men with the remaining 1.2% identifying as non-binary, self-describing, or preferring not to disclose. Age was a categorical variable with participants from 18 to over 65 (18–24 = 11.3%; 25–34 = 18.1%; 35–44 = 17.6%; 45–54 = 19.4%; 55–64 = 16.9%; 65 or older = 16.6%; Prefer not to say = 0.1%). The self-identified racial-ethnic breakdown was 76.5% White, 12.6% Black, 3.5% Asian, 3.4% multi-racial, with the remaining 4% identifying as another race or preferring not to disclose.
Measures
Big Five Domains
The 60-item BFI-2 also used in Study 1 measured extraversion (alpha = .83, M = 3.19, SD = 0.73), agreeableness (alpha = .81, M = 3.86, SD = 0.62), conscientiousness (alpha = .87, M = 3.77, SD = 0.73), negative emotionality (alpha = .90, M = 2.76, SD = 0.88), and open-mindedness (alpha = .82, M = 3.68, SD = 0.67).
Life Satisfaction
The five-item SWLS also used in Study 1 measured life satisfaction (alpha = .91, M = 4.16, SD = 1.58).
Self-Esteem
The single-item self-esteem scale (Robins, Hendin, et al., 2001) provided an alternative measure of global self-esteem (M = 3.39, SD = 1.27) using a 5-point response scale.
Procedure
Participants completed surveys created using Qualtrics. Participants were first presented with demographic questions for gender, age, and income to identify sampling quotas. Participants then completed the BFI-2, the single-item self-esteem measure, and SWLS in a fixed order and were then assigned to complete the RSES items in one of six conditions: 2-, 3-, 4-, 5-, 6-, or 7-point response options. After completing the self-esteem items, a survey quality item was presented (“How would you rate the quality of this survey”) with 6 response options (Very poor; Poor; Fair; Good; Very good; Excellent). A few additional demographic questions were presented, and the last two items were a yes/no “I have answered all questions honestly” and the Aust et al. (2013) seriousness check.
Study 2 Results
Does Increasing the Number of Response Options Increase Estimates of Internal Consistency?
As was the case for Study 1, alpha coefficients were calculated within each condition and reported in Table 1. The null hypothesis of equality of coefficients was rejected (chi-square = 21.00, df = 5, p = .0008). The follow-up test dropping the 2-point format yielded a null result (p = .17). Thus, it appeared that internal consistency estimates were only appreciably lower for the 2-point format compared with the others. The effect size for the difference between the highest alpha (.91 for the 5- or 6-point formats) and lowest alpha (.86 for the 2-point format) was 0.30 in a Cohen’s d-metric. The difference between the 3-point format (alpha = .89) and the 5- or 6-point formats (alphas = .91) was trivial (d = 0.13). As with Study 1, estimates of alpha were within an acceptable range across all formats but were higher for formats with more than 2-points.
Does Increasing the Number of Response Options Decrease Estimates of Central Tendency and Reduce Skewness?
Estimates of central tendency in POMP metric are reported in Table 2. Inferential tests using ANOVA yielded a significant result (F = 11.21, df = 5, 2,352, p < .001). Visual inspection of means suggested a drop between the 3-point format and the 4-point format. A series of post hoc comparisons using a Bonferroni correction suggested differences between the 2-point format and the 4- to 7-point formats and between the 3-point format and the 4- to 7-point formats. No other pair-wise comparison was statistically significant. Moreover, an inspection of the percentage of the sample reporting the highest possible score suggested that ceiling effects were more likely with the 2- and 3-point formats compared with greater response options. For example, roughly 34% of the participants in the 2-point condition had the maximum score whereas roughly 3% of the participants in the 7-point condition had the maximum score.
Does Increasing the Number of Response Options Impact Criterion-Related Correlations?
Correlations between self-esteem and the big five domains and life satisfaction for each response format are reported in Table 2. The correlations for the single-item self-esteem scale were .56, .63, .65, .60, .61, and .56 for the 2-, 3-, 4-, 5-, 6-, and 7-point conditions (overall r = .59), respectively. In general, the correlations were similar in magnitude across response options. This impression was formally tested using the Q statistic from meta-analysis for each set of correlations with a different criterion variable. The Q statistic was only statistically significant for conscientiousness (Q = 22.14, df = 5, p < .001). The correlation for the 2-point format appeared lower than the others as the Q statistic was not significant for conscientiousness when only considering the 3- through 7-point formats (Q = 2.00, df = 4, p = .74).
To further quantify differences, Cohen’s q was computed between the largest and smallest correlation with each criterion variable. These were all small according to conventions except for conscientiousness (extraversion: 0.13; agreeableness: 0.13; conscientiousness: 0.30; negative emotionality: 0.15; open-mindedness: 0.13; life satisfaction: 0.10; single-item self-esteem: 0.14). Moreover, no clear pattern was evident in the table such that the highest correlations were consistently observed when self-esteem was measured with more response options.
Exploratory Analysis: Did Gender Differences Vary by Condition?
An exploratory analysis to evaluate whether gender differences in self-esteem varied across the response formats was included in the preregistration. There was no evidence of differences. None of the tests for gender difference were significant in any condition. Moreover, a test of the homogeneity of effect size estimates was not statistically significant. Complete details are available upon request and can be tested using the data and code that accompany this article. The meta-analytic effect size estimate in the d-metric was .002 suggesting virtually no differences in average levels of self-esteem between those who identified as women and those who identified as men.
Exploratory Analysis: Did Reports of Survey Quality Differ by Condition?
Means for the quality of the survey item were similar across conditions (F = 1.72, df = 5, 2,330, p = .127). The smallest mean value was 4.75 (SD = 1.03) for the 6-point condition whereas the largest mean value was 4.94 (SD = 0.95) for the 3-point condition. This translated to a d-metric effect size of 0.19. The analysis was not preregistered but provided no indication that response options impacted judgments of survey quality. The overall ANOVA effect observed in Study 1 did not replicate in Study 2.
Study 2 Discussion
Several findings from Study 1 replicated in Study 2, which was larger and based on participants with wider age ranges. For example, the alpha estimate for the 2-point condition was the lowest of all conditions whereas there was no evidence that the other conditions differed from each other. Also mirroring the Study 1 results, increasing the number of response options reduced the number of participants with maximum scores. Unlike Study 1, however, a clearer pattern of null results with respect to associations between self-esteem and criterion-related variables emerged: there was little evidence for meaningful differences across the conditions. In addition, we conducted two sets of exploratory analyses: the first suggested no gender differences in RSES scores across response option conditions and the second revealed no differences in perceived survey quality across conditions. Thus, these results suggest few psychometric differences for the 4-, 5-, 6-, and 7-point formats. However, given that the 11-point response format has been recommended by some researchers (e.g., Leung, 2011), we conducted a third study to compare 2-, 5-, and 11-point formats using college student participants. The procedures for Study 1 and Study 3 were identical except for the use of a different set of response option conditions.
Study 3 Method
Participants
Participants were college students who indicated honest responding and that they took part seriously (Aust et al., 2013), as well as completed the self-esteem items. Data from up to 1,461 participants were analyzed in this report out of the 1,657 responses with unique ID numbers from the Department of Psychology research participant pool. Of the 1,657 responses, 1,476 participants responded affirmatively to the honesty question and 1,561 indicated that they took part seriously. A total of 1,463 participants passed both quality checks but 2 of those individuals were missing all self-esteem responses and were excluded. The sample was 71.9% women and 27.7% men with the remaining 0.5% identifying as non-binary, preferring not to disclose, or skipping the question. Participants age ranged from 18 to 52 years (M = 19.56, SD = 2.08). The self-identified racial-ethnic breakdown was 70.9% White, 14.1% Asian, 6.4% Black, 3.7% multi-racial, with the remaining 4.9% identifying as a different race, indicating a preference not to disclose, or skipping the question.
Measures
Big Five Domains
The BFI-2 measured extraversion (alpha = .86, M = 3.33, SD = 0.72), agreeableness (alpha = .77, M = 3.79, SD = 0.56), conscientiousness (alpha = .85, M = 3.59, SD = 0.66), negative emotionality (alpha = .89, M = 3.00, SD = 0.80), and open-mindedness (alpha = .84, M = 3.63, SD = 0.66).
Life Satisfaction
The SWLS measured life satisfaction (alpha = .88, M = 4.72, SD = 1.30).
HEXACO Domains
The HEXACO measured honesty-humility (alpha = .78, M = 3.22, SD = 0.51), emotionality (alpha = .83, M = 3.48, SD = 0.57), extraversion (alpha = .86, M = 3.28, SD = 0.60), agreeableness (alpha = .82, M = 3.01, SD = 0.54), conscientiousness (alpha = .84, M = 3.53, SD = 0.55), openness (alpha = .78, M = 3.15, SD = 0.54), and altruism (alpha = .56, M = 3.88, SD = 0.63).
Psychopathic Traits
The TriPM measured disinhibition (alpha = .86, M = 1.84, SD = 0.43), meanness (alpha = .90, M = 1.70, SD = 0.47), and boldness (alpha = .81, M = 2.57, SD = 0.42).
Study 3 Results
Does Increasing the Number of Response Options Increase Estimates of Internal Consistency?
Estimates of internal consistency are reported in Table 1. The null hypothesis for the equality of alpha coefficients was rejected (Chi-square = 35.11, df = 2, p < .001) and the follow-up test dropping the 2-point condition yielded a null result (p = .29). Thus, it appeared that internal consistency estimates were only appreciably lower for the 2-point condition compared with the other two conditions. The effect size difference between the highest alpha (.91 for the 11-point format) and the lowest alpha (.84 for the 2-point format) was 0.39 in a Cohen’s d-metric.
Does Increasing the Number of Response Options Decrease Estimates of Central Tendency and Reduce Skewness?
Estimates of central tendency in the POMP metric are reported in Table 2. Mean differences were observed from an ANOVA (F = 42.39, df = 2, 1,458, p < .001). A series of post hoc comparisons using a Bonferroni correction suggested differences among all conditions. Moreover, an inspection of the percentage of the sample reporting the highest possible score suggested that ceiling effects were more likely with the 2-point format. Roughly 30% of the 2-point condition had the maximum score whereas roughly 2% of the 5- and 11-point conditions had the maximum score.
Does Increasing the Number of Response Options Impact Criterion-Related Correlations?
Correlations between self-esteem and the big five domains and life satisfaction for each response format are reported in Table 3. The Q statistic was statistically significant for life satisfaction (p = .0151), but not for any of the big five domains. Cohen’s q between the largest and smallest correlation with each criterion variable were consistently small (extraversion: 0.10; agreeableness: 0.12; conscientiousness: 0.10; negative emotionality: 0.04; open-mindedness: 0.05; life satisfaction: 0.18). The Q statistic was only statistically significant for honesty/humility (p = .0420) when considering the HEXACO domains (and altruism scale; Table 4). Cohen’s q estimates were consistently small (honesty/humility: 0.16; emotionality: 0.02; extraversion: 0.12; agreeableness: 0.11; conscientiousness: 0.15; openness: 0.09; altruism: 0.07). Finally, Q statistics were not statistically significant for any TriPM domain (all ps > .30; q = 0.03 for disinhibition, 0.09 for meanness, and 0.06 for boldness; Table 5).
Exploratory Analysis: Did Reports of Survey Quality Differ by Condition?
There was no indication that means for the quality of the survey item differed across response conditions using ANOVA (F = 0.23, df = 1, 1,459, p = .634). The means were similar (2-point: 4.47; 5-point = 4.36; 11-point: 4.43, Pooled SD = 0.91). The difference between the smallest and largest means translated to a d-metric effect size of 0.12.
Study 3 Discussion
Consistent with the previous two studies, results from Study 3 revealed that a 2-point format produced lower alpha estimates compared with formats with more response options. There were few indications that the11-point format was appreciably superior to the 5-point format in terms of alpha estimates, estimates of central tendency, and number of participants at ceiling. Criterion-related correlations were similar across response formats.
Supplemental Analyses
We focused on coefficient alpha as our primary estimate of internal consistency in Studies 1 to 3. Some psychometric experts have raised concerns about computing alpha with ordinal responses (e.g., Gadermann et al., 2012; Zumbo et al., 2007; but see Chalmers, 2018). Zumbo et al. (2007) reported that coefficient alpha was a negatively biased estimate of theoretical reliability, particularly when there were few response options. A simple explanation is that attenuation occurs when coarsely categorized variables (item responses in this case) are used to estimate associations between continuous latent variables (see, for example, Aguinis et al., 2009) and such attenuation depresses estimates of internal consistency when using conventional alpha coefficients. To address this issue, we computed ordinal alpha coefficients using polychoric correlation matrices following procedures in Gadermann et al. (2012). Those results are reported in the supplement (Table S2). Consistent with Zumbo et al. 2007 and Müssig et al. (2021), estimates of ordinal alpha tended to be larger than the “standard” alpha estimates reported in Table 1. This difference was more pronounced with fewer response options. Moreover, the estimates of ordinal alpha were quite similar across conditions suggesting that any observed differences in alpha coefficients reported in Table 1 are potentially due to a negative bias in estimating internal consistency when scores are based on items with very few response options (i.e., two response options). A similar conclusion was drawn by Müssig et al. (2021).
In addition to concerns about the use of ordinal item responses when computing alpha, there is increasing recognition of the utility of alternatives to alpha such as omega coefficients (see, for example, Flora, 2020; McNeish, 2018; Revelle & Condon, 2019; see also Savalei & Reise, 2019 for commentary on this literature). These alternatives to alpha take the dimensionality and structure of a measure into account rather than assuming a unidimensional structure as is the case with alpha (see, for example, Schmitt, 1996). A complicating issue for determining the most appropriate model-based measure of internal consistency for the current research is that the structure of the RSES is a matter of ongoing debate (e.g., Alessandri et al., 2015; Donnellan et al., 2016; Gnambs et al., 2018; Huang & Dong, 2012). One perspective is that the RSES is essentially unidimensional or that a bifactor model with a general factor and orthogonal positively and negatively keyed method factors serves as a viable structure (Alessandri et al., 2015; Donnellan et al., 2016; Gnambs et al., 2018; Huang & Dong, 2012). Resolving debates about the structure of the RSES was beyond the scope of this report. However, to address potential concerns with our focus on alpha, we computed estimates of omega using a range of approaches including estimating coefficients from unidimensional and bifactor confirmatory-factor analytic models (see Flora, 2020 for a review). Bifactor models occasionally had convergence issues within some conditions for particular studies (likely due to sample size issues combined with estimation techniques for categorical indicators). Accordingly, we combined samples for each response-option condition across studies to maximize sample size for bifactor models. Results are reported in supplemental analyses (Tables S3 and S4) but generally pointed similar omega estimates across conditions when using estimation techniques for categorical indicators for confirmatory factor analytic models (i.e., Weighted Least Square Mean and Variance adjusted).
General Discussion
The current studies clarify how the number of response options impact the psychometric properties of the RSES. Consistent with previous work using personality trait inventories (Finn et al., 2015; Müssig et al., 2021; Simms et al., 2019), estimates of internal consistency using coefficient alpha were lowest for scores based on scales with 2-point response options compared with scores generated from items with greater response options. Alpha coefficients did not seem to differ substantially when comparing versions administered with three or more response options. Supplemental analyses using ordinal alpha (Gadermann et al., 2012) did not suggest internal consistency differences across conditions. Likewise, increasing response options seemed to have little systematic impact on criterion-related correlations. The distributions of scores, however, were impacted by the number of response options. Ceiling effects were more pronounced when using fewer than four response options (e.g., 34% max in the 2-point condition compared with 5% in the 4-point condition in Study 2). The exploratory tests of perceived survey quality yielded null and inconsistent differences and thus do not offer support for any specific conclusion about response options and participant ratings of survey quality.
Given the overall pattern of results from these three experiments, researchers may wish to administer the Rosenberg with at least four response options, especially if there are concerns about the distribution of scores and a desire to avoid ceiling effects. More specifically, we think it is worth recommending five scale points for the RSES for several reasons. First, five-point Likert-type scales seem to generate data that can be subjected to traditional psychometric analyses like item-based confirmatory factor analysis using standard maximum-likelihood estimation techniques (see, for example, Rhemtulla et al., 2012; see also Bollen & Barb, 1981). Second, five-point formats are consistent with the response scales for many big five trait inventories (e.g., Soto & John, 2017) and the HEXACO inventory (e.g., Lee & Ashton, 2018). Five response options are also common in the self-esteem literature (see, for example, Blascovich & Tomaka, 1991; Donnellan et al., 2015; Gray-Little et al., 1997). Thus, selecting five options may facilitate consistency in response formats within the same survey and with other studies in the literature. Third, we suspect that an 11-point response format for the RSES is cumbersome and there are hints that users find such a scheme more difficult to use for personality items when compared with formats with fewer options (Müssig et al., 2021). Finally, we favor a middle-point response option for self-esteem items given that we believe it is important to allow participants to explicitly express neutral or mixed feelings about the self. Providing such an option may better allow researchers to evaluate whether there is a bias toward neutral responding in some groups including those from collectivist cultures (see, for example, Schmitt & Allik, 2005).
Implications, Limitations, Future Directions
In general, our results seem to echo the general pattern of recent results reported by Simms et al. (2019) when they varied the response options for a big five measure and by Müssig et al. (2021) when they varied the response options for a big five and an emotion regulation measure. The literature is coalescing around the conclusion that somewhere between five and seven Likert-type response options is optimal for self-report measures of constructs like personality traits, global self-esteem, and emotion regulation. Researcher preference may play a role within this range. We suggest that researchers then consider the response formats of the other items in the survey and whether a neutral point is desirable. Simms et al. (2019) expressed a preference for 6-point formats (without a neutral point) whereas we noted a preference for 5-point formats with a neutral point. Simms et al. (2019) also concluded that “the most honest appraisal of our results is that it probably doesn’t matter much” (p. 564) when deciding between formats with an odd or even number of responses. This perspective is consistent with our view that preferences for particular response formats with Likert-type scales may involve a degree of subjectivity within a particular range of 5 to 7 points.
Although this study was based on relatively large sample sizes and Study 2 was preregistered, the current work had limitations. All variables were self-reported and all measures were administered over the internet. The underlying samples were likely to be familiar with survey research. It remains to be seen if these conclusions would hold when using more diverse samples in terms of culture and levels of education. We also caution against overgeneralizing from these results to constructs that are not trait-like and with response formats that are not Likert-type (i.e., variations of Disagree to Agree responses). It may be the case that response option effects are more pronounced in those cases.
Conclusion
We favor a 5-point scale with a neutral option for the RSES. We also acknowledge that the existing evidence does not provide strong or overwhelming evidence that response options other five are unequivocally problematic for the RSES. In fact, some might be surprised at how well 2-point versions of the RSES performed here in terms of internal consistency and criterion-related correlations. When viewing the totality of the existing evidence, we suspect that the number of response options matters less than having a decent-sized set of well-written items to measure core constructs. Indeed, the current findings suggest that for most applications of the RSES, using four or more response options will yield similar psychometric properties as indicated by internal consistency estimates, normality of score distributions, and associations with criterion-related variables.
Supplemental Material
sj-docx-1-asm-10.1177_10731911221119532 – Supplemental material for How Does the Number of Response Options Impact the Psychometric Properties of the Rosenberg Self-Esteem Scale?
Supplemental material, sj-docx-1-asm-10.1177_10731911221119532 for How Does the Number of Response Options Impact the Psychometric Properties of the Rosenberg Self-Esteem Scale? by M. Brent Donnellan and Andrew Rakhshani in Assessment
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
