Abstract
The item wording (or keying) effect is respondents’ differential response style to positively and negatively worded items. Despite decades of research, the nature of the effect is still unclear. This article proposes a potential reason; namely, that the item wording effect is scale-specific, and thus findings are applicable only to a particular measure involved in an investigation. Using multiple scales and several methods, the present study provides strong and converging support for this hypothesis. In light of the results, we reinterpret major findings of the item wording effect and propose possible future directions for research.
Keywords
Researchers often include both positively worded and negatively worded items to measure a given construct. Positively worded items usually measure the presence of a positive construct (such as social equality), while negatively worded items usually measure the absence of such a construct. However, respondents may not give logically consistent answers to positively and negatively worded items, despite the items’ equivalent (but opposite) content. Such inconsistency is often referred to as the item wording (or keying) effect (C. C. S. Kam and Meyer 2015).
Despite decades of research, no conclusion has been drawn on the exact nature of the item wording effect: Studies often show different empirical results with distinct conclusions. For example, some researchers found the item wording effect to correlate significantly with social desirability response style (Rauch, Schweizer, and Moosbrugger 2007), while others failed to replicate this finding (DiStefano and Motl 2006, 2009). The current article addresses a perplexing question in the literature: Why has previous research been unable to give a conclusive answer on the nature of the item wording effect? Specifically, why do findings from one study not generalize to another?
To answer these questions, we first explain the item wording effect and common methods used to investigate it. We then identify the inconsistent findings in the literature and provide a possible explanation. Finally, we describe an empirical study that tests the explanation.
Item Wording Effect
Many survey instruments include both positively and negatively worded items to measure a construct. For instance, when measuring participants’ extroversion, researchers use positively worded (e.g., “I am an extrovert”) and negatively worded (e.g., “I am an introvert”) items to capture the positive and negative pole of the same construct (Peabody 1967). However, when factor analyses are conducted, the positively and negatively worded items usually load on separate factors. This result led some researchers to conclude that positively and negatively worded items measure distinct (albeit correlated) constructs (e.g., Chang, Maydeu-Olivares, and D’Zurilla 1997; Credé et al. 2009; Robinson-Whelen et al. 1997). Other researchers attributed the two-factor phenomenon to a response style: Participants respond to positively and negatively worded items differently in general (e.g., Horan, DiStefano, and Motl 2003; Magazine, Williams, and Williams 1996; Marsh 1996; Motl and DiStefano 2002; Tomás and Oliver 1999; Tomás et al. 2013). The latter group of researchers regards the item wording effect as a methodological artifact.
Either way, it is known that the item wording effect is quite prevalent in survey scales (C. Kam and Meyer 2012). Earlier researchers found the effect for self-esteem (Carmines and Zeller 1979; Marsh 1996), optimism (Marshall et al. 1992; Plomin et al. 1992), anxiety (Vagg, Spielberger, and O’Hearn 1980; Vautier et al. 2004), and organizational commitment (Magazine et al. 1996). Later findings suggested that the effect exists for other common psychological measures such as the Big Five personality dimensions (Biderman et al. 2011; see also Biderman et al. 2013). As a result, Biderman et al. (2011) concluded that the item wording effect is ubiquitous across survey instruments and thus cannot be ignored.
To examine the nature of the item wording effect in a survey instrument, early researchers simply correlated summed scores of the survey’s positively worded items with external variables. They then did the same with negatively worded items. Differential patterns of correlations may shed light on the nature of item wording. For example, Carmines and Zeller (1979) investigated whether scores from positively and negatively worded self-esteem items correlated differently with external criteria. No significant differences were found, and Carmines and Zeller argued there was no substantive difference between positively and negatively worded items. Marshall et al. (1992) correlated optimism and pessimism items with external variables and found that optimism correlated stronger with extroversion and that pessimism correlated stronger with neuroticism. Marshall et al. concluded that optimism and pessimism items are distinct constructs.
More recently, using the more sophisticated technique of multitrait-multimethod confirmatory factor analysis (e.g., correlated trait-method minus one, CT[M − 1] technique; Eid 2000), researchers have decomposed item variance into (a) variance due to trait and (b) variance associated with the use of positively or negatively worded items (e.g., Alessandri et al. 2011; Vautier et al. 2004). This technique—which can be used with single- and multiple-trait scales—is a substantial improvement over the previous procedure of Carmines and Zeller (1979), because it allows researchers to correlate the unique variances in positively or negatively worded items with a wide range of external variables, such as social desirability, approach and avoidance motivation system, and self-enhancement. A significant correlation between the external variables and the wording factor could provide information regarding the nomological network of the item wording effect (i.e., its relationship with external associates; Cronbach and Meehl 1955). For example, Vecchione et al. (2014) extracted an item wording factor from positively worded items on an optimism scale and found that the wording factor correlated with respondents’ tendency to self-enhance. Vecchione et al.’s finding thus suggested that the item wording effect was related to participants’ motivation to present themselves in an overly positive manner. Quilty, Oakman, and Risko (2006) found that the item wording factor associated with negative self-esteem items was related to participants’ avoidance motivation. Lindwall et al. (2012) found that the same wording factor was positively correlated with depression and negatively correlated with life satisfaction. Depressed and dissatisfied individuals may thus be more likely to endorse negatively than positively worded self-esteem items.
Inconsistent Findings on the Nature of Item Wording Effects
Although there is little debate that the item wording effect is pervasive in survey instruments (Biderman et al. 2011), there is no consensus on its nature. The nomological network findings of the item wording effect have never been consistently replicated. Take, for example, the case of social desirability response style as a plausible explanation of the effect. Rauch et al. (2007) argued that participants may have different social desirability response styles to positively and negatively worded items. The response styles may thus cause a survey respondent to answer positively and negatively worded items differently, artificially producing a factor due to item wording (cf. DiStefano and Motl 2006). In their empirical study, Rauch et al. (2007) found that the artificial factor from positively worded optimism items correlated significantly with a facet of social desirability response style (r = .35), meaning that the factor could have arisen from participants’ tendency to report themselves in an overly positive manner. Summarizing their findings, Rauch et al. (2007) stated, “Deviation from unidimensionality of observed scores does not imply deviation from unidimensionality of optimism when method effects are incorporated in the model” (p. 1597). These researchers therefore attributed the emergence of the “method” or wording effect as simply a methodological artifact. In a more recent study, Alessandri et al. (2011) investigated how the wording factor unique to positively worded optimism items (i.e., not pessimism items that are negatively worded) is related to external variables. Consistent with the argument by Rauch et al., Alessandri et al. found optimism items to be strongly predicted by egoistic bias (a close correlate of social desirability response style). They argued that the wording effect, in short, is merely a response style: It does not represent any substantive construct other than social desirability response bias.
The claim that the item wording effect is nothing more than a methodological artifact, however, is not universally accepted. DiStefano and Motl (2006) tried to predict the wording effect on Rosenberg’s self-esteem scale from social desirability response style, but the response style was not a significant predictor. DiStefano and Motl (2009) later split their participants by gender and found similar results—the wording factor related to the use of negatively keyed self-esteem items was unrelated to social desirability response style. Given these inconsistent findings regarding the nomological network of the wording factor, it is almost impossible to ascertain its status. For at least three decades now, researchers have been examining the nature of the item wording effect; we have yet to understand it.
The current work suggests that the inconclusive status of the item wording effect is due to its scale-specific nature. Researchers have often implicitly or explicitly assumed that correlational findings related to the wording effect in a particular survey instrument (such as self-esteem or optimism) are reasonably generalizable to findings from other survey instruments (Quilty et al. 2006). For example, DiStefano and Motl (2009) found that constructs signaling fear of negative evaluations (e.g., self-consciousness) were correlated with the wording effect in a particular self-esteem measure. Based on this result, they concluded that “the concern over possible negative evaluation by others seems to play a role in explaining method effects among negatively worded items” (p. 460). They did not limit the scope of their conclusion to the self-esteem measure in their investigation.
However, the assumption regarding the generalizability of correlational results is seldom upheld. DiStefano and Motl (2006) modeled two negatively worded factors: one from a self-esteem instrument and the other from an anxiety instrument. They found the two factors to be only weakly correlated (r = .37), suggesting that wording factors from different instruments are possibly not the same or similar. Pohl and Steyer (2010) also found the correlation of the wording factors to be weak (r = .35) between a calmness measure and an alertness measure. Although further investigations are required to check how consistent these findings are across scales, these early results inspired us to formulate two hypotheses: (1) Item wording factors from different scales are not identical because they may be only weakly correlated with each other and (2) if the former hypothesis is true, how a wording factor from a scale related to external variables will differ from one survey instrument to another. Previous empirical studies do not provide conclusive evidence regarding the nature of the item wording effect in general because each used a single specific scale.
The Present Study
DiStefano and Motl’s (2006) preliminary findings suggested that the factors associated with item keying direction, and their correlations with external variables, are specific to a particular survey scale: The results would not be generalizable to another scale. While DiStefano and Motl found a low correlation between wording factors from two specific scales, the current research aims to conduct a more comprehensive investigation of this issue. To test the first hypothesis—that item wording factors are scale-specific—we examine the intercorrelations among the wording factors extracted from different survey instruments. Low correlations would support the hypothesis. To examine the second hypothesis—that the nomological network correlations related to wording effects from different instruments are also scale-specific—we correlate social desirability response style with the various wording factors. If these correlations vary substantially from one scale to another, the second hypothesis would be supported. Social desirability response style was chosen because it has often been used as an external construct to validate the nature of item wording factors (e.g., DiStefano and Motl 2006; Rauch et al. 2007).
However, even if the two hypotheses are supported, our argument apparently contradicts findings in another line of research, showing that the item wording effect can be stable over time. For example, Marsh, Scalas, and Nagengast (2010) extracted an item wording factor for positively worded self-esteem items and another one for negatively worded self-esteem items; these item wording factors were significantly correlated across multiple one-year periods (rs = .39–.65 across time; see also Motl and DiStefano 2002). Using an optimism measure, Vecchione et al. (2014) similarly found that the item wording factor across three two-year periods were significantly correlated (rs = .33–.38). These studies clearly demonstrated the stability of item wording effect. Do these findings contradict our earlier two hypotheses regarding the instability of the wording effect across distinct measures?
We believe the answer to be no, as we conjecture that the item wording effect has both a stable and an unstable characteristic. We speculate that the effects are stable within the same measure across time (as in Marsh et al. 2010; Motl and DiStefano 2002; Vecchione et al. 2014) and possibly across disparate samples, even though the effect varies across measures (as implied by the findings of DiStefano and Motl 2006). This led to the third hypothesis: Item wording factors should have a similar pattern of correlations with external constructs across two distinct samples (i.e., between-sample stability), even though these correlations fluctuate widely from one focal measure to another within any one sample (i.e., between-scale variability). To examine this hypothesis, we tested a second sample of participants with survey instruments partially overlapping those of the first sample. Hypotheses and expected results are summarized in Table 1.
Hypotheses in the Current Study.
Method
Sample 1
One thousand four hundred and six students (947 females, 452 males, and 7 unidentified) were recruited from a large Canadian university. They completed an online survey for partial course credit. The mean age was 18.47 years (SD = 2.19). The data collection had obtained approval by the ethics board of the university. All participants provided informed consent to participate.
Online Scales
Participants in sample 1 completed the following scales:
International Personality Item Pool (IPIP)
The instrument consisted of five scales measuring Big Five personality factors (Goldberg et al. 2006), namely, openness to experience (Cronbach’s α = .73, 95% CI [.70, .75]), conscientiousness (Cronbach’s α = .79, 95% CI [.77, .81]), extroversion (Cronbach’s α = .88, 95% CI [.86, .89]), agreeableness (Cronbach’s α = .76, 95% CI [.73, .78]), and neuroticism (Cronbach’s α = .85, 95% CI [.83, .87]). Participants completed the measure on a 5-point Likert-type scale (1 = strongly disagree, 5 = strongly agree). Each personality factor was measured by five regular-keyed items and five reverse-keyed items and was theoretically unidimensional.
Social Dominance Orientation (SDO)
The instrument included 16 items (Cronbach’s α = .90, 95% CI [.89, 92]), measuring participants’ preference for social group inequality (Pratto et al. 1994). Sample items are “superior groups should dominate inferior groups” (negatively worded) and “all groups should be given an equal chance in life” (positively worded). Half of the items were positively and half negatively worded. Participants completed the measure in a 7-point Likert-type scale (1 = strongly disagree, 7 = strongly agree).
Team Meeting Attitudes (TMAs)
The instrument included eight items (Cronbach’s α = .91, 95% CI [.90, .92]), measuring participants’ general favorability toward team meetings (O’Neill and Allen 2012). Sample items are “most decisions should be made during group meetings” (positively worded) and “meetings are overrated” (negatively worded). Half of the items were positively and half negatively worded. Participants completed the measure in a 5-point Likert-type scale (1 = strongly disagree, 5 = strongly agree).
Balanced Inventory of Desirable Responding (BIDR)
Paulhus (1991) developed a popular social desirability instrument with two facets: impression management (Cronbach’s α = .79, 95% CI [.77, .81]) and self-deception (Cronbach’s α = .66, 95% CI [.63, .69]). It has been suggested that impression management reflects the conscious aspect of social desirability (but see Uziel 2010), whereas self-deception reflects the less conscious aspect of the construct (Li and Bagger 2007). The original version of the BIDR has 20 items for each facet of social desirability, half of which are reverse keyed. Due to concerns from the ethics board, however, one item was removed from each facet. Participants completed the measure in a 7-point Likert-type scale (1 = strongly disagree, 7 = strongly agree). Although Paulhus originally suggested that item scores be dichotomized, later researchers discovered more favorable psychometric properties when the raw scores are treated as continuous indicators (C. Kam 2013; Stöber, Dette, and Musch 2002).
Sample 2
To determine whether correlations of social desirability with wording factors are consistent across samples, a second sample of introductory psychology students (N = 1,254) who completed the IPIP, the SDO, and the BIDR was recruited. Although samples 1 and 2 have similar demographics (i.e., university students), they were collected four years apart. Data in sample 2 have been published in a previous study (C. C. S. Kam and Meyer 2015); we borrowed the data to examine the replicability of the findings in sample 1.
Analysis Strategies
All statistical analyses were conducted in the R console Version 3.1.0 (R Development Core Team 2013), with lavaan library package for structural equation modeling (SEM) analysis (Rosseel, 2012). All models were conducted with robust maximum likelihood (MLR) estimator, because it does not assume multivariate normality in the data. Model comparison was based on the appropriate scaled χ2 statistics for MLR estimator (Satorra 2000). Missing data were treated with multiple imputation with the Amelia library package in R program. All SEM analyses were conducted with item-level indicators.
First, we examined the consistency of item wording factors across measures in sample 1. In the first model (Figure 1, model a), a wide range of measures were included and allowed to correlate with one another. All construct variances were set as unity. In the second model, for each measure, we modeled negative wording factors and studied their intercorrelations (Figure 1, model b). Construct and wording factors were restricted to be uncorrelated for reasons of identification. The variances of all construct and wording factors were set as unity. The decision to model a wording factor on negatively worded items only is consistent with the framework of the CT(M − 1) model (Eid 2000). We did not model both the positively and negatively worded factor for two reasons. First, previous researchers have suggested that the simultaneous modeling of both factors (i.e., correlated trait-correlated method [CTCM] model) leads to overextraction of the trait variance within the wording factors (Marsh 1989). Second, previous investigators have suggested that variances related to item wording come mainly from negatively rather than positively worded items (e.g., Lindwall et al. 2012; Quilty et al. 2006), possibly because negatively worded items can cause interpretational difficulty (Swain, Weathers, and Niedrich 2008; van Sonderen, Sanderman, and Coyne 2013). If the second model fits better than the first, the effect related to item wording generally exists in the data set. If the intercorrelations among wording factors are weak, our Hypothesis 1 that wording factors are scale-specific would be supported. To further test the scale-specificity hypothesis, we tested a one-wording factor model in which all wording factors were lumped together to form one-wording factor (Figure 1, model c). If Hypothesis 1 is correct, the one-wording factor model should fit the data worse, again meaning that item wording factors are distinct across disparate measures.

Models in the current study. Model a = baseline model; model b = multimethod factor model; model c = one-method factor model; model d = examination of the latent correlation between social desirability response style subscales and other constructs. O = Openness to experience; C = Conscientiousness; E = Extroversion; A = Agreeableness; N = Neuroticism; S = Social dominance orientation; T = Team meeting attitude; IM = impression management subscale of social desirability; SD = self-deception subscale of social desirability. The subscript “p” represents positively worded items; the subscript “n” represents negatively worded items; the subscript “m” represents a method factor related to item wording. Due to space limitation, we do not show all item indicators under a construct. All positively and negatively worded items for openness, for example, are shown as Op and On. Disturbances and item residuals are not shown for clarity. All analyses were conducted at the item level. The latent constructs of impression management and self-deceptions were allowed to covary in model d.
Next, we examined the consistency of correlations between negative wording factors and external constructs in sample 1. The external constructs were the two facets of social desirability, namely, impression management and self-deception. Because participants are from the same source, a different pattern of correlations with different item wording factors cannot be attributed to sample variations. If different patterns of relationships between these wording factors and social desirability response styles are found across distinct substantive constructs (e.g., extroversion, openness, and SDO), Hypothesis 2 would be supported. Such results would signify that nomological network investigations for the wording factor of one measure cannot generalize to the wording factor of another.
Finally, we examined the correlations between negative wording factors and social desirability in sample 2 and compared the magnitude of these correlations with those of sample 1 using multigroup SEM technique (Figure 1, model d). Although measures in sample 2 overlapped with sample 1, they were made years apart. If the nomological network investigation results are indeed scale-specific, as predicted by our hypothesis, we would expect the negative wording factors from the same scale to correlate with social desirability response style facets similarly even across two different samples. In contrast, if the nomological network investigation results are sample-specific, we would expect the negative wording factors from the same scales to correlate with social desirability facets differently across two distinct samples. The scaled χ2 difference test (Satorra 2000) would show whether the corresponding correlations in the multigroup SEM were statistically invariant between two samples.
Before conducting this correlation invariance test, we need to ensure that the scales have the same factor loadings (i.e., metric invariance) between our two heterogeneous samples of respondents. The common χ2 difference test is known to be oversensitive to sample size for this purpose. Alternative fit indices are thus employed. Based on the recommendation by Chen (2007), Δ comparative fit index (CFI) ≥ −.010, together with Δ root mean square error of approximation (RMSEA) ≥ .010 or Δ standardized root mean square residual (SRMR) ≥ .030, would violate metric invariance.
Results
Intercorrelations Among Wording Factors
Hypothesis 1
The first hypothesis states that wording factors among various scales are scale-specific. Therefore, we would expect that the wording factors from various survey instruments to be weakly correlated with one another. We first set up a model without wording factors (see Table 2 and Figure 1, model a), and the fit of the model was good. Then we modeled wording factors from each measure and allowed them to be intercorrelated (Figure 1, model b). The fit of this multiwording factor model was a significant improvement over the previous model, based on scaled χ2 statistics (Table 2, model 1 vs. model 2). The intercorrelations among the construct factors and among the wording factors are shown in Table 3. The correlations of the wording factors range from virtually 0 to moderate [−.02, .54], supporting the hypothesis that item wording effect is generally scale-specific. When we assumed all the negatively worded items loaded together on only one-wording factor (Figure 1, model c), this unidimensional method model fits significantly worse than the previous multidimensional method model, based on scaled χ2 difference test statistics (Table 2, model 2 vs. model 3). Taken together, all these results suggest that wording factors among different measurement instruments do not measure the same construct. Our first hypothesis is thus supported.
Model Fit.
Note: N = 1,406. Scaled χ2 difference test was used for model comparison; the better model is bolded and underlined. RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual.
Construct Intercorrelations and Method Factor Intercorrelations in the Multimethod Factor Model.
Note: N = 1,406. r = item scores have been reversed so that a higher score represents a more positive (favorable) construct.
*p < .05. **p < .01. ***p < .001.
Correlations of Wording Factors With External Variables
Hypothesis 2
Based on the first hypothesis that the nature of wording effects is scale-specific, the second hypothesis states that their correlations with external variables (such as social desirability response bias) would also be scale-specific and show varying magnitude. We included the two latent factors of social desirability response styles (impression management and self-deception) and allowed them to correlate with the item wording factors. Notice that we continued to use items (instead of parcels) as observed indicators in these models. The correlations between social desirability response bias and a substantive construct are shown in the top panel of Table 4. The two facets of social desirability response style demonstrate from moderately negative correlations to slightly positive correlations with item wording factors (rs = [−.55, .11] for impression management; rs = [−.32, .16] for self-deception) in sample 1. This means that even though all data came from the same participants in sample 1, results from one scale are hardly generalizable to another scale. For decades, researchers used nomological network investigations from various scales to illuminate the nature of the item wording effect. The present finding helps explain why a consistent conclusion was never achieved.
Correlations Between Social Desirability Response Style and Item Wording Factors.
*p < .05. **p < .01. ***p < .001.
Hypothesis 3
The third hypothesis was that if item wording effects are scale-specific but not entirely sample-specific, wording factors from the same survey instruments would show a similar pattern of correlations with external constructs across two distinct participant samples. Therefore, we first conducted a preliminary analysis on how item wording factors correlated with social desirability response style in sample 2 and compared the results with those of sample 1. Samples 1 and 2 data were analyzed separately because they did not contain exactly the same set of scales. The magnitude of these correlations, shown in Table 4, appears to be surprisingly similar across samples. For example, the wording factor from agreeableness correlated strongly with impression management but not with self-deception in both sample 1 (rs = −.55 for impression management, −.08 for self-deception) and sample 2 (rs = −.63 for impression management, −.01 for self-deception).
To formally test the similarity of these sets of correlations between samples, we followed up the previous analysis with a two-group comparison model. All measurement scales were included in the comparison model except the TMA scale, as it was available only in sample 1 (see also Figure 1, model d). Both configural invariance (χ2 = 31,195.24, df = 10328, p < .001, RMSEA = .04, RMSEA 95% CI [.040, .041], SRMR = .06) and metric invariance (ΔCFI = −.002, ΔRMSEA < .001, ΔSRMR < .001) were achieved in the two-group comparison model (see Table 4), suggesting that the scales were interpreted similarly across the two samples. The correlations of the keying factors with impression management (Δχ2 = 7.92, Δdf = 6, p = .24) and self-deception (Δχ2 = 6.48, Δdf = 6, p = .37) were also statistically invariant between samples. These findings thus lend strong support to the third hypothesis: Wording factors from the same measures show a similar pattern of correlations with external constructs across samples.
As a supplemental analysis, we estimated the correlation between an overall wording factor and social desirability. We did so by including a second-order factor with all the wording factors from individual scales as indicators before correlating it with the two facets of social desirability. The correlations with impression management and self-deception were far from perfect (−.49 and −.24 in sample 1; −.69 and −.08 in sample 2), meaning that processes other than social desirability response style may be involved in the item wording effect.
Discussion
After decades of research, we still do not fully understand the item wording effect. The current study illuminates a potential reason for this state of affairs—namely, that the item wording effect is scale-specific. Several findings are noteworthy. First, even among the same group of respondents, item wording effects extracted from different scales are heterogeneous. Second, the extent to which wording effects correlate with external variables depends on the measurement instrument used. In the current study, some instruments showed strong correlations with social desirability response style whereas others showed weak or nonsignificant ones. Thus, findings with one scale do not necessarily generalize to those with others. Third, across two samples, for any given scale, an item wording factor showed a consistent relationship with an external variable.
Overall, the item wording effect showed both strong consistency (in the examination of the same survey instrument across samples) and large variability (in the examination of different survey instruments within a sample). Therefore, drawing a conclusion about the nature of the item wording effect from only a single survey instrument is almost futile, as the results are unlikely to be generalizable to other survey instruments. The current finding thus reveals a serious limitation in previous research in this area, which brings us to two important questions: How should we reinterpret previous findings on the item wording effect? And where do we go from here?
How Should We Reinterpret Previous Findings?
We suggest that researchers be fully aware of the scale specificity of the item wording effect in their investigations. Currently, researchers tend to assume that the item wording effect found in a particular survey instrument is generalizable to another instrument, leading to many inconsistent findings. As mentioned, Alessandri et al. (2011); Herzberg, Glaesmer, and Hoyer (2006); Rauch et al. (2007); and Vecchione et al. (2014) investigated the nature of item wording factors with a popular optimism measure, Life Orientation Test–Revised (LOT-R). They found that the wording factor associated with positively worded items was positively correlated with social desirability response style (Rauch et al. 2007), emotional stability (Alessandri et al. 2011), and egoistic bias (Alessandri et al. 2011; Vecchione et al. 2014), and negatively correlated with depression (Vecchione et al. 2014; see also Herzberg et al. 2006, for similar findings). Therefore, their results suggested that in conditions where the same measure (LOT-R) is used, wording effects associated with optimism items are related positively to self-enhancement and negatively to psychological disorder. Contrary to these researchers, we do not believe that these results are generalizable to other survey measures. Social desirability response style has been found to be unrelated to the wording effect associated with Rosenberg’s self-esteem scale (DiStefano and Motl 2006, 2009). Inconsistent findings across different measures seem to be the rule rather than the exception, implying that the nature of item wording factors in the optimism and self-esteem measures are qualitatively different.
Where Do We Go From Here?
If, as suggested here, the item wording effect is scale-specific, we envision two lines of future research. The first is to examine the nature of the effect associated with just one particular survey instrument. For example, what is the nature of the wording factor associated with LOT-R? This kind of research may be particularly important in order to illuminate the nature of measurement instruments commonly used for research and assessment. For example, given the consistent finding that the factor related to item wording in LOT-R is strongly correlated with social desirability response style and self-serving bias (Alessandri et al. 2011; Herzberg et al. 2006; Rauch et al. 2007; Vecchione et al. 2014), the factor may indeed represent a certain kind of response bias. However, researchers should be aware that comparable empirical findings from one survey instrument may not be found with another instrument. For example, the wording factor related to Rosenberg’s self-esteem scale has been found not to correlate with social desirability response style (DiStefano and Motl 2006).
The second line of research is to examine the general item wording effect by including multiple scales. What is the nature of the wording factor associated with a wide range of survey instruments? What is the nomological network associated with this general wording effect? Can the result reasonably be generalized to survey instruments not included in a study? Researchers may model multiple wording factors from individual survey instruments and a second-order latent factor with multiple wording factors as indicators. The researcher can then correlate the second-order factor with potential antecedents or consequences of the item wording effect. In addition, researchers could study the moderators for the correlation between a wording factor and a given external correlate. For instance, even if an external construct such as social desirability response style has a relationship with an overall wording factor, the relationships between social desirability response style and individual item wording factors (i.e., wording factors from individual scales) likely vary from null to strong (as in the current study). Researchers could usefully follow up this result by studying the moderating process behind these varying correlations. Although previous research has not examined these questions, we encourage future researchers to investigate them across a wide range of survey instruments in addition to popular candidates such as self-esteem and optimism.
In terms of replication research, the present results imply that when a researcher studies the general effect of item wording on survey instruments, using multiple samples of distinct respondents may not by itself ensure an adequate test of generalizability. As shown here, findings from the same measurement instrument may be well replicated across different samples (provided the sample size is large enough, as here) without there being replicability across different scales. Researchers should thus be wary of generalizing the item wording effect from only one survey instrument.
If the item wording factors from different measurement instruments are qualitatively different, implications follow for factor modeling. Our findings suggest a complex, multidimensional view of the item wording effect. Specifically, when examining the overall effect of item wording with external constructs, researchers may consider a second-order factor model in which items from each scale load on independent item wording factors, which in turn load on an overall latent factor. The independent item wording factors capture variances from items of distinct scales, whereas the overall latent factor captures common variance among individual wording factors.
In addition to examining the correlates of an item wording factor, researchers may examine the structure of the factor more closely. Previous researchers usually allow the factor loadings of an item wording factor to vary, assuming that item wording has differential impact on individual items (e.g., Alessandri et al. 2011; DiStefano and Motl 2006, 2009; Marsh et al. 2010; Rauch et al. 2007; Vecchione et al. 2014). However, little is known about the actual structure of the item wording factor. Is it possible to model an item wording factor with a more parsimonious structure, such as a tau-equivalent structure (i.e., identical factor loadings) without losing much information? Advancing knowledge regarding the factor structure of the item wording effect is important because it will further improve our understanding of the effect’s nature.
As in any study, the current research has limitations. We employed university students, which limits generalizability to other populations. We included only one scale—social desirability response bias—as an external correlate of the item wording effect. Similarly, we employed the CT(M − 1) model rather than other multitrait-multimethod models such as the CTCM. CTCM has been advocated by some researchers (e.g., Castro-Schilo, Widaman, and Grimm 2013; Lance, Noble, and Scullen 2002) while criticized by others (e.g., Geiser, Koch, and Eid 2014; Pohl, Steyer, and Kraus 2008). Our position is not to advocate one model over another but simply to demonstrate the results using a method. We speculate that results may vary with populations other than students, scales other than social desirability, and methods other than CT(M − 1). However, the main conclusion—that item wording effects can differ across measurement instruments—would nonetheless remain the same.
Conclusion
The current study provides an answer to the perplexing question of why three decades of research has not yielded a thorough understanding of the item wording effect (Carmines and Zeller 1979). The answer is that the item wording effect differs across survey instruments; consequently, results from one study, using a given instrument, will not generalize to another study, using a different instrument. We trust that future researchers will keep this in mind.
Footnotes
Acknowledgment
The data were collected when the author was a graduate student at the University of Western Ontario. This work was performed in part at the High Performance Computing Cluster (HPCC) which is supported by Information and Communication Technology Office (ICTO) of the University of Macau.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work has been financially supported by Multi-Year Research Grant (MYRG2015-00076-FED) offered by the University of Macau.
