Abstract
A sample of 516 participants responded to the Balanced Inventory of Desirable Responding (BIDR) under answer honest and instructed faking conditions in a within-subjects design. We analyze these data with a novel application of trifactor modeling that models the two substantive factors measured by the BIDR—self-deceptive enhancement (SDE) and impression management (IM), condition-related common factors, and item-specific factors. The model permits examination of invariance and change within subjects across conditions. Participants were able to significantly increase their SDE and IM in the instructed faking condition relative to the honest response condition. Mixture modeling confirmed the existence of a theoretical two-class solution comprised of approximately two thirds of “compliers” and one third of “noncompliers.” Factor scores had good determinacy and correlations with observed scores were near unity for continuous scoring, supporting observed score interpretations of BIDR scales in high-stakes settings. Correlations were somewhat lower for the dichotomous scoring protocol. Overall, results show that the BIDR scales function similarly as measures of socially desirable functioning in low- and high-stakes conditions. We discuss conditions under which we expect these results will and will not generalize to other validity scales.
Keywords
Faking in questionnaire responses is considered such a problem that remarkable levels of effort are devoted to addressing it. Validity scales are among the most common methods for detecting dissimulation on self-report questionnaires, despite some controversy regarding the extent to which they can detect dissimulation and be used to correct for it (McGrath et al., 2010; Morey, 2012; Rohling et al., 2011). The Marlowe Crowne Social Desirability Scales (MCSDS), for instance, which are based on the Minnesota Multiphasic Personality Questionnaire (MMPI), differentiate attribution, or claiming desirable characteristics, from denial, or disclaiming undesirable characteristics (Millham, 1974). Leite and Beretvas (2005) reviewed the ways that the MCSDS are used in practice. They are commonly used in three different ways. First, substantive trait scores are correlated with social desirability scale scores to see whether the correlations are high, indicating socially desirable responding impacted the assessment process. Second, factor analysis is sometimes used to check whether the construct of interest and social desirability are empirically distinct. Finally, the substantive assessment scores for respondents with high social desirability scores are sometimes ruled invalid.
Another well-known scale for detecting social desirability that is used in similar ways to the MCSDS is the Balanced Inventory of Desirable Responding (BIDR: Paulhus, 1984; Paulhus et al., 1995). In early writings, Paulhus traced the history of researchers who have distinguished two forms of socially desirable responding, self-deceptive enhancement (SDE) and impression management (IM), in questionnaire responses. SDE refers to when a respondent really believes their inflated responses, whereas IM occurs when a respondent consciously inflates responses (Paulhus, 1984). According to Gignac (2013), the former behaviors are observable only to the self while the latter are observable to others. Paulhus identified scales that had historically been used as markers of each form of dissimulation and noted that factor analysts who initially discovered these tendencies referred to them as alpha and beta (e.g., Block, 1965; Wiggins, 1964). Paulhus (1984) factor analyzed the scale scores for questionnaires thought to be markers of alpha and beta to create the BIDR. Readers may refer to Paulhus’s original work or to Leite and Beretvas (2005) for a brief history of the origins of the BIDR, including its psychodynamic roots that were subsequently dropped.
The BIDR scales have now been extensively evaluated, including testing under instructions to fake and not to fake. This work has primarily occurred at the observed variable level with classical test theory rather than with latent variable models. Despite an abundance of research articles on the BIDR, the number of articles that explicitly focus on the measurement invariance of the BIDR under low- and high-stakes conditions is limited to a few papers at most. Yet, the factorial invariance of the BIDR across settings where people are not expected to dissimulate and where they might is critical to the validity of the measure. For the scales to be useful in measuring the extent of socially desirable responding, the measurement parameters of the confirmatory factor models (i.e., loadings and intercepts or thresholds) ought to be the same between honest and faking conditions with only the population heterogeneity parameters (i.e., latent factor means and variances) changing between the conditions. That is, to measure the extent of SDE and IM, the scales should “work” equally well under low- and high-stakes conditions. This is even more important for social desirability scales than it is for other constructs. After all, it is social desirability scales that researchers purport to be measures of faking. If they functioned differently when people faked, it would be analogous to a ruler functioning differently when measuring shorter as opposed to longer distances. Whether or not the measurement parameters vary as a function of instruction conditions is an open question and we explore this as a research question rather than a directional hypothesis. There have been few comprehensive investigations of measurement invariance for validity scales across honest and faking conditions to our knowledge.
In fact, we found just one instance where a study examined measurement invariance for the full BIDR across low- and high-stakes conditions (Li & Reb, 2009). This study used a multiple group approach with a within-subjects design, which violates the assumption that the groups are independent populations and renders the conclusion challenging to interpret. Instead, a single-group longitudinal invariance model should be applied with within-subject designs, as discussed by authors, including Chan (1998), Marsh and Grayson (1994), and Liu et al. (2017). A second study using a between-groups design found support for measurement invariance for the IM scale of the BIDR. This study was designed to examine the effect of anonymity versus confidentiality assurances (Miller & Ruggs, 2014). However, it omitted the self-deception scale. The main objective of this article is to appropriately examine the measurement invariance of the BIDR across honest and faked responses using a within-subjects instructed faking design. In particular, we will determine whether any changes in item means and covariances between conditions affect measurement parameters, population heterogeneity patterns, or both.
Before invariance over experimental conditions can be examined, one needs to obtain a well-fitting measurement model for the BIDR. Few multiple factor psychological measures are truly orthogonal and it is likely that the BIDR factors of SDE and IM are correlated. In situations where there are three or more substantive factors, potential measurement models would include a higher-order factor model, a correlated factors model, or hierarchical models.
Higher-Order Models
Higher-order models model the relationships between factors with a higher-order factor that explains the variance in lower-level factors. As Paulhus (1984) originally proposed a two-factor model of socially desirable responding, a higher-order factor model is not identified without modeling additional variables or including model constraints beyond typical identification constraints when three or more latent variable indicators are available. This leaves the correlated factors model and hierarchical factor models (Holzinger & Swineford, 1937; Markon, 2019; Reise, 2012) as plausible models.
Correlated Factor Models
Correlated factor models represent the relationships between measured factors with nonzero correlations. The correlated factors model has not proved to be the best-fitting model for the BIDR in past research, which has raised concerns about the BIDR structure (e.g., Gignac, 2013; Leite & Beretvas, 2005). We expect that this is because of at least two reasons. First, we expect that a general latent factor representing nonuniform response biases might be necessary (Brown et al., 2017). Second, we expect that unique item factors might be required, albeit that item-specific factors are not identifiable in typical self-report single-occasion responses (e.g., Rao, 1955).
Hierarchical Models
Hierarchical models assume that all factors influence items directly, as opposed to indirectly via subordinate latent factors in the case of higher-order factor models. In the bifactor model, one form of hierarchical models, a general factor is fitted along with an orthogonal subset of homogeneous group factors accounting for variance not explainable by the general factor. The group factors can themselves be correlated or uncorrelated (Holzinger & Swineford, 1937; Markon, 2019; Reise, 2012).
Caution has been suggested in interpretation of bifactor models as revealing substantive group factors as opposed to its more conventional use for examining the extent to which a measure provides a unidimensional score. Reasons include the difficulty of interpreting the general factor as a causal factor; the tendency of bifactor models to improve fit by modeling construct-irrelevant variance, and the fact that good model fit does not indicate validity of the latent variables in the bifactor model (Bonifay & Cai, 2017). Nonetheless, these are popular models and have been applied with BIDR data.
One application of bifactor modeling to the BIDR was presented by Gignac (2013) who tested an extensive array of models using a low-stakes sample that included Paulhus and Reid’s (1991) updated BIDR structure, where SDE is split into SDE and self-deceptive denial. Gignac compared the fit of numerous models where the data were dichotomously scored and “continuously” scored (i.e., treated as ordinal rather than binary, in both cases he modeled the data using a diagonally weighted least squares [DWLS] estimator).
In that study, the best-fitting model for the BIDR’s continuous scoring was a hierarchical model that included (a) a general social desirability factor; (b) orthogonal-specific factors for SDE and IM; and (c) a method factor that modeled all items that were negatively keyed. Considering the earlier concerns about giving substantive interpretations to methods factors, readers might think twice today about the validity of a bifactor representation of the BIDR with substantive group factors. At the same time, the bifactor model is not up to the complexity of the task of modeling BIDR responses in an instructed faking design, where the same participants answer honestly and under faking instructions.
However, while the modeling task becomes more complex due to the introduction of a within-subjects repeated-measures design, possibilities are opened by the additional experimental condition that enable identifying different sources of variance in the response process. Namely, unique item factors are now identifiable with two measurement occasions corresponding to participant responses under each instruction condition. Moreover, the latent mean differences between experimental conditions can be identified and, if appropriate, interpreted. In this article, we capitalize on this opportunity to present a novel application of trifactor modeling (Bauer et al., 2013), which is itself an extension of the bifactor model.
Trifactor Models
Bauer et al. (2013) presented trifactor modeling in the context of multi-informant designs, suggesting that responses of multiple informants answering about a single construct could be explained by a general factor measuring the construct of interest, rater factors measuring the unique perspective of each rater source on the latent construct, and item-specific factors that measure unique variance associated with each item across the informant groups.
We adapted Bauer et al.’s approach in the following ways. We modeled two substantive BIDR factors in each experimental condition. Hence, the model included honest SDE and faked SDE and honest IM and faked IM factors. These are the ultimate constructs of interest when one uses the BIDR because they reflect the extent of respondents’ SDE and IM under low- and high-stakes assessment conditions. Different to Bauer et al. (2013), we incorporated a mean structure for the BIDR substantive factors, permitting the interpretation of the experimental difference in latent means resulting from the faking intervention. To this end, we imposed strict measurement invariance across conditions and fixed the means of SDE and IM in the honest condition to zero while estimating them freely in the faking condition. The measurement invariance constraints are described in the “Methods” section. Next, we modeled further dependencies in item responses in each condition with two method factors—thus, all 40 BIDR items rated under the honest condition loaded on a “general Honest” factor and all 40 BIDR items rated under the faking condition loaded on a “general Faking” factor. Brown et al. (2017) suggested that such general factors can be used for capturing response biases. If the factor loadings are similar across items, they can be considered uniform distortion (i.e., response styles), such as acquiescence, whereas if they are different across items (e.g., follow the pattern of positive and negative loadings in the BIDR’s balanced design), they can be considered nonuniform biases, such as socially desirable responding. However, in the present study the latter is measured directly via the BIDR, so the general factors are intended to pick up any remaining variance due to response styles. Luckily, the BIDR’s balanced design provides a unique opportunity to separately identify the substantive factors SDE and IM (with half of the items expected to load negatively), and the method factors as response styles (with all items expected to load positively). Finally, dependency that is due to the same item being answered in two instructional conditions was modeled as item-specific factors, operationalized as correlated errors of the same item across the two conditions. This parameterization is mathematically equivalent to the item-specific factors with both factor loadings set to unity and freely estimated variance, producing identical fit. This trifactor model variation is represented graphically in Figure 1.

Diagram of the Single Class Within-Subjects Trifactor Model.
While modeling method-related general factors on top of the substantive factors often cause problems with model identification (e.g., Podsakoff et al., 2003), the repeated-measures design in this study allows identification of substantive factors by imposing strict measurement invariance across conditions, which in turn allows assigning the mean differences between BIDR items in honest and faking conditions to the latent SDE and IM mean shifts rather than to the method factors. This would not be possible without the repeated-measures design. To further facilitate this separate identification, as is recommended in the bifactor modeling literature (e.g., Reise, 2012), we set the method factors uncorrelated with the substantive factors. The method factors, however, should correlate with each other if they are to capture the same person’s response styles under different instructions. Similarly, the substantive factors should correlate with each other if they are to capture the same person’s standing on SDE and IM under different instructions.
With reference to the concerns with interpretation of bifactor models mentioned earlier, it is important to note that we do not make any causal interpretation of the BIDR factors that were not originally proposed by Paulhus. The hypothesized substantive factors in each condition are still IM and SDE. The trifactor model is simply a technical adaptation that allows us to interpret the impact of faking on (a) the measurement properties of the self-deception and IM scales under high-stakes conditions, and (b) if appropriate, as determined by the invariance of the item parameters across instruction sets, the differences in latent means because of the experimental manipulation. However, this is not necessarily to say that method factors in the trifactor model could not be given a substantive interpretation that is not based on response styles at a later point given appropriate evidence.
Trifactor Mixture Models
In any experimental intervention, there is a chance that some participants will not follow instructions and we anticipate that in the faking condition, which imposes a significant cognitive burden on participants, some participants do not actually fake but respond normally. Hence, the distribution of scores in the faking condition may show heterogeneity as a result a mixture of honest and faked responding. We accommodate for this scenario with an adapted trifactor mixture model, which identifies two latent classes of individuals (Clark et al., 2013), such as people who followed the faking instructions and those who did not. Within the “compliers” class, the default trifactor model with the honest and faking conditions applies, and within the “noncompliers” class, both repeated measures in the trifactor model are modeled as the honest condition. Kim and von der Embse (2021) described the integration of trifactor modeling and mixture modeling in the context of Bauer et al.’s (2013) original formulation of the trifactor model. Here, we extend the application of trifactor mixture modeling to the repeated-measures designs to allow modeling simultaneously the dimensional structure of the BIDR across faking and nonfaking while also identifying classes of individuals who complied (compliers) with the instructions and those who did not (noncompliers).
Method
Participants
Here we report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study. Participants in this study were (a) 213 professional colleagues and students in the network of the author team and professionals in the working community in the United Kingdom who consented to participate after the survey was advertised to professional networks and on the LinkedIn website, and (b) 303 respondents from a panel survey company called, Cint, who were representative of British working adults. The study was approved by the ethics committee at the first author’s institution. Sample (a) received the instruction to fake first, followed by the instruction to respond honestly, and this order was reversed in Sample (b). The two samples were similar in some but differed in other potentially important respects. The largest demographic group was White in both cases, but the student sample was younger on average and its modal highest education was higher. We present a detailed demographic breakdown of the samples in Table 1, including age, gender, education, occupation, and ethnicity.
Sample Demographics.
There were significant differences in observed scale scores between the samples, presented in Table 2, that are impossible to clearly attribute to the different sampling mechanisms or the fact that the first sample was asked to answer honestly before faking rather than vice versa. Our conjecture, although impossible to prove, is that the bigger shifts in means between conditions for the student sample were because they first had the opportunity to answer the items honestly before being asked to fake. Nonetheless, in both samples, the means shifted due to the experimental conditions in the same direction. We analyze the sample overall, as each sample alone is too small to analyze with the trifactor model set up we use. Most interpretations of sample size requirement for confirmatory factor models require at least 400 participants, particularly with the maximum likelihood with robust standard errors (MLR) estimator with missing data (Savalei & Bentler, 2005; Yuan & Bentler, 2000) and a highly parameterized model.
Observed Score Scale Means and Standard Deviations.
Note. Continuous refers to the sample statistics when the continuous scoring protocol for the BIDR is followed; dichotomous refers to the descriptive statistics when the dichotomous scoring protocol of the BIDR is followed. BIDR = Balanced Inventory of Desirable Responding; SDE = self-deceptive enhancement; IM = impression management.
Measures
In this article, we analyze and report on Paulhus’s (1984) BIDR scales administered as part of a larger study. Paulhus discussed the use of both 5-point and 7-point Likert-type scale variations as well as dichotomous scoring schemes. In this study, we administered the BIDR with a 5-point scale ranging from strongly disagree (1) to strongly agree (5). In addition to demographic data questions described under the “Participants” section and the BIDR measures, participants completed a maladaptive personality measure, the G50 (Guenole, 2015), which is not analyzed here. The exact design reported in Guenole et al. (2018) was completed under honest and faked conditions.
Design and Procedure
Assessment stakes (low vs. high) were experimentally manipulated within subjects. In the “low stakes” (“honest”) condition, respondents were instructed to answer honestly as follows: “You are now going to be answering questions about your personality. We would like you to answer questions in this section as honestly and as accurately as possible. Please answer truthfully, all responses are anonymous.” In the “high stakes” (“instructed faking”) condition, respondents were asked to respond as though they were applying for a job they really liked, with the following instructions: “In this section we would like you to ‘fake’ your answers to the questions. In other words, answer as if you want to make your best impression to get a job you really want.”
Many research studies employing instructions to fake specify a particular job type so that all respondents are faking toward the same profile (Robie et al., 2007; Wetzel et al., 2021). The counter argument to using this type of instruction with heterogeneous samples is that this would likely lead to different people being better positioned to fake for the nominated job role than others, in addition to having very different motivations and attitudes toward the nominated job profile, which may lead to the lack of interest or motivation to follow the faking instruction in many participants. Moreover, we expect that heterogeneity of job roles that participants imagine should not be detrimental for the analysis of BIDR data, which reflect general rather than job-specific considerations of socially acceptable behaviors and unlikely virtues.
Analysis
All data and scripts used in this article are available at the following link: https://figshare.com/articles/dataset/Guenole_N_Brown_A_Lim_V_in_press_Can_faking_be_measured_with_dedicated_validity_scales_Within_Subject_Trifactor_Mixture_Modeling_applied_to_BIDR_responses_Assessment_/14725074. The measurement model fitted was a variation of the trifactor model first reported by Bauer et al. (2013) and discussed above. We identified the unit of measurement for the BIDR scales by identifying a strong loading item in each scale (the referent) in each condition and setting its factor loading to one in each condition.
Estimation
There can sometimes be confusion when BIDR models are described as continuously scored, which is contrasted with dichotomous scoring conventions that were proposed by Paulhus (1984), and when researchers discuss whether model parameters were estimated using continuous or categorical estimators. We considered 5-point Likert-type data in this study because we were interested in modeling the response process and identifying various variance sources. In this section, therefore, references to continuous and categorical refer to parameter estimation methods for the 5-point Likert-type data that we analyze, not to any BIDR scoring protocols. It is possible to model these response data as continuous or ordered categorical, using MLR or a diagonally weighted least squares (DWLS) estimator, respectively, implementing the appropriate identification constraints for repeated-measures measurement invariance (Liu et al., 2017). On one hand, these data are certainly ordinal, suggesting DWLS. On the other hand, it is unlikely that response tendencies underlying the observed variables (assumed in all limited information estimators such as DWLS) are multivariate normal, particularly in the instructed faking condition.
Inspection of the item distributions for the continuous scoring protocol, presented in Figure 2, reveals an increase in endorsement of extreme categories in the instructed faking condition, making the distributions of item responses heavily skewed. This likely reflects the respective skewness of the underlying response tendencies (which represent utilities or psychological values that people feel toward the items). We can see this same pattern in the observed scale totals in upper panel of Figure 3. The categorical analysis will, however, treat the increased frequencies in extreme categories by widening the boundaries of these categories to preserve the normality of the response tendency, thus totally distorting the estimates for thresholds and polychoric correlations. A similar point was made by Robitzsch (n.d.). Using ML with continuous data shifts the assumption of multivariate normality to the observed level, but MLR is robust to nonnormality in the observed variables. As we expect that that the continuous response interpretation is more representative of the actual response process due to the expected latent utility nonnormality, we analyzed these data with the maximum likelihood estimator with robust standard errors in Mplus 8.6 (Muthén & Muthén, 2009).

Observed Item Distributions (After Reverse Scoring) Under Honest (Left Panel) and Faking (Right Panel) Instructions (1 = Strongly Disagree, 5 = Strongly Agree).

Observed Score Distributions for Continuous and Dichotomous Scoring Under Honest and Faking Conditions.
Data Cleaning and Missing Data
Some participants failed to complete some or all of the honest or faking instruction conditions. As a result, we undertook the following data screening analyses. First, we removed the respondents who had missing data for more than 50% of questions in either the honest or the instructed faking condition. Even optimistic interpretations of missing data treatment approaches are expected to struggle with greater missing data than this. Second, we eliminated any participants who completed the entire survey (including the additional maladaptive survey items) in less than 10 minutes. We deemed it implausible that a respondent could pay due attention to the questionnaire in such a short time. All pairwise samples following this data screening were greater than 95% and missing data were handled with full information maximum likelihood (Enders & Bandalos, 2001).
Model Fit
We first examined the fit of the baseline model from which all invariance tests would be conducted. We examined chi-square, the root mean square error of approximation (RMSEA; Steiger, 1998), the comparative fit index (CFI; Bentler, 1990), the Tucker–Lewis index (TLI; Tucker & Lewis, 1973), and the standardized mean root square residual (SRMR; Jöreskog & Sörbom, 1989). A significant p value for the chi-square indicates rejection of the fitted model. For the other indices, values observed are compared against established cut-offs (Hu & Bentler, 1998, 1999). For RMSEA, .05 has been suggested as indicating close fit, and .08 for adequate fit. For CFI, a cut-off .95 is considered for good fit and .90 for adequate fit. For SRMR, a value of .08 or less is considered acceptable. In addition to looking at these global fit statistics and indices, we inspected the size, sign, and significance of the parameter estimates themselves (e.g., item loadings and latent correlations), as well as the correlation residuals and modification indices (Kline, 2015).
Measurement Invariance
We implemented longitudinal mean and covariance structures (LMACS; Chan, 1998) approach appropriate for the repeated-measures design employed in this study. In this approach, the repeated-measures factor invariance model is identified by fixing the factor loading of a single referent item to one at each timepoint (i.e., condition), constraining the intercept of the referent item equal across conditions/times, fixing the latent factor mean of the first timepoint (condition) to zero, and freely estimating it in the other. The referent item for each factor was identified as an item that loaded strongly in each condition. To test for invariance at the item level from this baseline, we used the free-baseline method described by Stark et al. (2006).
This approach imposes additional item constraints to the free-baseline model one by one, constraining each item’s loading and intercept simultaneously across conditions, and comparing fit with the free baseline that has only identification constraints in place. If the change in fit is statistically significant, then measurement noninvariance, or differential item functioning (DIF), is detected. We adopted a p value of .01 as our criterion for statistical significance given that we were making many comparisons in total in this sequence of analyses. Each item has three loadings, one for the substantive factor (either SDE or IM), one for the condition factor, and one for the item specific factor. We conducted item-level invariance analyses on the substantive factor loading (and intercept) only as only the invariance of the constructs measured by the BIDR itself can be plausibly expected. The invariance for the condition factor is not assumed and the invariance of the item-specific factors is given by their fixed to unity loadings. Tests were conducted item by item and we applied Satorra and Bentler’s (2001) correction when calculating chi-square between nested models.
Mixture Modeling
We estimated the mixture model variation of the trifactor model using the MLR estimator that permitted comparison of the single-class trifactor model with partial measurement invariance we report with a two-class trifactor mixture model. Important differences between the complier class and the noncomplier class are that in the noncomplier class the latent means of SDE and IM in the faking condition are 0, just as they are for these factors in the honest condition. In contrast, an increase in the latent scores is expected between the same construct in different conditions for the complier class. Second, given that there is no change for the noncomplier class, the correlation between corresponding factors, SDE honest and SDE faked, IM honest and IM faked, is expected to be near 1, indicating no change in the participant rank ordering across conditions. In contrast, we expect correlations substantially less than 1 between corresponding factors across conditions in the complier class. Given our a priori anticipation of two latent classes, this application is considered a confirmatory trifactor mixture model.
To compare the fit of trifactor mixture models, for which chi-square-based fit indices are not available, to the ordinary trifactor models, we used information criteria—Akaike’s information criterion (AIC; Akaike, 1987) and the Bayesian information criterion (BIC; Schwarz, 1978). When alternative models are compared, the model with smallest AIC/BIC is considered best and the AIC difference of 10 or greater with the alternative model is interpreted as “providing no support” for the alternative model (Burnham & Anderson, 2004).
Finally, we examined entropy, which reflects the class separation. The higher the entropy, the clearer the class separation, and values of .80 and above have been suggested as indicating strong class separation (Asparouhov & Muthén, 2014). We consider all of these in evaluation of the confirmatory trifactor mixture model.
Results
Global Fit
Fit statistics for all models that we discuss in the following sections are presented in Table 3 to facilitate model comparison. The baseline (unconstrained) trifactor model included the two BIDR substantive factors in each condition, condition-related common factors, and specific factors representing the same item asked across conditions. We only included constraints required to identify the model including the latent mean difference between conditions. The fit for model, estimated with the continuous MLR estimator, was as follows: χ2 = 4,807.552, df = 2,953, p ≤ .01, RMSEA = .035 (90% confidence interval [CI] = [.033, .037]), CFI = .848, TLI = .838, SRMR = .053. While chi-square was significant, RMSEA and SRMR indices indicated adequate fit while incremental fit indices fall short of conventional cut-offs for good fit. The trifactor baseline model for the 80 item responses (40 BIDR items × 2 conditions) fitted significantly better than a correlated factors model with no condition factors or specific factors, for which the fit was χ2 = 7,924.564, df = 3,074, p ≤ .01, RMSEA = .055 (90% CI = [.054, .057]), CFI = .603, TLI = .592, SRMR = .079. Adding a negatively keyed method factor, which Gignac (2013) reported for his best-fitting model, marginally worsened the fit of the trifactor model. We concluded that the trifactor model of item responses treated as continuous was plausible.
Fit Statistics for Trifactor Models.
Note. The trifactor continuous model and the trifactor ordinal model used a continuous maximum likelihood estimator with robust standard errors and a diagonally weighted least squares estimator, respectively. The fit reported for the trifactor partial invariance model is the model where measurement invariance constraints are added to the trifactor continuous model. χ2 = chi-square; DF = degrees of freedom; RMSEA = root mean square of approximation; CFI = comparative fit index; TLI = Tucker–Lewis index; SRMR = standardized root mean square residual.
Readers may be interested in the fit of the ordinal trifactor model, which was as follows: χ2 = 4,850.728, df = 2,994, p ≤ .01, RMSEA = .035 (90% CI = [.033, .036]), CFI = .928, TLI = .924, SRMR = .057. 1 Gignac’s (2013) best-fitting model to the raw response data (i.e., without dichotomization) was χ2 = 1,379.21, df = 680, RMSEA = .047 (90% CI = [.046, .051]), CFI = .840, TLI = .817, and SRMR was not reported. Our model, which includes an additional 40 items representing the instructed faking condition, fitted better than Gignac’s model for honest responses only to 40 items on every fit criterion. It appears that with our nonnormal response data the ordinal estimator overfits the data by “normalizing” data that are actually nonnormal. Supporting the idea that the ordinal model is overfitted, Savalei (2021) has discussed the tendency of incremental fit indices to overestimate the fit of models with ordinal data.
Correlation Residuals
We examined the model’s residual correlations next. We considered the absolute value of residuals to see any that were above an absolute value .10, following Maydeu-Olivares and Shi (2017). Just 7% of the n × (n − 1) / 2 = 3,160 residual correlations were above .10. Of those that were larger than .10, the median (and mean) correlation was .12 and the maximum was .19, for a correlation residual between unrelated items across conditions. Other correlation residuals were for seemingly unrelated items within conditions—both within and across scales. Given that these represent a very small proportion of correlation residuals, that their absolute values were not excessive, and these modifications were not anticipated a priori, we did not incorporate these empirically driven revisions in our measurement invariance testing that we report below.
Modification Indices
In contrast to the residual correlations as indicators of model deficiencies in explaining particular intervariable covariances, modification indices provide more direct advice on where changes to the model might be required by pointing to specific parameters. Modification indices did not, on this occasion, indicate substantive changes that would improve the fit that we could have a priori anticipated. For example, the largest modification indices all pointed toward modifications that were contrary to Paulhus’s (1984) theory, such as allowing items to load on alternate factors or across conditions, or to changes that had no theoretical validity basis, such as residual correlations between seemingly nonrelated items.
Item Parameters
We also inspected the parameters of the model paying most attention to the item loadings in terms of sign, size, and significance and interpreting factor correlations in the same way. In Paulhus’s (1984) original specification, the first item of the self-deception scale is positively phrased, the second is negatively phrased, and the remainder continue alternating sign in this manner. The first IM scale item, in contrast, is hypothesized to be negative, the second is positive, and the remainder continue to alternate in this pattern. Inspection of the model estimated loadings on the substantive SDE and IM factors indicated that this pattern held for all except the seventh item of the self-deception scale, which should have been positive but was in fact weakly negative. We reserve presentation of the substantive factor loadings until the measurement equivalence section where we present the equated loadings.
General Factor Loadings
Brown et al. (2017) offered suggestions about how to interpret methods factors such as the general factors in this article. They suggested such general factors can be considered random additive effects of bias, for instance, that have been incorporated in the modeling of the response process (Böckenholt, 2012). The nature of the distortion can be uniform (e.g., acquiescence and leniency factors discussed by Maydeu-Olivares & Coffman, 2006) or nonuniform (e.g., the ideal employee factor described by Klehe et al., 2012). Nonuniform distortion would be implied by nonequal factor loadings across items, whereas the uniform forms of distortion would be suggested if the loadings on the common factor were equal across items.
The standardized loadings for the general factor in each condition for the trifactor model are presented in Table 4. A first observation is that the range of these loadings in each condition is narrow, from −.10 to .45. Variation is apparent in the loadings from inspection, and a formal test of the difference between the models where these loadings were freely estimated and a model where they were constrained to be equal significantly worsened fit. However, the loadings do not follow the pattern of an ideal employee factor because both the desirable and undesirable items measuring the SED scale have positive factor loadings. This pattern is less pronounced for the IM items, where the desirable items have mostly near-zero loadings and the undesirable items have mostly positive loadings. The near-zero loadings for the desirable IM items provide reassurance that the general factors do not capture the “faking” variance because if this were the case, the desirable IM items would be affected most as ultimate indicators of IM behavior. This is good news because in the trifactor model, the change in substantive factors SDE and IM is expected to capture the faking effect. It seems likely that because there is restricted variability of loadings, the general factor in both conditions is a mix of content-dependent acquiescence and, despite that we screened fast response times, inattentive responding. We return to ways that the inattentive responding might be eliminated in future research in our discussion.
Standardized Factor Loadings for Experimental Condition (Method) Factors.
Note. Keywords in columns may be used to match items to the original BIDR items. Keywords are presented for honest condition only as they are the same for the faking condition. General honest refers to the loadings of items on the general factor in the honest condition. General faking refers to loadings of items on the general factor in the faking condition. HS = honest condition for self-deceptive enhancement; HI = honest condition for impression management; FS = faking condition for self-deceptive enhancement; FI = faking condition for impression management; BIDR = Balanced Inventory of Desirable Responding.
Measurement Invariance for Substantive Factors
With the baseline model established, we proceeded with item-level invariance tests that simultaneously examined loadings and intercepts. These results are presented in Table 5. These results indicate partial invariance because while for the majority of items the p values are below .01, there are three items for SDE and three for IM that showed significant chi-square difference tests. Six items might seem like a large number, but it should be considered as a proportion of the total of 40 BIDR items (15%). We conducted a series of follow-up measurement invariance tests to examine whether differences on the slope or intercept gave rise to the significant combined tests reported in Table 5. We first conducted a loading invariance test, and only if the metric (loading) invariance test was nonsignificant, conducted a metric (intercept) invariance test. Following these additional tests for each of the items that were initially identified as noninvariant, we attempted a qualitative examination of causes for the difference in item functioning. It is important to note that there may not be any obvious reason for the measurement noninvariance, and hence, we cautiously offer here potential reasons. If the BIDR scales were to be reviewed, we would recommend expert panel reviews that consider all items but focus on the items that showed noninvariance items as a starting point.
Item Level Measurement Invariance Results.
Note. SD is self-deceptive enhancement; IM is impression management. χ2 is the chi-square model fit; χ2 diff is the difference in chi-square between the baseline model where identification constraints only are imposed and the model where an item is constrained to have the same loading and intercept across honest and faking conditions. SD20 is the last item of the SD scale that was selected as the anchor item because it had strong loadings across conditions whereas Item SD1 did not. The first item of the IM scale did have strong loadings in both conditions and was therefore selected as the referent for this scale. There is no χ2 diff for the first anchor item listed in each scale because that item provides the model fit from which all other differences are calculated.
Item is not invariant with regard to intercept; **item is not invariant with regard to factor loading. Degrees of freedom for the baseline chi-square in Row 1 of the table is 2,953 while degrees of freedom for the nested models is 2,955. WLSMV correction value was 1.13. Keywords in columns may be used to match items to the original BIDR items. BIDR = Balanced Inventory of Desirable Responding.
For the SDE scale, the first item to exhibit noninvariance was Item 1, “My first impressions of people usually turn out to be right.” The loading was invariant but there was noninvariance on the intercept, with the expected item score for a person scoring zero on SDE lower in the faking condition. One possible reason may be that agreeing with this statement is seen as arrogant and undesirable in high-stakes employment situations. The next SDE item exhibiting noninvariance was Item 3: “I don’t care to know what other people really think of me.” With this item, the noninvariance was due to the loading. Interestingly, Item 3 went from being a good indicator of the SDE construct in the honest condition to almost unrelated to SDE in the faking condition. A possible explanation would be that under high-stakes conditions, social desirability considerations do not apply to this item and response decisions are based solely based on IM considerations (which would result in almost universal rejection of the item). Finally, on the SDE scale, Item 13, “The reason I vote is because my vote can make a difference” again showed the loading noninvariance. In contrast to Item 3, however, the strength of the loading went from moderate in the honest condition to strong in the faking condition. It seems that in a high-stakes setting, the item triggered more consideration for SDE than in a low-stakes setting.
There were also three items that showed noninvariance on the IM scale. First, Item 5, “I sometimes try to get even rather than forgive and forget” showed loading noninvariance. The loading was more strongly negatively related to the underlying IM factor in the faking condition, indicating its greater importance for impression-management considerations in a high-stakes setting. In contrast, the final two items on the IM scale showed intercept noninvariance. These were Item 6, “I always obey laws, even if I’m unlikely to get caught,” and Item 10, “I always declare everything at customs.” Both items had lower intercepts in the faking condition, indicating perhaps less enthusiasm for producing extreme scores where simply hiding illegal behavior would do, in comparison with other items in high-stakes settings. We note that these minor intercept decreases are after controlling for the very large IM score inflation from the honest to the faking condition and that the means of these items are much higher in the faking condition.
We estimated a final partial invariance model for these data based on the measurement invariance results, allowing item loadings and intercepts to freely vary for six items where invariance analyses produced a significant decrement in fit. The fit of this model was χ2 = 5,053.594, df = 3,018, p < .001, RMSEA = .036 (90% CI = [.034, .038]), CFI = .833, TLI = .826, SRMR = .057. After achieving partial invariance, the standardized latent means in the faking condition were 1.163 on the self-deception scale and 1.322 on the IM scale. Hence, the positive latent means represent experimental effects in the positive direction, with both SDE and IM showing large increases (score inflation) in the faking condition. The standardized loadings for the substantive factors for the partial invariance model are presented in Table 6.
Standardized Factor Loadings for Substantive BIDR Factors in Honest and Faking Conditions.
Note. SD is self-deceptive engagement; IM is impression management.
Item is not invariant with regard to factor loading. For all items except those marked with ** the unstandardized loadings are equated, loadings in the table differ across conditions only due to standardization. For unstandardized loadings, please see the online materials.
Factor Score Determinacy
We examined the factor determinacy for the substantive factors—the correlation between the factor score estimates and the latent traits they represent (Beauducel, 2011; Guttman, 1955). The closer factor determinacy coefficients are to 1, the better factor scores represent the latent factors, and cut values based on whether scores are used for research (.80: Gorsuch, 2014) or individual assessment (.90: Grice, 2001) have been proposed. The factor score determinacies for the trifactor model for the complete data pattern were all high as follows: honest SDE .901; faked SDE .906; honest IM .958; and faked IM .956. Moreover, for all of the numerous missing data patterns factor determinacy was similarly high. These values meet even the strictest recommendations for individual assessment discussed by Grice.
Mixture Modeling
When we examine the BIDR observed score distribution plots for continuous responses in the faking condition, we see a clear bimodal distribution, suggesting a mixture of subpopulations might account better for the response patterns in that condition. However, this is not the case in the honest condition. A plausible explanation is that under the faking condition, not all participants follow the instruction and think of ideal responses; some simply respond in the normal fashion (honestly), which is less cognitively demanding. To test this hypothesis, we allowed two latent classes in the final equated trifactor model as described earlier. The first class is a class of “compliers” and the second class is a class of “noncompliers.” In the mixture model, the factor means, variances, and covariances are allowed to differ in the class on noncompliers, but only in the faking condition. All parameters related to the honest condition are the same across both classes because people in both classes complete the honest condition normally.
The trifactor mixture model has only 16 parameters more than the equated trifactor model and it fits decisively better. Both the AIC (114,852 vs. 114,587) and the BIC (116,134 vs. 115,938) were much lower for the two-class model. The mixture model also has very interpretable average response profiles and the expected correlations of near 1 between respective constructs even though these were freely estimated. The model also has good class separation, indicated by an entropy value of .88. With the mixture model, the bimodal distributions of the latent variables are clearly explained as an artifact of two subpopulations present—compliers and noncompliers, and as a result, there was a higher faking effect than in the standard trifactor model, represented by the standardized mean differences of 2.235 for SDE and 2.335 for IM between honest and faking conditions. This is because 30.6% of participants actually belong to the noncompliers class and their honest responses were dragging the overall effect of the faking instruction down in the single class model.
Factor Correlations for the Mixture Model
Substantive factors of the BIDR, SDE and IM, were allowed to correlate with one another within conditions and across conditions. For the complier class, the within-condition correlation for the substantive factors was .27 for honest instructions and .79 for faking instructions. The within-condition correlation for the noncomplier class was also .27 for the honest instructions and was .25 for the faking instructions. For the complier class, the across-condition correlations between corresponding substantive factors was .14 for SDE and .20 for IM, whereas in the noncomplier class, these values were 1.06 for SDE and 1.03 for IM. We note that for noncompliers these correlations are expected to be unity, and while they are estimated as slightly greater than unity the confidence intervals for both correlations include one.
General factors were set to be orthogonal to substantive and specific factors but were allowed to correlate with each other across conditions. The correlation between the general factors in the honest and faking instruction was .54 for the complier class and 1.02 for the noncomplier class. Once again, these are expected to be unity for noncompliers by design and while they are estimated greater than unity confidence intervals about these estimates include the expected value of 1. Item-specific factors were orthogonal to all other factors as well as being orthogonal to one another and are not interpreted.
Observed to Latent Scale Correlations
We expect that most readers will use the BIDR observed scores, so a natural question might be, “What is the relationship between observed BIDR subscale scores and latent representations of the same dimensions?” We calculated the correlation of the estimated factor scores in the best-fitting model, trifactor mixture model, with the observed scores, revealing the results presented in Table 7. The correlations between the estimated factor scores for SDE and IM and the respective 5-point Likert-type scores were all in excess of .95. This indicates that for practical purposes, the 5-point scoring system will produce ordering of people similar to that described by the SDE and IM substantive factors modeled in this article. On the contrary, Table 7 reveals that the dichotomously scored BIDR and the estimated factor scores are somewhat different. For the SDE scale, the correlation between latent and dichotomous scores was .53 and .74 respectively for the honest and faking conditions, whereas the correlation between latent and dichotomous scores for the IM scale was .80 and .91 respectively. The lower correlations for the SDE scale across conditions, relative to the IM scale, are likely due to the BIDR dichotomous scoring protocols. For the SDE scale, this protocol sees reverse-scored items appropriately recoded, and then coding one, two, three, or four as zero, and values of five recoded as one. The IM scale, on the contrary, sees one, two, and three recoded as zero, and four and five recoded as one. Such dichotomization of the IM scale is likely more representative of the threshold differentiating high- and low-stakes responses (with three bottom response options being almost exclusive markers of the low stakes); therefore, it captures more information about the latent traits and aligns more closely with the continuous scoring in the trifactor model. The fact that the dichotomous scoring less accurately reproduces the continuous score distributions is evident in the distributions presented in Figure 3.
Correlations Between Dichotomous Observed Scores, 5-Point Observed Scores, and Trifactor Mixture Model Factor Scores.
Note. Dichotomous refers to the observed scores scored using the BIDR dichotomous scoring protocol; 5-point refers to the observed scores scored using the BIDR Likert-type scoring protocol with five scale points. Fs = faking self-deceptive enhancement; fi = faking impression management; hs = honest self-deceptive enhancement; hi = honest impression management; h gen = general factor in honest condition; f gen = general factor in faking condition; BIDR = Balanced Inventory of Desirable Responding.
Discussion
Socially desirable responding is a ubiquitous threat to validity in psychological assessment, so it is important that we have techniques available to identify when such responding occurs. Among the most widely used approaches are validity scales, and perhaps the most routinely used validity scales is the BIDR. Early efforts to validate the BIDR’s purported two-factor structure have produced unsatisfactory fit. In fact, Leite and Beretvas (2005) commented with respect to the two-factor structure of the BIDR that “It seems that until the structure of responses to the MCSDS and the BIDR can be better clarified, researchers should be careful when attempting to correct scores of other scales based on SDB scores” (p. 152). In this study, we were able to tease apart the factors that were preventing adequate fit on the BIDR in earlier studies using a trifactor modeling approach and to establish partial measurement invariance for the BIDR across honest responding and instructed faking—a simulated high-stakes situation. We also showed that the factor scores, estimated with high determinacy, correlated near unity with observed variable counterparts. This means that the BIDR can be used to quantify the extent of faking in an assessment.
Methodological Contributions
In terms of methodological contributions, we demonstrated how Bauer et al.’s (2013) trifactor model can be extended from a multiple raters design to study change between experimental conditions for within-subjects designs. The trifactor model we fitted included substantive factors capturing the extent of SDE and IM in each experimental condition, method factors capturing response styles in each condition, and item-specific factors, and it enables identification of condition-related change on the substantive factors due to the response instructions. Identification of the trifactor model was permitted by having two instructional conditions—namely, the ability to identify specific item variance in repeated administration designs. The remaining (common) item variance was partitioned into the substantive effects with their mean shift across conditions and the effect of response style (which turned out to be mostly acquiescence).
We also demonstrated how mixture modeling can be applied to the trifactor model (or to any factor model suitable for the task in hand) to account for heterogeneity in the distribution of observed scores in the faking condition. In our experience, this is not a rare event when research participants do not fully comply with experimental instructions that are cognitively taxing. In such cases, it is possible to identify the latent class of noncompliers in fully confirmatory fashion, by constraining their model parameters in the faking condition to be equal to the honest condition. This approach allows estimating the faking effect more accurately, by disallowing the noncompliers to influence the result. In our samples, noncompliance was estimated at 30%, which would have a strong influence on both the model fit and the substantive results if not controlled.
While we operationalize item-specific factors in our models as correlated residuals, it is also possible to estimate these effects as latent variables that are orthogonal to each other and all other factors in the model, which achieves identical results. Whereas the correlated residuals approach is simpler syntactically, the latent variables approach might be preferred if covariates are expected to predict the item-specific factors, or if the item-specific factors themselves are to be used in prediction of other variables. In our online materials, we present the baseline trifactor model estimated with both correlated residuals and with specific factors ways, achieving identical fit.
Practical Implications of These Results
To inform use of the BIDR in applied settings, we report correlations of the factor scores from the best-fitting model, a trifactor mixture model, with their observed variable counterparts in each experimental condition. We show that the 5-point scoring scheme yields scores that correlate highly (above .95) with their trifactor model counterparts, whereas the dichotomous scoring departs from them substantially. These results give some confidence regarding the interpretation of the observed BIDR scores derived from the 5-point scoring schemes as representative of the same factors in the latent variable model. Because the latter were shown to be mostly measurement invariant across low- and high-stakes instructions, and sensitive to these instructions in term of the mean shift, the observed 5-point sum scores will possess similar properties and can be recommended for use in practice. The dichotomous BIDR scores, however, are not supported by the same evidence and need further investigation with respect to their construct validity.
On the broader question of using validity scales such as BIDR in practice several points are important to note. The first is that these results give some confidence regarding the use of the BIDR summated scores or the factor scores estimated and saved from respective latent variable models as a check for the extent of faking in applied assessment settings. For instance, at the individual level, observed scores that exceed the recommended cut-offs might prompt particular care to examine consistency between how individuals describe themselves in response to different assessment methods as well as self-other discrepancies on constructs assessed with multiple methods. Where noninvariant items are identified, it would be best to remove them from the sum score. At the group level, the score distributions and the means can serve as good indicators of the extent of SDE/IM in the population of test takers in the current assessment context.
Second, while we provide support for the use of the observed BIDR scores as measures of faking, this does not imply that they could or should be used to “correct” any substantive assessment scores. Many authors have warned against such “corrections” using manifest scores from validity scales, as the observed scores carry at least some “substance”—that is, variance due to stable personality attributes such as neuroticism and conscientiousness (Ones et al., 1996). Indeed, results of this study show that the 5-point summated BIDR scores in the faking condition correlate weakly but significantly (.17 for both IM and SDE) with the same scores in the honest condition. Therefore, using these scores to partial out the “faking” variance will lead to removing some variance related to stable personality characteristics.
Expected Generality of These Results
A broader question is whether this result will generalize to other validity scales and what conditions must be met for a scale to be valid in measuring faking. From a psychometric perspective, we must recognize that the BIDR scales aim to assess psychological constructs and they do so with reasonable measurement properties (e.g., unidimensionality and reliability). For the results reported here to hold for other validity scales, it is likely important that the scales also target psychological constructs with scales that have strong psychometric properties. The MCSDS for instance, which is well grounded theoretically, aims to measure a homogeneous construct, behaviors with low base rate probabilities and high social desirability (Lambert et al., 2016). The MCSDS items are most similar to the BIDR IM items and the construct captured by them in a high-stakes situation is likely similar to that of IM (see also Uziel, 2010). We anticipate the results of this study would likely hold for this scale. The case for validity scales that are not targeted to capture self-presentation behaviors is not so clear. For instance, the “Cannot say” scale of the Minnesota Multiphasic Personality Inventory, which is a count of omitted items, is likely to measure a response style and not a purposeful behavior, making it difficult or even impossible to generalize these results. Assuming items to measure purposeful self-presentation behaviors, for these results to generalize it is further important the observed variable rating scales directly reflect the modeled responses. We have seen here, for instance, that as the scoring procedure is coarsened, the correlations between the model-based factor scores and the observed variable counterparts depart substantially from unity.
Limitations and Future Directions
One potential limitation is that we combined a nonprobability student sample and a second nonprobability convenience sample, and there were differences in the samples on demographic background variables and on observed scale scores. Although nonprobability samples are common, there may be some concern due to the combined student and working adult samples. Panel respondents of working adults can be different to nonpanel respondents in unknown ways that attenuate relationships between variables and it is thought that this is sometimes due to being experienced experiment participants (Chandler & Shapiro, 2016). This seems not to have happened in the current study given that the largest experimental effects, at the observed level, were for the student sample that was not a panel sample. We expect that the ability to see the items once and answer honestly improved the ability to fake. However, fitting this model to a more homogeneous sample of working adults should be a future direction to check the generalizability of the conclusions about the degree to which mean levels change due to the faking instruction and to investigate the fit of the repeated-measures trifactor model.
Other limitations to this study include that we collected data across two samples for counterbalancing and it would have been desirable to randomly assign participants to the respond honest first versus faking first conditions. Nonetheless, aside from differences in mean scores on scales across the subsamples, the data were able to be modeled in ways that suggest the results are generalizable. Some might also feel that a professional sample from a single organization might be preferrable to a potentially heterogeneous community sample. Yet, heterogeneous samples are characteristic of many assessment settings such as high-stakes job selection.
In this study, our experimental condition factors appeared to be a mix of acquiescence and inattentive responding. This is even though we screened out those respondents who completed in times we deemed too fast to represent diligent responding. One future approach might be to try a direct instruction question that is not scored, such as “choose 5 for this question,” and eliminate any respondents who fail this check. This might offer a more concrete indicator of inattentive responding. Finally, our study used an instructed faking design and it is not certain that people will fake in precisely the same in an induced faking situation as they would in a real high-stakes assessment context. It is difficult to overcome this limitation, however, outside of the context of an instructed faking design such as that we adopt here.
Future directions could include generalizing the adapted trifactor approach to study multiple group models to, for instance, compare invariance across populations of interest. Another major direction is exploring external validities of the method factors and the substantive factors by adding covariates to the trifactor model. This line of investigation might also link specific factors to covariates to interpret their meaning (rather than viewing specific factors as technical devices to allow accurate estimation of other factors). It would also be worthwhile to conduct simulation studies to examine the performance of the within-subjects trifactor model, including examining the relative performance of the correlated residual and specific factor implementations. Researchers may also wish to examine the properties of factor scores for the within-subjects trifactor model, as has recently been reported for the original multiple informant trifactor formulation (Curran et al., 2021).
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
