Abstract
This research analyzes the effectiveness of the list experiment and crosswise model in measuring self-plagiarism and data manipulation. Both methods were implemented in a large-scale survey of academics on social norms and academic misconduct. As the results lend little confidence about the effectiveness of the methods, researchers are best advised to avoid them or, at best, to handle them with care.
Introduction
Eliciting accurate prevalence estimates of sensitive characteristics, such as drug usage, antisocial behavior, or misconduct in a survey environment is prone to misreporting (Krumpal 2013; Tourangeau et al. 2000; Tourangeau and Yan 2007). Survey methodologists thus seek to develop question designs that allow capturing sensitive behavior indirectly to circumvent issues related to social desirability pressures. Two popular methods are the list experiment (LE) (Droitcour et al. 1991; Miller 1984) and the crosswise model (CM) (Yu et al. 2008).
Even though both formats are frequently used to measure sensitive characteristics (Droitcour et al. 1991; Gilens et al. 1998; Hoffmann and Musch 2016; Höglinger and Jann 2018; Höglinger et al. 2016; Holbrook and Krosnick 2010; Jann et al. 2012; Johann and Thomas 2017; Korndörfer et al. 2014; Roberts and John 2014; Thomas et al. 2016; Tsuchiya et al. 2007; Wolter and Laier 2014), an increasing literature raises concerns about the effectiveness of these methods (Aronow et al. 2015; Coffman et al. 2017; Coutts and Jann 2011; Droitcour et al. 1991; Hoffmann et al. 2017; Höglinger and Diekmann 2017; Jerke et al. 2019; Kiewiet de Jonge 2015; Landsheer et al. 1999; Schnell and Thomas Forthcoming; Wolter 2012; Wolter and Laier 2014). The methods “impose a higher cognitive burden on respondents” (Jerke et al. 2019:320); are prone to produce false negative and positive results (Höglinger and Jann 2018; Wolter and Laier 2014); and their effectiveness may depend on the sample quality (Schnell and Thomas Forthcoming).
On the Logic of the List Experiment and Crosswise Model
The traditional LE is based on a split sample design, where the control group receives a short list of unobtrusive items; the treatment group gets the same list plus a sensitive item. Participants are asked to indicate how many of the items apply to them without disclosing which ones. The mean difference of the long and short list provides the prevalence estimate of the sensitive characteristic. Variations of the traditional LE exist (Aronow et al. 2015; Blair and Imai 2012; Glynn 2013; Li and Van den Noortgate 2019), but have only been applied infrequently.
The CM features an unobtrusive and a sensitive question. As opposed to providing separate answers to each question, respondents are asked to give a joint answer to both. Either the response to the questions is the same or it is different. The prevalence of the sensitive characteristic can be estimated, if the researcher knows the probabilities of the unobtrusive question and both questions have binary response codes (Jerke et al. 2019; Krumpal et al. 2015). Variations of the CM, including changes to the selection of the unobtrusive question, have been proposed but not yet tested (Diekmann 2012; Höglinger 2016).
If validation with a true value is impossible (Landsheer et al. 1999), a direct question (DQ) can be asked in a separate split sample to test whether the LE or CM produces better estimates than does the DQ. A better estimate according to the commonly applied “more-is-better” assumption is a higher prevalence estimate for socially undesirable behavior and, in reverse, a lower estimate for socially desirable characteristics (Umesh and Peterson 1991). Yet, recent research indicated that this assumption might be undermined by respondents deliberately lying or cheating (e.g., by disregarding the rules) (Höglinger and Jann 2018; Höglinger et al. 2016; Jerke et al. 2019; Walzenbach and Hinz 2019).
Method
The Zürich Survey of Academics (Rauhut et al. 2020) enquired about recent developments in academia in Austria, Germany, and Switzerland. The project conducted a census of academics in Austria and Switzerland and drew a 50%-probability sample of scholars in Germany. In total, 15,972 academics at 236 universities were interviewed (Austria: n = 2,832; Germany: n = 8,228; Switzerland: n = 4,912). The overall response rate was 11.33% (Austria = 10.08%; Germany = 10.44%; Switzerland = 14.44%).
The LE and CM were designed to measure socially undesirable academic misconducts: (1) Submitting the same results without indicating it (self-plagiarism) and (2) intentionally altering the data to confirm the research question (data manipulation) (Fanelli 2009). The exact question wording is presented in the Appendix A.
Both items are considered as sensitive: Roughly 95% indicated that they felt uncomfortable admitting to data manipulation (Austria: 94.88%; Germany: 95.10%; Switzerland: 93.74%); two-thirds said the same about self-plagiarism (Austria: 67.38%; Germany: 65.33%; Switzerland: 67.98%).
The wording and all items were pretested in expert discussions and cognitive interviews, indicating that commonly used unobtrusive questions with known probabilities, such as asking about someone’s birthday, raised mistrust among respondents (Jerke et al. 2019; Rauhut et al. 2020). To circumvent this, a context-related statement question was asked in a separate split sample and serves as an estimator for the unobtrusive question. This is a novel approach. To avoid the loss of privacy protection in the LE due to floor and ceiling effects, the unobtrusive items vary in their prevalence and some of the items correlate negatively (see Appendix B) (Glynn 2013; Jerke et al. 2019).
Both designs were integrated toward the end of the overall questionnaire to avoid early breakoffs. The computer-assisted web interviews were programmed to randomly assign respondents to treatment and control groups: Group 1 received both CM questions, Group 2 the LE’s long lists, Group 3 the LE’s short lists, Group 4 the DQs.
We calculated the prevalence estimate π^ along with the respective standard errors SEπ^ for both items and methods on the full sample. The underlying equations can be found in Appendix C. We excluded respondents indicating that they had never published, as academics without publication experience are unlikely to self-plagiarize. We also drop those saying that they have no experience with statistical data analysis as they are unable to manipulate quantitative data.
For robustness, we (1) re-ran the analysis on the full sample; (2) conducted cross-country analysis to uncover potential variation; and (3) re-estimated the results for self-plagiarism by respondents’ perceived level of sensitivity, comparing those feeling highly uncomfortable admitting to the misconduct with those feeling less uncomfortable. This last check was only performed for self-plagiarism, as almost all respondents rated data manipulation as highly sensitive, generating too little item variance. The rationale for this check is that the perceived item sensitivity may impact the effectiveness of the method when high-frequency behavior is concerned (Wolter and Laier 2014).
Results
Estimates of Self-plagiarism (SP) and Data Manipulation (DM).
The robustness checks confirm our results: The analyses reveal similar patterns for the full sample, across countries—with the exception of Austria, where the LE arguably works better following the more-is-better assumption—and when splitting the sample by respondents’ perception of sensitivity. 1
Discussion and Conclusion
The LE and CM are popular methods to help circumvent issues of misreporting and are believed to have the potential to better estimate sensitive characteristics. But there are concerns that the methods are mostly successful when low quality samples are used (Schnell and Thomas Forthcoming), that they are cognitively too challenging (Jerke et al. 2019), and that they produce substantive numbers of false positives and negatives (Höglinger and Jann 2018). Effective implementations may also be tied to the level of sensitivity (Thomas et al. 2016; Wolter and Laier 2014).
Our results show that both methods fail to generate higher prevalence estimates for self-plagiarism. While the LE does not work for data manipulation, the results indicate that the CM potentially performs well for data manipulation. However, we are wary that this positive result may be biased by the occurrence of false positives (Höglinger and Jann 2018; Höglinger et al. 2016).
Our results support a growing body of literature raising concerns about the effectiveness of the LE and CM. The occurrence of negative prevalence estimates in the LE conditions might be a relic of the design or related to different sample populations (Tsuchiya et al. 2007). However, as we employ a highly educated sample of academics, we may conclude that the effectiveness of LE is not an issue of respondents’ cognitive skills. The same argument applies to the CM (Jerke et al. 2019). We would carefully note that asking about academic misconduct in general may be too sensitive, as the associated sanctions can be severe. Any survey question format might fail to predict too sensitive characteristics. This article challenges Roberts and John (2014), who emphasize that future studies on scientific misconduct should consider using specialized questioning techniques, such as LE and CM. In line with prior research (e.g., Gelman 2014; Glynn 2013; Hinsley et al. 2019; Jerke et al. 2019; Schnell and Thomas Forthcoming), we conclude that researchers are best advised to avoid these methods or, at least, to handle them with care.
Footnotes
Appendix A: Question Wording
Appendix B: Correlations Unobtrusive Items LE
Correlations Unobtrusive Items LE Data Manipulation (V87).
| Item 2 | Item 3 | Item 4 | |
|---|---|---|---|
| Item 1 | 0.0347 | −0.0119 | 0.0886 |
| Item 2 | −0.1280 | −0.0834 | |
| Item 3 | 0.1646 |
Appendix C: Prevalence Estimates and Variances
Acknowledgments
We thank the anonymous reviewers as well as Alexander Ehlert, Isabel Raabe, and Justus Rathmann for their concise comments and constructive feedback on our work. Co-authors in alphabetical order. Study Design: Julia Jerke, David Johann, Heiko Rauhut, Kathrin Thomas, Antonia Velicu. Coding and Analysis: Julia Jerke, David Johann, Kathrin Thomas, Antonia Velicu. First draft: Julia Jerke, Heiko Rauhut, Kathrin Thomas, Antonia Velicu. Revisions: David Johann, Kathrin Thomas, Antonia Velicu. Final approval of the paper: Julia Jerke, David Johann, Heiko Rauhut, Kathrin Thomas, Antonia Velicu.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research is supported by the Swiss National Science Foundation (SNSF), Starting Grant “CONCISE” BSSGIO 155981 of Heiko Rauhut.
