Abstract
Complete and accurate survey data are key input for research, policy, and decision making in many disciplines. However, survey respondents do not always fully cooperate, such that they skip some items or overuse the “don’t know” answer option. Evasive answer bias reflects different information than overall survey response rates, leading to item missing data and causing substantial inaccuracies in survey results. Using data from the World Values Survey, this article identifies the magnitude of the problem, then relies on individual data and country-level cultural values to derive patterns of and reasons for this evasive answer bias. While skipping answers happens less often in collectivistic and low power distance cultures, the choice of the “don’t know” option is not significantly influenced by any cultural dimension. Across countries, the effect of cultural values is stronger for female than for male respondents. Accordingly, cross-cultural researchers are advised to use advanced imputation rather than deletion methods for handling missing data.
Keywords
Introduction
Sample surveys are an important basic research method and tool in the social sciences, fueling an industry with a multibillion-dollar annual turnover (Wright and Marsden, 2010). Increasingly, government agencies and economic enterprises require complete and accurate data to respond to their constituents’, customers’ and other stakeholders’ needs. They use data to learn about society, and provide their patrons with targeted information, advertisements, and means for self-expression, often in exchange for discounts or small gifts (Messner, 2004; Steiner and Maas, 2018). An important source of these data are the patrons themselves (Cazier et al., 2007), such as when they complete survey questionnaires (Bowling, 2005). To understand and access international markets, organizations increasingly conduct surveys across national borders too (e.g. Bellman et al., 2004; Brown et al., 2007).
Despite this spread, people may react negatively to their data being collected, profiling efforts, and unpredictable or unwanted uses of their information (Cazier et al., 2007). Growing exposure to marketing efforts and over-surveying have intensified such negative perceptions (Gummer, 2019; Presser and McCulloch, 2011) and diminished respondents’ willingness to cooperate, which in turn raises concerns about survey accuracy (Fuchs et al., 2013; Gummer, 2019; Wagner and Stoop, 2018). That is, improving survey response rates is critical, but response rates represent only a part of the issue (Durand et al., 1983).
In addition to nonresponse bias, evasive answer bias manifests itself in item nonresponse, that is, in skipping questions or overusing the “don’t know” option (Kamakura and Wedel, 2000). To be clear, evasive answer bias is different from nonresponse bias (Couper and De Leeuw, 2003). While the latter relates to survey response rates (aka, unit response rates), the former describes the phenomenon of nonresponse at the item level (aka, item missing data). This article is exclusively about evasive answer bias at the item level.
Furthermore, respondents’ answers may reflect their conscious or unconscious biases (Osborne and Blanchard, 2011), such as consistency bias, social desirability bias, or biases reflecting inflated views of one’s own performance (Harzing et al., 2012). Respondents also exhibit systematic tendencies in their responses, regardless of the content (Baumgartner and Steenkamp, 2001), including acquiescence, disacquiescence, middle, and extreme response styles (Harzing, 2006; Watkins and Cheung, 1995).
Unmotivated survey respondents “can cause substantial misestimation of results, biasing results towards the null hypothesis” (Osborne and Blanchard, 2011: 1), leading to data quality problems. Evasive answer bias hides actual responses that could be meaningful for analysis (Little and Rubin, 2002). Moreover, researchers often resolve item missing data with list- or pairwise deletion (Myers, 2011). If the bias is random, the cleaned data set is representative of the sample and the statistical inferences remain valid. But if the bias can be explained systematically by specific factors, the cleaned data is no longer representative of the sample and the resulting inferences will be problematic (Jensen et al., 2010).
Similar to differences in consumer behavior linked to national cultures (Malhotra et al., 2005), international surveys often reveal response style biases that vary across countries, cultures, or languages (Harzing, 2006; Harzing et al., 2012; Messner, 2017). Accordingly, multigroup confirmatory factor analysis has become a method of choice for establishing measurement equivalence in cross-cultural research (Welzel et al., 2019). Yet, “there is a lack of research that investigates how cultural differences across countries affect consumers’ willingness to share personal information” (Schumacher and Maas, 2019, CC41), and little is known about how such differences affect unit or item response rates (Lyness and Kropf, 2007). For example, differences in answer bias between countries could look like differences in content, leading researchers to derive inaccurate insights from the incorrectly assumed differences (Van Rosmalen et al., 2010). Nor does extant research detail how the influence of individual demographic characteristics and psychological traits shifts across cultures (Jans, et al., 2018; Roxas and Stoneback, 2004). Attempts to link cultural and survey nonresponse theory are relatively new in extant research (Jans, et al., 2018; Johnson et al., 2010), and so Zhou et al. (2017) suggest researching whether respondents’ cultural and socialization background influence their answer behavior.
In response, the current study seeks to explore the reasons for evasive answer bias, examine trends across countries and cultures, and reveal how individual-level drivers might be moderated by cultural values. Observational data from the World Values Survey (WVS; Inglehart et al., 2014) combined with the GLOBE study’s cultural framework (House et al., 2004) help demonstrate how personal characteristics and cultural predispositions influence, in sometimes conflicting ways, two measures of evasive answer bias: skipping questions and choosing the “don’t know” answer. The article’s final section then details implications. It advises cross-cultural researchers to address missing data in international surveys through advanced imputation rather than deletion methods.
Resource exchange and leverage saliency theory
Resource exchange theory demonstrates how people’s willingness to disclose information relates to the benefits they receive (Brinberg and Wood, 1983). When completing surveys, for each item, respondents typically undergo a four-step process: (1) comprehend the question, (2) recall information from memory, (3) map retrieved information to the question asked, and (4) communicate a response (Bowling, 2005; Tourangeau, 1984). But respondents also may elect to skip certain questions (Warner, 1965) or choose an “I don’t know” answer option (Dobronte, 2014). In particular, they tend to avoid questions that require a longer view into the future, cover remote or controversial topics, are too revealing, come along with complex instructions (Converse, 1976), or require absolute answers (Shoham, 1998).
Leverage saliency theory proposes that respondents’ characteristics determine their likelihood of cooperating in a survey (Groves et al., 2000; Gummer, 2019), and that different people place a different level of importance to survey requests (Seifert, 2008). Previous research indicates that age, education, gender, and income determine the number of non-substantive responses people provide (e.g. Craig and McCann, 1978; Feick, 1989; Francis and Busch, 1975), though research into item nonresponse phenomena has been hindered by a lack of sufficient data (Ferber, 1966), leaving limited insights into random and evasive responders (Furnham et al., 2015). Still, the possibilities of modern technologies and associated risks of data fraud have propelled information disclosure topics into the limelight in current academic research and management practice (Steiner and Maas, 2018).
Hypotheses
Overview of hypotheses.
This table lists the hypotheses. The results are recorded separately for the unanswered questions (NOANSW) and questions marked with “don’t know” (DNTKNW).

Multilevel research model and variables. This is the multilevel research model. The hypotheses H1 to H6 examine individual-level effects (level 1); H7 and H9 model group-level direct effects (level 2); H8 and H10 test cross-level interactions between level 2 and 1. Table 1 provides further information about the hypotheses; Table 2 records the variables and their abbreviations.
Individual-level effects
Gender identity theory argues that gender phenomena are multifactorial, consisting of biological sex, gender role attitudes, and psychological traits (McCabe et al., 2006; Spence, 1993; Spence and Helmreich, 1978). In general, women show more expressive, interpersonally oriented, and caring psychological traits than men. They are more emotional and adopt different perspectives. Men instead tend to be more ambitious, autonomous, aggressive, confident, and proactive (e.g. Kidder and Parks, 2001). Bossuyt and Van Kenhove (2018) find that the gender difference in assertiveness remains robust even after controlling for social desirability bias. Even during the COVID-19 pandemic, women suffered greater emotional and life distress than men (Ding et al., 2021).
According to gender socialization theory, social factors and influences also cause distinct moral developments in women and men, leading to differences in value orientation (Glass and Cook, 2018). In a similar situation, women and men respond differently (Betz et al., 1989). This theory also maintains that men prefer success and competition to following rules, whereas women prefer to perform tasks properly while maintaining amicable relationships (Roxas and Stoneback, 2004). Female respondents in turn may underreport undesirable answers (Collins, 2000; Randall and Fernandes, 1991; Schoderbek and Deshpande, 1996).
Ethics studies generally suggest that women are more ethical than men (Collins, 2000; Grimshaw, 1991; McCabe, Ingram, & Dato-on, 2006), and men are more likely to break rules (Roxas and Stoneback, 2004). Deciding how to respond when one does not quite know the answer is akin to an ethical dilemma. Differences between female and male respondents could thus be an enduring effect of their differential socialization (Rapoport, 1982). Yet, most surveys measure only ethical intentions, which are highly susceptible to social desirability bias (Dalton and Ortegren, 2011); studies testing or measuring actual ethical behavior are extremely rare (Bossuyt and Van Kenhove, 2018).
In the context of social and political science, an additional factor comes into play. Women and men differ with respect to the issues that interest them (Campbell and Winters, 2008; Coffé, 2013). In political science research, numerous studies find that women exhibit lower levels of political knowledge than men (Bennett and Bennett, 1989; Dolan, 2011). They are sometimes attributed to gender differences in social roles and responsibilities, and access to resources associated with political engagement (Verba et al., 1997). While these differences are not always large, they are consistent and persistent (Dolan, 2011). And depending on the content domain of a survey, appropriate levels of political knowledge can be important for effective participation. Taken together:
H1: Female (male) respondents evade questions more (less). Other demographic characteristics also might anticipate differences in survey response behavior (Scholz and Zuell, 2012). For example, willingness to participate in surveys changes with age (Gummer, 2019; Gergen and Back, 1965). For certain domains, especially politics, younger respondents tend to be less experienced than more mature respondents, because they have spent less active time participating in these domains. Mature respondents tend to be more deliberate and careful in their answers, and thus more likely to admit a lack of knowledge, so they may refuse to answer questions more so than younger respondents. This effect continues with old-age, albeit for different reasons. For example, a health survey reports lower response rates among respondents older than 65 years than among younger respondents (Brazier, et al., 1992; Jenkinson et al., 1993). Similarly, in epidemiologic questionnaires, Eaker et al. (1998) identify old age as the most important predictor of partial nonresponse, such that the mean percentage of missing answers is 5% higher in the 70–79 years age group than in the 30–49 years cohort. Consumer settings concur, such that Ferber (1966) cites a below-average item nonresponse rate among older people. Accordingly,
H2: Older (younger) respondents evade questions more (less). Comparatively low knowledge and experience also may encourage a self-perception of personal views as unusual or socially unacceptable. Especially when questions ask about public opinion, “lower-educated groups in the society are at some disadvantage” (Converse, 1976: 515). If less experienced people feel intimidated or threatened by survey questions (Krosnick and Presser, 2018), they likely exhibit nonresponse behavior to certain questions (Catania et al., 1986; Leigh and Martin, 1987). Beyond a lack of access to specific information, respondents with lower income levels may sense a status at a peripheral level of society, excluded and alienated, which they express by refusing to be interviewed or to answer all the questions (Francis and Busch, 1975). In contrast, if respondents have relevant considerations available in memory to answer a question, they need to construct an overall answer that integrates those considerations. This cognitively demanding process may exceed some respondents’ motivation or ability (Krosnick et al., 2002), such that they seek ways to avoid this work. Such avoidance may be more common among people with lower income and education levels. While education is one of the main determinants of income distribution across countries (De Gregorio and Lee, 2002), income as the consumption component of education (del Mar Salinas-Jiménez et al., 2011) is often found to be a stronger predictor than education in research models (e.g. Sabanayagam and Shankar, 2012). In turn:
H3: Respondents with higher (lower) income evade questions less (more). Careless responses, signaling insufficient effort (Curran, 2016), also can lead to psychometric problems and affect correlational outcomes, especially between regular and reverse-scored items (Kam and Chan, 2018; Kam and Meyer, 2015). Yet the need to check for careless responses often is overlooked by researchers (Kam and Chan, 2018). The lack of screening in turn prevent insights into its prevalence or impact (Maniaci and Rogge, 2014). Furthermore, conscientiousness as a personal trait has been linked to valid, non-random responding (Furnham et al., 1998; Furnham et al., 2015; Hülsheger and Maier, 2010). Existing studies mainly address insufficient effort responding (or satisficing) and study attrition (Ward et al., 2017), without considering the potential influence of the evasive answer bias. Therefore, the next hypothesis suggests a route to investigate the effect of conscientiousness on people’s tendency to evade survey questions:
H4: Respondents who are generally thorough (careless) evade questions less (more). As organizations solicit more customer information to gather data, enhance their offerings, and better target their marketing efforts, customers are becoming increasingly protective of the data and information they are willing to disclose (Zimmer et al., 2010). Privacy concerns inhibit information sharing (Ogan et al., 2017); trust can overcome those risk perceptions and encourage increasing information disclosure (Cazier et al., 2007; Dinev and Hart, 2006; Levin and Cross, 2004). Trust refers to a person’s willingness to be vulnerable to someone else (Cazier et al., 2007), including sharing personal information or opinions, which places the person doing the sharing in a potentially vulnerable situation. Because trust is the “variable most universally accepted as a basis for any human interaction or exchange” (Gundlach and Murphy, 1993: 41), it can support information disclosure (Steiner and Maas, 2018) and reduce the likelihood of opportunism (Bradach and Eccles, 1989), thereby increasing the probability that a person will enter a vulnerable situation through voluntary information sharing. In a survey setting, the interviewer has a prominent role, with considerable effects on survey responses (Loosveldt and Beullens, 2014), according to the trust that respondents develop in this interviewer. In this sense, people who generally trust others may exhibit less evasive answer behavior:
H5: Respondents who generally trust (do not trust) other people evade questions less (more). Consumer trust also might inhere to the organization collecting the data. In online settings, the organization even may become the primary recipient of consumer trust (Chow and Holden, 1997), because of the lack of a human representative. Regardless of how the questionnaire is administered, in the absence of sufficient time, cognitive capacity, or willingness to evaluate risk consciously, trust in the organization becomes a heuristic that decreases perceived risks (Vischers and Siegrist, 2008) and increases respondents’ motivation to disclose information (Kehr et al., 2015). Therefore:
H6: Respondents who generally trust (do not trust) major companies evade questions less (more).
Group-level effects
Jennings and Farah (1980) reported strong cross-national differences in evasive answer bias, which they could explain only partly with individual-level variables. Almost two decades later, De Leeuw and De Heer (2002) identified relatively large differences in refusal rates to household surveys across countries. Another 20 years later, Purdam et al. (2020) noticed considerable variation in the “don’t know” response across European countries. Yet, little research has sought to establish how individual-level characteristics manifest across cultures (Roxas and Stoneback, 2004) or how cultural values as group-level effects influence the evasive answer bias (Schumacher and Maas, 2019; Zhou, et al., 2017).
Leung et al. (2005) define national culture as the values, beliefs, norms, and behavioral patterns of a national group. A country’s culture shapes its people’s perceptions, dispositions, and behaviors (Triandis, 1989). Acquiring culture is a “slow process of growing into society” (Hofstede, 1991: 4), also referred to as socialization (Furth, 1990; Hofstede, 2001), by which people “learn various patterns of interaction that are based on the norms, rules, and values of their culture” (Gudykunst, et al., 1996: 510). A “complex configuration of values” (Woodside et al., 2011: 785) refers to standards that guide attitudes, actions, and the presentation of the self to others (De Mooij, 2017).
In turn, perceptions of ethical dilemmas, choices of alternative actions, and evaluations of their consequences influence decision making (Hunt and Vitell, 1986). Bartels (1967) defines ethics as “a basis for judgment in personal interaction,” pertaining “to the fulfilment or violation of expectations” (p. 21). Cultural values are critical in shaping ethical perceptions (Ho, 2010; Martin et al., 2009; Roxas and Stoneback, 2004), and a substantial body of research compares the ethical sensitivities of people from different countries (Collins, 2000), such that the current “call for globalization of ethics indicates there are nontrivial differences” (Roxas and Stoneback, 2004: 151). If a person perceives a set of alternatives, it evokes both deontological (right or wrong) and teleological (related to the goal) evaluations. In the deontological evaluation, the decision maker evaluates the rightness or wrongness of the behavior implied by each alternative, comparing it with a set of predetermined behavioral norms and values that reflect beliefs about honesty, cheating, and treating others fairly, as well as confidentiality, anonymity, and deceptiveness. In a teleological evaluation, the decision maker examines the consequences of the action with respect to its consequences for stakeholders, their probability, the desirability of the consequences, and the importance of each stakeholder (Hunt and Vitell, 1986). When answering survey questions, the most salient stakeholders are the interviewer and the organization conducting the survey. Because interviews involve social interactions with another person, respondents likely take social norms into account when responding (Bowling, 2005). Zey-Ferrell and Ferrell (1982) indicate that the organizational distance and relative authority of the stakeholder influences people’s beliefs and behaviors in ethical decision-making situations.
In-group collectivism is “the degree to which individuals express pride, loyalty, and cohesiveness in their organizations or families” (House and Javidan, 2004: 12). Dealing with the relationship between self and group, it is “perhaps the most important dimension of cultural difference in social behavior across diverse cultures of the world” (Triandis, 1989: 60). Its central element is the distinction between independent and interdependent self-construal (Aaker and Lee, 2001; Hofstede, 2001), that is, the image of self as separate from others (Singelis and Sharkey, 1995). As a cultural value, collectivism emphasizes ties between people. Collectivists tend to be interdependent, recognize their duties and obligations to others (Triandis, 1995), interact in a cooperative mode (Doney et al., 1998), and generally conform (Bond and Smith, 1996). Because they value relational harmony, they likely adjust to fit the environment in which they find themselves (Morling et al., 2002). If they are being interviewed, they adjust accordingly and try to answer every question. Because of their strong emphasis on avoiding conflict, collectivists likely expect to deploy a conflict-reducing technique (Forbes et al., 2011), such as skipping fewer questions, because they know answering the questions is expected of them in an interview situation, and failing to meet these expectations could cause friction. This effect may be especially powerful if the survey interview is conducted in the interviewee’s home, such that the interviewer takes the role of a guest or in-group member of society (Brewer and Kramer, 1985; Robbins and Krueger, 2005; Sumner, 1906). Conversely, if the survey is conducted in an out-group setting (e.g. shopping mall) or anonymously through electronic tools, antagonistic behaviors such as consciously skipping questions may seem more tolerated (Triandis, 1995). Collectivists are then likely to avoid answering socially sensitive topics. In a face-to-face interview, however, the immediate pressure caused by the presence of the interviewer is likely stronger than the desire to avoid a more remote and abstract social conflict.
On the other hand, people from individualistic societies prioritize themselves and their own well-being, so if they do not know the answer to a survey question, or if answering the question requires too much effort, they may just evade the question altogether. These considerations prompt the following hypothesis:
H7: Respondents from countries with predominantly collectivistic (individualistic) values evade questions less (more). Many previous studies have suggested that ethical differences between women and men are consistent worldwide, yet social values can influence ethical values (Chen et al., 2016; Turiel, 1994). In a certain social situation, value differences between women and men cause them to adhere to their cultural values in different ways (Taras et al., 2010). Gender traits also may be stimulated by cultural values with similar qualities (Chen et al., 2016). Moreover, women’s ethical judgments often are not fixed but vary with the context, unlike the more abstract, rule-based ethical judgments that men tend to exhibit, by abstracting the ethical problem away from an interpersonal situation to find an objective way to choose a course of action (Lee et al., 2000). Beekun et al. (2010) show that women’s choices of actions in response to ethical dilemmas are significantly affected by individualism, whereas men’s ethical decisions are more universal and unrelated to culture. In summary, female ethics appear influenced more by cultural values than do male ethics:
H8: The effect of female gender on evasive answer bias is stronger in countries with predominantly individualistic values. Power distance is the degree to which people “expect and agree that power should be stratified and concentrated at higher levels” (House and Javidan, 2004: 12). Some societies play down these differences; others allow the differences to surface and increase (Armstrong, 1996). Power distance beliefs appear related to decreased information transparency and disclosure but greater conformity and agreeableness. First, Turilli and Floridi (2009) define transparency in terms of information visibility, that is, “the choice of which information is to be made accessible” (p. 105), and Vaccaro and Madsen (2006) add the “degree of completeness of information” (p. 146) to this definition. Jain and Jain (2018) instead associate transparency with the process of disclosing information that can enable others to make judgments. Because transparency encourages openness, it increases privacy concerns (Ball, 2009). A preference for such information transparency and disclosure is negatively impacted by high power distance values (Hofstede, 2001; Hofstede et al., 2010; Jain and Jain, 2018), considering that “a hallmark of a high power distance cultures is … the consequent control of information” (Jain and Jain, 2018: 137). People may fear that survey responses, even if officially anonymous, could be monitored by government agencies that draw inferences from their responses. Therefore, in environments with less freedom, nonresponse may be more likely (Jensen et al., 2010). Conversely, low power distance cultures prefer flat hierarchies and endorse open access to information (Jain and Jain, 2018), implying that survey respondents from low power distance cultures likely answer more questions. Second, Taras et al. (2010) find in a meta-analysis that power distance is positively related to conformity at the individual level and to agreeableness at the country level. People with high power distance beliefs tend to enter into role-constrained interactions with authorities; they are more strongly regimented by the relative position of a superior (Lee et al., 2000). In a dyadic relation, a more powerful other can restrict the available choices and make a person conform to role expectations (Kahn et al., 1964; Zey-Ferrell and Ferrell, 1982). Survey respondents in high power distance cultures may be afraid or unwilling to express disagreement (Hofstede, 2001; Khatri, 2009), and their nonresponse offers a shield from possible reprisals (Jensen et al., 2010); they would rather evade a question than say something that might be contrary to perceived expectations. Therefore,
H9: Respondents from countries with predominantly high (low) power distance values evade questions more (less). Finally, women try to change rules to preserve relationships (Gilligan, 1982), but changing rules is more difficult in a high power distance culture. With respect to whistleblowing in audit firms, Taylor and Curtis (2013) find that men are relatively less sensitive to variations in power distance than women. The last hypothesis therefore states:
H10: The effect of female gender on evasive answer bias is stronger (weaker) in countries with predominantly high (low) power distance values.
Research context
Data set
This study uses the aggregated data set of the World Values Survey (WVS; Inglehart et al., 2014) with 348,430 respondents from 100 countries, constructed from various surveys and waves between 1981 and 2014. The WVS is an international academic project studying human values, self-descriptions, and attitudes across the globe. The WVS data are used widely by cross-cultural psychologists, anthropologists, and social scientists (Minkov, 2012). These data are collected in face-to-face interviews at respondents’ places of residence by professional organizations. The samples are representative of people residing in private households in each country. Questionnaires are translated into any languages that serve as a first language for more than 15% of the population (WVS, n.d.). However, overall response rates are not consistently reported across countries (Inglehart, 2000).
Operationalization of culture
To classify culture, this study turns to the in-group collectivism and power distance dimensions from the GLOBE framework (House et al., 2004). Introduced in 2004, the GLOBE framework is a useful cross-cultural framework (Smith, 2006). For both dimensions, this study considers societal to-be values, rather than as-is practices, because people are mainly driven by cultural values, whereas organizations reflect the influences of practices (Hofstede et al., 1990; Hofstede and Peterson, 2000). Though not without criticism (Hofstede, 2006), GLOBE shapes contemporary international research on cultural value differences (Beugelsdijk et al., 2017; Stahl and Tung, 2015). Notably, the GLOBE power distance measure refers to control of others, being conceptually different from Hofstede’s version of power distance, which implies acceptance and expectations of exhibitions of power and authority (De Mooij, 2017; Hofstede, 2006).
Research model and variables
Study variables.
This table records the provenance of the variables, listed in alphabetical order.
In these formulae, CNTNOANSW counts the number of questions not answered by the respondent (coded −2 in WVS). Then ALLQUE = 1367 denotes the total number of possible WVS questions; CNTMISSNG is the number of system missing questions for reasons unknown or not documented by WVW (coded −5); CNTNOTAPP and CNTDNTKNW count the number of questions that the respondent indicates are “not applicable” (coded −3) or “don’t know” (coded −1), respectively; and CNTNOTASK indicates the number of questions not asked in the survey (coded −4). For example, for a 64-year-old female respondent from Poland with unique WVS-ID 6160240223,
Descriptive statistics and observations
Among all WVS respondents, 56.966% skip at least one question (198,487 respondents with NOANSW > 0). The item nonresponse average is 1.319%, with standard deviation of 3.244. Then 70.585% indicate they cannot answer at least one question (246,012 respondents with DNTKNW > 0). For an average of 2.812% of items (standard deviation 5.043), respondents choose the “don’t know” option. Figure 2 reveals the frequency distribution of both types of evasive answer bias across all countries; Table 3 specifies the mean percentages at country level. Overall, the magnitude of evasive answer bias in the WVS appears substantial and worth investigating. Frequency of evasive answer bias in WVS (by respondent). These histograms show how many respondents (y-axis Frequency) have evaded what percentage of questions (x-axis Evasive answer bias). The figures are given as a percentage of all WVS respondents, across countries. Sample sizes, country-level means, and GLOBE cultural values. This table lists the number of respondents per country, the average percentage of questions skipped (NOANSW) or marked with “don’t know” (DNTKNW). It further gives the GLOBE values for collectivism (GCO2SV) and power distance (GPOWSV).
Ferber (1966) shows that the pattern of item nonresponses to different questions in surveys is very similar for various types of questions. Even though, the interpretation of the “don’t know” answer (and to some extent also the skipping of questions) is related to the kind of question asked. It can both be an evasive reaction or a legitimate response, that is, a declaration of a lack of knowledge (Presser, 1981). Figure 3 examines the frequency of both biases for a subset of 130 questions of the WVS that have been answered by more than 200,000 respondents. On average and across countries, a question has been skipped by 1.114% respondents (standard deviation 0.014), and 2.759% respondents (standard deviation 0.0321) have answered a question with don’t know. In Figure 3 and for both biases respectively, only 5.384% and 6.153% of questions are in the long right tail (defined as exceeding the average plus two standard deviations). These are sensitive questions (Krumpal, 2013) about the willingness to fight for one’s country (WVS item code E012), political positioning (E033, E115, E179WVS), and confidence in unions (E069_05), movements (E069_14, E069_15), and institutions (E069_20). The research model and the individual-level hypotheses H2 and H3 include demographic factors and individual information to help control for these effects. Frequency of evasive answer bias in WVS (by item). These histograms show – in percentage – how often items (y-axis Frequency) have been evaded by how many respondents (x-axis Evasive answer bias). The figures are given as a percentage of 130 WVS items answered by more than 200,000 respondents.
Figure 4 examines the existence of a potential questionnaire fatigue effect over time for all respondents across all countries. While the cumulative percentage of NOANSW nearly follows the trendline, DNTKNW shows a slight time-bound effect in that initially less questions are answered with “don’t know.” Taken together and following Ferber (1966), a view across all questions is justified to examine a respondent’s evasive bias. Fatigue effect in answering questions. This plot shows the influence of fatigue on evasive answer bias, separately for questions answered with “don’t know” (DNTKNW, black line) and questions not answered (NOANSW, grey line). The dashed lines are the trendlines through the origin.
Analysis and results
To simultaneously test individual-, cultural group-level effects, and cross-level interactions on the number of unanswered questions (NOANSW) and questions marked with don’t know (DNTKNW), two-level linear models using full maximum likelihood were separately calculated in HLM 7.03 software. All predictors were group mean centered at level 1 and grand mean centered at level 2; for the spotlight analysis (Spiller et al., 2013), all predictors were uncentered to ease interpretation (Baguley, 2009; Enders, 2013). The HLM notation by Raudenbush and Bryk (2002) is used; all variables are abbreviated according to Table 2 to avoid ambiguity.
Models and predictor variables.
Not all WVS countries are covered by the GLOBE study, and not all level-1 variables are available in all WVS countries and waves. This table provides a breakdown of the available level-1 and level-2 units in each research model used in this study, together with the included explanatory variables. Variable abbreviations are available in Table 2.
Dependent variable NOANSW
Unconditional model
The examination of the homogeneity of the means of NOANSW across all 100 countries and 348,530 respondents (Model A) produces the sample sizes and means listed in Table 3. Levene’s test indicates unequal variances, F(99; 348,430) = 526.263, p < .001, and the Kruskall-Wallis test reveals statistically significant differences in NOANSW among countries, H(99) = 65,834.566, p < .001. Removing countries not covered by the GLOBE study results in an unconditional model (Model B) with 51 countries and 233,525 respondents:
Level 1: NOANSW ij = β 0j + r ij .
Level 2: β 0j = γ 00 + u 0j .
The average effect is γ 00 = 1.345, SE = 0.145, p < .001. There is a relatively small variation between countries with 8.6% of the variation in item completion occurring between countries and about 91.4% within countries (intraclass correlation coefficient, ICC = 0.086). Yet, this variation is statistically significant (u 0 = 1.100, p < .001), and so, it is reasonable to proceed with the multilevel linear model.
Individual-level effects
HLM context, Model C
This is the HLM context for Model C (see Table 5), in which GENDER serves as a level-1 predictor. The effect size measure f
2
relates to variance explained for the overall model, and is computed as
Level 1: NOANSWij = β0j + β1j*(GENDERij) + rij.
Level 2: β0j = γ00 + u0j; β1j = γ10 + u1j.
On average and across countries, GENDER is positively and statistically significantly related to NOANSW, with an average effect of γ 10 = 0.193, SE = 0.040, p < .001, in support of H1. Female respondents leave more questions unanswered than male respondents, and the variation between countries is statistically significant, u 0 = 1.030, p < .001.
The variables AGE, entered as a level-1 predictor (Model E), reveals a positive and statistically significant link to NOANSW, with an effect of 0.010, SE = 0.003, and p = .003. The effect of GENDER is 0.196, SE = 0.041, p < .001. This outcome aligns with H2, such that older respondents leave slightly more questions unanswered than younger respondents. The variation between countries is u 0 = 0.995, p < .001.
When INCOME is entered as an additional level-1 predictor (Model F), GENDER and AGE continue to be positively and statistically significantly related to NOANSW (GENDER 0.148, SE = 0.039, p < .001; AGE 0.007, SE = 0.002, p = .011). Yet INCOME is negatively and statistically significantly related to NOANSW (−0.049, SE = 0.013, p < .001). In support of H3, respondents who earn higher income levels leave fewer questions unanswered.
When GENDER is supplemented with THORGH as an additional level-1 predictor (Model D), the effect of GENDER falls to 0.138, SE = 0.059, p = .038, and the effect of THORGH is −0.035, SE = 0.015, p = .045. This result confirms H4, because people who are generally thorough leave fewer questions unanswered.
Regarding the effect of trust on questionnaire completion, in Model G, the binary variable TRUST1 is not statistically significantly related to NOANSW, with an effect of 0.014, SE = 0.059, and p = .803. Moreover, the variable TRUST2 (Model H) reveals an effect that is not significant (0.012, SE = 0.021, p = .570). Similarly, the variable TRUST3 (Model J) exhibits no significant effect (0.024, SE = 0.028, p = .391). These findings conflict with H5 and H6 and indicate instead that the level of trust respondents have in others and in organizations does not significantly influence the percentage of questions they answer.
Group-level effects
HLM context, Model K
Level 1: NOANSW ij = β 0j + β 1j *(GENDER ij ) + r ij , and
Level 2: β 0j = γ 00 + γ 01 *(GCO2SV j ) + u 0j ; β 1j = γ 10 + γ 11 *(GCO2SV j ) + u 1j .
Thus, NOANSW depends on the level of GCO2SV (γ
01
= −0.608, SE = 0.287, p = .039), see Figure 5. In support of H7, people from collectivistic countries leave fewer questions blank than people from individualistic countries. Between male and female respondents, the average effect for NOANSW is represented as an increase to γ
10
= 0.192, SE = 0.037, p < .001. The cross-level interaction of GCO2SV on GENDER/NOANSW is γ
11
= −0.280, SE = 0.080, p = .001. As Aguinis et al. (2013) recommend, Figure 6 presents a depiction of the cross-level interaction at low (GCO2SV = 5.15), medium (5.68), and high (6.25) collectivism values with a spotlight analysis (all variables uncentered). According to Gelfand et al. (2004), these values reflect the midmost values of three societal in-group collectivism value bands in the GLOBE study. The gender effect is weaker for respondents from high collectivism countries, in support of H8. Group-level effects on unanswered questions. These scatterplots illustrate the cross-level interactions of the level-2 predictors collectivism (GCO2SV) and power distance (GPOWSV) for unanswered questions (NOANSW). Every dot represents a country. The trendlines are calculated without outliers (t-test < 0.01; Kuwait and Morocco; shown as hollow dots): NOANSW = −0.370 GCO2SV + 3.134, R2 = 0.102 and NOANSW = −0.423 GPOWSV + 3.413, R2 = 0.043.
When adding AGE and INCOME into the model (Model L), the cross-level interaction of GCO2SV on AGE/NOANSW and INCOME/NOANSW becomes clearly non-significant. None of the cross-level interactions of any trust-related variables (TRUST1, TRUST2, and TRUST3) on either the intercept or slope of the GENDER/NOANSW relationship is statistically significant either (Models M, N, and O).
Next, GPOWSV can be added as an alternative level-2 predictor (Model P):
Level 1: NOANSWij = β0j + β1j*(GENDERij) + rij.
Level 2: β0j = γ00 + γ01*(GPOWSVj) + γ02*(GCO2SVj) + u0j; β1j = γ10 + γ11*(GPOWSVj) + γ12*(GCO2SVj) + u1j.
HLM context, Model Q.
This is the HLM context for Model Q (see Table 5). Collectivism (GCO2SV) serves as a level-2 predictor for both intercept and slopes; power distance (GPOWSV) serves only as a level-2 predictor for the slopes. For effect size (f 2 ) interpretation refer the legend to Table 5; variable abbreviations are available in Table 2.
Level 1: NOANSWij = β0j + β1j*(GENDERij) + rij.
Level 2: β0j = γ00 + γ01*(GCO2SVj) + u0j; β1j = γ10 + γ11*(GPOWSVj) + γ12*(GCO2SVj) + u1j.
In this case, GCO2SV has a statistically significant negative effect on NOANSW (γ
01
= −0.608, SE = 0.287, p = .039). Between male and female respondents, the average effect of NOANSW is represented as an increase to γ
10
= 0.194, SE = 0.036, p < .001. The cross-level interaction of GPOWSV on GENDER/NOANSW is positive, γ
11
= 0.145, SE = 0.073, p = .055, and that of GCO2SV is negative, γ
12
= −0.210, SE = 0.096, p = .033. Figure 7 depicts the cross-level interaction through a spotlight analysis at low (GPOWSV = 5.15), medium (5.68), and high (6.25) power distance values (always medium GCO2SV = 5.68; predictors entered uncentered). In line with Carl et al. (2004) bands, these GPOWSV values reflect the extreme ends and midpoint of the societal power distance. The results confirm H9; an increase in power distance exerts a negative effect on questionnaire completion. This effect is stronger for female than for male respondents, so H8 also receives support. Cross-level interactions, Model K. For Model K (see Table 5) and unanswered questions (NOANSW), this diagram depicts the cross-level interaction through a spotlight analysis at low (GCO2SV = 5.15), medium (5.68), and high (6.25) collectivism values. All variables are uncentered. Cross-level interactions, Model Q. For Model Q (see Table 5) and unanswered questions, this diagram depicts the cross-level interaction through a spotlight analysis at low (GPOWSV = 5.15), medium (5.68), and high (6.25) power distance values. Collectivism is at medium level (GCO2SV = 5.68). All variables are uncentered.

Dependent variable DNTKNW
Unconditional model
Table 3 details the country-level means for DNTKNW. Similar to the results for NOANSW, Levene’s test indicates unequal variances for DNTKNW, F(99; 348,430) = 692.700, p < .001, and the Kruskall-Wallis test confirms statistically significant differences between the countries, H(99) = 68,678.910, p < .001. The ICC in Model B is 0.109, and variation between countries is statistically significant, u 0 = 2.822, p < .001. The average effect is γ 00 = 2.263, SE = 0.233, p < .001.
Individual-level effects
Model C (Table 5) confirms H1; GENDER is positively and statistically significantly related to DNTKNW, with an average effect of γ 10 = 0.677, SE = 0.104, p < .001. Variation between countries is statistically significant, u 0 = 2.813, p < .001. The results reveal that AGE is positively and statistically significantly related to DNTKNW in Model E, with an effect of 0.023, SE = 0.004, p < .001. Because the effect of GENDER is 0.674, SE = 0.104, p < .001, the result confirms H2. The variation between countries in this analysis is u 0 = 2.779, p < .001. In addition, INCOME is negatively and statistically significantly related to DNTKNW in Model F, at −0.192, SE = 0.027, p < .001. Both GENDER and AGE continue to be positively and statistically significantly related to NOANSW (GENDER 0.559, SE = 0.087, p < .001; AGE 0.017, SE = 0.003, p < .001), in support of H3.
In Model D, the effect of GENDER is 0.474, with SE = 0.130, and p = .003, but the effect of THORGH is not significant, −0.047, with SE = 0.027, and p = .115. In this case, H4 must be rejected. In Models G and H, neither TRUST1 nor TRUST2 is statistically significantly related to DNTKNW, with effects of 0.092, SE = 0.094, p = .333 and −0.013, SE = 0.069, p = .851, respectively. But TRUST3 (Model J) reveals a significant effect, 0.111, SE = 0.030, p < .001, indicating no support for H5 but confirmation of H6. Respondents who generally trust organizations choose the “don’t know” answer option less frequently. However, trust in other people plays no significant role.
Group-level effects
In the cross-level model (Model K, Table 6), DNTKNW does not depend on the level of GCO2SV as a level-2 predictor with statistical significance (γ 01 = −0.342, SE = 0.657, p = .604). The average effect for DNTKNW between male and female respondents again reveals an increase to γ 10 = 0.677, SE = 0.102, p < .001, but the cross-level interaction of GCO2SV on GENDER/NOANSW is not statistically significant, such that γ 11 = −0.198, SE = 0.266, p = .460. Similarly, in Model R, DNTKNW does not depend on the level of GPOWSV with any statistical significance (γ 01 = 0.044, SE = 0.551, p = .936; γ 11 = −0.047, SE = 0.173, p = .787). Thus, H7–H10 are all rejected, and cultural values appear to have no significant influence on how often respondents choose the “don’t know” answer option.
Robustness tests
A series of robustness tests informs the results. First, Model KR2 is assessed with Hofstede’s individualism dimension (HOFIND) as an alternate proxy of culture, which correlates with GCO2SV at −0.265, p = .068 across the countries of this study. Accordingly, NOANSW continues to depend on HOFIND with statistical significance (γ 01 = 0.013, SE = 0.004, p = 0.010), but the cross-level interaction of HOFIND on GENDER/NOANSW has lost significance (γ 11 =0.001, SE = 0.001, p = .264). The effects for DNTKNW continue to be nonsignificant. In Model QR2, GPOWSV is also replaced with HOFPOW. HOFIND has a statistically significant effect on NOANSW (γ 01 =0.013, SE = 0.004, p = .004), but the cross-level interaction of HOFPOW and HOFIND on GENDER/NOANSW are no longer significant (γ 11 = 5.510 × 10−4, SE = 0.001, p = .678; γ 12 =0.001, SE = 0.001, p = .264). In summary, only the cultural effects of collectivism can be replicated with Hofstede’s dimensions.
Second, the other cultural dimensions of the GLOBE study—institutional collectivism, assertiveness, future orientation, performance orientation, gender egalitarianism, and humane orientation—potentially could explain further variance. Ashraf et al. (2017) argue for limiting analyses to culture dimensions that are strongly tied to the research focus, yet arguably, “ignored cultural factors … are contributing as much if not more to the observed effects” (Sivakumar and Nakata, 2001: 556). In iterative replacements of GCO2SV with other GLOBE dimensions in Model K, future orientation, performance orientation, and humane orientation exhibit statistical significance, but they lose significance in combination with GCO2SV.
Third, as Frank et al. (2013) suggest, it is useful to quantify the robustness of the effects using the application by Rosenberg et al. (2018). Because the underlying populations of Models A through R are different and the WVS data set is collected by many different interviewers with potentially different interviewing styles, this test is especially important. To invalidate the effect of GENDER on NOANSW in Model K, about 62.230% of the data would have to be replaced with samples for which there is no gender effect; and to invalidate the main (slope) effect of GCO2SV, about 7.481 (44.001) percent of the data would have to be due to bias, which is improbable.
Fourth, as the distributions of NOANSW and DNTKNW are highly skewed (see Figure 2), the models in Table 4 are rerun with a Poisson model. Because of the huge data size, HLM 7.03 software runs out of memory and often stops the computation after six or eight iterations. Yet, the resulting coefficients and their significance are very similar to the models with assumed normal distributions.
Discussion
Summary
Transparency is a ubiquitous notion, ranging from a legal requirement in tax declarations to an implicit expectation in surveys (Jain and Jain, 2018). Without complete and correct survey data, researchers and practitioners cannot draw meaningful conclusions. But when survey respondents evade answers consciously or unconsciously, they violate the requirement of transparency.
Considering the high nonresponse rates to items in surveys such as the WVS, “it is important to learn about the determinants of nonresponse behavior” (Riphahn and Serfling, 2005: 522). This article examines the evasive answer bias in surveys, which might manifest as skipping questions or consciously choosing a “don’t know” answer option. Both measures correlate only very loosely, indicating that they are distinct manifestations of evasive answer bias. They relate to demographic and attitudinal variables at the individual respondent level, as well as to cultural values at the national level (see the overview of hypotheses in Table 1).
First, the individual-level variables exert similar influences on both types of evasive answer bias, such that, female respondents (H1), older respondents (H2), and respondents with higher incomes (H3) evaded questions more than male, younger, and lower income respondents. Respondents, who characterize themselves as generally careless, skip more questions than conscientious respondents (H4). Though, conscientiousness is not significantly related to how often the respondents choose the “don’t know” answer option. Hence, including a “don’t know” answer option in surveys is not a catch-all option for careless responses. It rather provides a veritable measure that can reduce noise when respondents lack a strong opinion (non-attitude reporting; Dobronte, 2014) or do not want to share their opinions.
This research also reveals that skipping answers is not an expression of a lack of trust in either the interviewer or the organization conducting the survey (H5 and H6). Though, the frequency of the “don’t know” option is related to respondents’ general lack of trust in organizations. Trust in other people, which includes the interviewer, does not seem to influence respondents’ evasive answer behavior, but more “don’t know” answers may signal respondents’ lack of trust in the organization conducting the survey.
Second and at the country level, cultural values exert distinct effects across the two dependent measures (H7 to H10). How often respondents choose the “don’t know” answer option does not seem related to any of the cultural dimensions. In contrast, skipping answers happens less often in both collectivistic and low power distance cultures (H7 and H9), likely because people in these cultures tend to be more transparent in terms of their information sharing (Turilli and Floridi, 2009) and more ethically determined by their devotion “to the fulfilment or violation of expectations” (Bartels, 1967: p. 21). Across countries, the effect of cultural values is significantly stronger for female than for male respondents (H8 and H10).
Implications
First, evasive responses must be measured and monitored by researchers who should then “take this information into consideration when interpreting the score” (Charter, 2000: 315), with the clear recognition that “we can learn not only from truthful responses…, but also from the systematic patterns of nonresponses” (Jensen et al., 2010: 1496). These recommendations are particularly critical if important decisions, with substantial implications for stakeholders, rely on survey results in settings in which the survey respondents have not been suitably motivated (Osborne and Blanchard, 2011). Developing an understanding of the mechanisms that drive item nonresponse in surveys can support the further development of techniques to either reduce evasive answer bias or meaningfully analyze it (Riphahn and Serfling, 2005) and thereby substantively increase the managerial value of surveys.
Second, evasive responses cause missing data through a nonrandom process. Yet, researchers typically do not spend much time implementing effective strategies to address this issue. “A tacit understanding that missing data is a trivial nuisance seems to be the rule” (Myers, 2011: 298). The most frequently chosen and default approach in many statistical packages is listwise (aka, casewise) deletion (Tanguma, 2000), that is, simply discarding any respondent with a missing measurement on any item. But listwise deletion disproportionally reduces the effective sample size for certain groups of respondents, such as female, older and more well-off respondents (refer the hypotheses examined in this article, Table 1). Besides, disproportionally more responses from females are lost in countries characterized by individualistic and high power distance values (H8 and H10). Pairwise deletion is an alternative and lesser-used method, which discards cases only when the estimate requires a response to an item. This also results in biased estimates. For example, in regression analysis, different respondents are included in the estimate of each separate coefficient. Mean substitution imputes the mean of a variable in missing places and returns a complete data set with all respondents. Unfortunately, this approach artificially deflates the variation of a variable and potentially changes the value of regression estimates. Given the biases caused, researchers are generally advised to avoid the aforementioned methods and utilize statistically more appropriate approaches (Newman, 2009).
Then, multiple imputation (Rubin, 1996) replaces each missing value with a probability sampled set of values, ultimately creating several and slightly different data sets. The statistical analysis would then be performed separately on each data set and the results combined. Hot deck imputation (Andridge and Little, 2010; Myers, 2011) or k-nearest neighbor (Fix and Hodges, 1989) are somewhat easier to use; both attempt to find similar respondents in the data set and predict the missing data from them. Whereas all missing data treatments are naturally imperfect (Newman, 2014), the latter two approaches allow for retention of the complete data set and perform realistic imputations based on answers observed by other respondents. It is specifically recommended that cross-cultural researchers use imputation rather than deletion approaches on data sets from international surveys.
Third, greater understanding of the relationships of individual-level and group-level variables with answer bias in surveys also increases the chances of designing surveys that provide valid responses. The results of this study reveal, for both researchers and practitioners, key differences and similarities in people’s reactions when it comes to answering survey questions. Cross-cultural researchers specifically need to recognize that “off the shelf” questionnaires are not necessarily appropriate for measuring what they want to measure in a certain target group (Jenkinson et al., 1993). It may be necessary to deploy questionnaires that reflect considerations of the respondents’ demographics and cultural background, rather than generic ones. To identify questions that create potential item nonresponse problems, questionnaire pretesting should be done across cultural fault lines and devote particular attention to female and older participants, as well as those with lower income.
Limitations and further research directions
Although prior literature contains some studies of response bias in surveys, relatively few aim at understanding evasive answer bias. The importance of accurate data in the advanced information age suggests the critical need for further exploration of this issue. The current study taps the large WVS data set to measure two elements of respondents’ evasive answer bias. However, across WVS waves and countries, various questions have been modified or dropped, so the percentages of evasive answer bias had to be calculated and compared. Absolute numbers could not be obtained. For that reason, it is not practical to calculate a respondent’s evasive bias for thematic groups of WVS questions. Because the calculation of the response bias would be based on fewer questions, this would reduce the explanatory power of the analysis. Evasive answer bias is an irregular and deviant behavior, and its detection needs as large and robust a data set as possible. While examining the effect for different question types needs to be assigned to future research, few other databases offer such a massive number of comparable respondents and questions as the WVS though. Future research should also investigate reasons for evading items. A suggestion would be to follow up the WVS survey with a qualitative interview asking why the respondent has evaded questions.
Furthermore, the WVS Association, as a global network of social scientists, is widely perceived as a benevolent organization, so respondents to such social surveys might express less evasive response bias than they would in surveys conducted by for-profit marketing organizations. This prediction should be investigated in further research. The household survey also involves interviews conducted inside respondents’ homes; they complete the paper-and-pencil survey in the presence of the interviewer. Although trust in other people did not significantly affect evasive response bias in this study, it would be interesting to determine if the absence of the interviewer or an outgroup scenario (Forbes et al., 2011), created by surveys conducted outside respondents’ homes, affects evasive response bias. Similarly, differences between paper-and-pencil questionnaires and electronic surveys, longer and shorter surveys could be investigated in continued research.
Footnotes
Acknowledgements
The author is very thankful to an anonymous reviewer for suggesting to describe the study’s implications for data cleaning of international surveys. Moreover, the author gratefully acknowledges generous support and funding from the Darla Moore School of Business and the Center for International Business Education and Research (CIBER) at the University of South Carolina.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
