Abstract
We address methodological challenges in cross-cultural and cultural psychology. First, we describe weaknesses in (quasi-)experimental designs, noting that cross-cultural designs typically do not allow any conclusive evidence of causality. Second, we argue that loose adherence to methodological principles of psychology and a focus on differences, while neglecting similarities, is distorting the literature. We highlight the importance of effect sizes and discuss the role of Bayesian statistics and meta-analysis for cross-cultural research. Third, we highlight issues of measurement bias and lack of equivalence, but note that recent large-scale projects involving researchers across many countries from the beginning of a study have much potential for overcoming biases and improving standards of equivalence. Fourth, we address some implications of multilevel models. Cultural processes are multilevel by definition and recent statistical advances can be used to explore these issues further. We believe this is an area where much theoretical work needs to be done and more rigorous methods applied. Fifth, we argue that the definition of culture and the psychological organization of cross-cultural differences as well as the definition of cultural populations to which research findings are generalized requires more attention. Sixth, we address the scope for anchoring cross-cultural research in biological variables and by asking multiple questions simultaneously, as advocated by Tinbergen for classical ethology. Bringing these discussions together, we provide recommendations for enhancing the methodological strength of culture-comparative studies to advance cross-cultural psychology as a scientific discipline.
Psychology is facing unrest—lack of replicability of landmark studies, weak standards in conducting research and an overreliance on seemingly arbitrary standards to judge whether a finding is worth reporting (e.g., p < .05, or p < .01), combined with weak theoretical foundations are issues that have been raised in our discipline. The question is, how should we deal with this in cross-cultural psychology? We could drift toward pessimism and despair, on one hand, or react with rejection and denial, on the other hand. Yet, these discussions and the scrutiny by the public of the field of psychology also carry within them the opportunity to critically reflect on our research practices and improve them. Cross-cultural psychology is claimed to be a field of science, that is, it aims at using scientific principles in a quest to uncover regularities and differences in behavior across contexts. Behavior differences between diverse groups are easily detected by even the most casual observer. Note diversity in dress, food, constructing shelters, or any other behavior that clearly demarks individuals from one corner of the world from those in another part. At the same time, the processes and functions underlying overt behavior are much more difficult to probe and examine. The core concept in addressing this diversity is “culture,” but determining what “culture” signifies and what role any form of it plays in this mix is a serious challenge.
The gold standard in psychological science for understanding causal relations and prediction has been the experimental paradigm. Individuals are randomly assigned to various treatments that are (ideally) completely under control of the experimenter, and the differences in behaviors (measured scores or behavioral observations) are then statistically analyzed. As we emphasize later on, the assumptions on which such an experimental design is based by and large cannot be met in culture-comparative research. Furthermore, the experimental approach has come under scrutiny because of how it has been implemented and practiced in psychology. First, the way in which experimental research tends to be conducted has been criticized over several decades, but particularly in recent years (e.g., Asendorpf et al., 2013; Bakker, van Dijk, & Wicherts, 2012). Simmons, Nelson, and Simonsohn (2011), writing on “false-positive psychology,” mention three reasons for false-positive findings, namely, flexibility in data collection, in analysis, and in reporting. These challenges are not directed at the logic of the experimental model per se, but at the lack of rigor with which it is being applied, commonly termed “researcher degrees of freedom.” Second, the way in which psychology is determining the presence of a finding worth reporting has been questioned. Differences between score distributions are analyzed, most often by testing the null hypothesis that there will be no statistically significant difference between means, or some other statistic. A finding is considered worth reporting if nominally p ≤ .05, which refers to the probability that, given the chosen statistical model, the statistical summary of the data is equal to or more extreme than the observed data (Wasserstein & Lazar, 2016). The statistical model often assumes that there is no difference, the so-called null hypothesis. In our case, the most common hypothesis is that there is no difference in means or correlations between samples. This chase for statistical significance has been criticized extensively. In the area of health, Ioannidis (2005) stated, [A] research finding is less likely to be true when the studies conducted in a field are smaller; when effect sizes are smaller; when there is a greater number and lesser preselection of tested relationships; where there is greater flexibility in designs, definitions, outcomes, and analytical modes; when there is greater financial and other interest and prejudice; and when more teams are involved in a scientific field in chase of statistical significance. (p. 696)
To provide an overview, we address what we see as five methodological issues that need to be addressed in culture-comparative studies. These are (a) the questionable status of the experimental model in cross-cultural research, (b) the degrees of freedom that researchers have in pursuing desirable outcomes and the consequences for null-hypothesis testing, (c) the difficulties of differentiating between valid differences and bias in data, (d) the complication that culture in psychology is addressed at the individual level and at the level of populations, and (e) the need to define what is meant with “culture” in a given study and how cultural populations are selected and sampled. We then turn to the exciting developments in cultural neuroscience. We argue that it has an important role to play because it can help us to add much-needed biologically oriented perspectives to cross-cultural research in the social sciences, but we also note concerns about some recent work in this nascent field that need addressing. At the theoretical level, we plead for cross-cultural research that addresses multiple perspectives simultaneously, following classic work by Tinbergen (1963). Bringing these various points back together, we list a set of recommendations for making findings more robust and less vulnerable to alternative interpretation.
The Experimental Model in Cross-Cultural Psychology
A true experiment, or randomized experiment, is based on two assumptions that are difficult to meet in cross-cultural research. The first assumption is that subjects, or participants, are assigned to treatment conditions randomly, this is also referred to as the assumption of exchangeability. The assumption is violated when participants are nested in groups and, by implication, assigned to the treatment represented by their own group, as is invariably the case for groups defined in terms of culture. Methodologists have coined the term “quasi-experiment” to denote the nonrandom nature of this natural experiment; within the cross-cultural literature, “static group comparison” has also been used (e.g., Van de Vijver & Leung, 1997, 2011). However, using naturally occurring variation between cultural groups greatly increases the burden of proof. If we compare two groups of participants drawn from different societies, an observed difference can be due to myriad variables, which are not the target of the study and several of which will be hard to rule out as providing a possible explanation. The second assumption in true experiments is that the experimenter has control over—and thus can manipulate—the experimental conditions (or treatments). In our field, most (“cultural”) conditions are hardly open to manipulation. 1
The challenge in quasi-experimental designs is to determine the plausible cause–effect relationships. In cross-cultural research, there are treatments that can be fairly well defined; some being factors in the external context, such as climatic conditions (Van de Vliert, 2009) and presence or absence of sufficient iodine in the drinking water and its consequences for cognitive functioning (Bleichrodt, Drenth, & Querido, 1980), and others relating to specific practices common in a population, such as the direction of reading and writing (Román, El Fathi, & Santiago, 2013). However, in most cross-cultural research, psychologists are interested in the role broad social or mental “treatments” play in psychological processes (e.g., an interdependent self or values acquired during socialization). Such treatments can only be surmised retrospectively; their presence in the past is typically inferred from available information rather than demonstrated. For example, individuals raised in an East Asian country are often assumed to be more collectivistic and less individualistic compared with individuals raised in Western Europe or North America. Yet, the theoretical rationale that justifies such assumptions is obscure or relies on circular logic: Asians are more collectivistic, therefore, they are more modest, express emotions less explicitly, pay more attention to the social context, form smaller networks of tight-knit members, and so forth, which then implies that they are more collectivistic. What is required is an explicit test of the antecedent conditions that presumably make one group more collectivistic, which, in turn, then can be linked to other psychological outcomes.
So far, we paid more attention to mean differences, but it is important to note that the experimental model is not limited to the analysis of differences in mean scores between samples. Hypotheses can also be formulated in terms of relationships between variables; for example, testing the null hypothesis that the correlation between two variables is the same in two populations. There are various possible extensions: Multivariate analyses such as factor analysis can be conducted to identify psychological structures, and differences between cultures can be modeled in terms of intervening or mediating effects (which requires partial correlation). Insofar as “culture” is part of a causal chain of argument, the same arguments apply equally to complex and simple designs. Below, we return to the use of multivariate analyses and discuss their use for validation purposes before mean comparisons (the issue of cross-cultural equivalence).
Certainly, it is possible to attempt to measure assumed mental states that are thought to be causal in a cause–effect sequence. This is akin to a manipulation test in experimental research (Smith, Fischer, Vignoles, & Bond, 2013). Nevertheless, unmeasured “third” variables are particularly potent threats to inappropriate interpretations (e.g., a third variable may influence both the manipulation check and the dependent variable). All in all, the interpretation of results pointing to cross-cultural differences requires careful consideration. At the same time, there is also the threat that existing differences are not found, for example, due to small samples (resulting in low power; Button et al., 2013) or reference group effects (i.e., due to respondents’ judgments being anchored in their own group rather than a common standard; Heine, Lehman, Peng, & Greenholtz, 2002). Experimental procedures may also fail to elicit the predicted differences if the manipulation was not culturally appropriate or sensitive enough (see also discussions by Boer, Hanke, & He, 2018).
Academic Degrees of Freedom in Cultural Research and Null-Hypothesis Testing
Researchers have considerable degrees of freedom in conducting their studies. Many decision points arise that are often implicit and may not be documented well, hindering the possibility to effectively replicate studies. In addition to these implicit biases in the research processes, there are explicit choices, which researchers may take to achieve significant results that are considered worth publishing.
One of the most telling observations is the overemphasis on differences in our field. The “chase of statistical significance” mentioned by Ioannidis (2005) had been illustrated by Brouwers, Van Hemert, Breugelmans, and Van de Vijver (2004), who analyzed hypotheses and outcomes in a set of 80 culture-comparative articles published in the Journal of Cross-Cultural Psychology. In 69% of these articles, the authors only had formulated hypotheses postulating differences between cultures. In contrast, the findings in the majority of articles (71%) pointed to similarities (invariance) as well as differences. There was not a single article in the entire set where only cross-cultural invariance had been predicted and/or was found, although the distributions of differences and invariant results compellingly suggest that there should have been at least some such articles. The most likely explanation is that studies finding “no differences” do not get published, pointing to publication bias. 2
Simmons et al. (2011) identified a number of procedures that researchers routinely use to increase their chances to find significant differences. Probably, the most relevant to cross-cultural psychology are (a) to sample many groups, increasing the probability of some significant findings; (b) to use small samples, which increase the chance of error fluctuations; and (c) to include multiple measures of the construct of interest and then emphasize those variables that show significant differences. It can be tempting to claim “partial support” or “partial validation” for conjectures, rather than to accept rejection. However, we should remind ourselves that, according to Popper (1959), a theory to the effect that raven are black is falsified when a single nonblack raven is observed. “Partial support” implies that a theoretical position has boundary conditions that need to be addressed. Given the prevalence of confirmation bias of hypotheses in psychology, we suggest that until the theory has been modified and the modification tested, the negative evidence should be taken as a sign that the theoretical proposition is overall incorrect and needs to be modified. But we also emphasize that true exploratory research is valuable, and rigor in hypothesis testing does not prevent in any way ad hoc interpretations of unexpected findings; our concern is with post hoc hypotheses as well as uncritical acceptance of some hypotheses, whereas others are rejected.
An additional concern for cross-cultural research with a somewhat paradoxical effect is the possible presence of cultural bias in data (see the next section). If there is an even slight bias, increasing the size of samples will lead to a higher probability of finding a statistically significant difference. A larger data set will lead to more precise estimates of mean scores, but in the presence of bias, a mean will include the effect of bias (Malpass & Poortinga, 1986).
Criticisms of research practices in psychology and elsewhere have often focused on the consequences for null-hypothesis statistical testing (NHST), taking as a lead the unrealistically high rate of rejection of the null hypothesis that defies belief (e.g., Fanelli, 2012). One way in which this has been addressed is by asserting the importance of effect size (Cohen, 1994; Matsumoto, Grissom, & Dinnel, 2001). A statistical significant outcome of NHST is then a minimal condition for the interpretation of differences; more relevant is the size of an effect, reflecting the proportion of explained variance. In fact, when sizable cross-cultural differences are anticipated, a small observed effect size should be qualified as a negative outcome. Cohen (1994) famously has issued the warning: All psychologists know that statistically significant does not mean plain-English significant, but if one reads the literature, one often discovers that a finding reported in the Results section studded with asterisks implicitly becomes in the Discussion section highly significant or very highly significant, important, big! (p. 1001)
A step away from NHST is the use of Bayesian analysis in which a prior probability is specified for a hypothesis on the basis of existing information, and a posterior probability is derived reflecting the change in the prior probability informed by the evidence gained from a study. Thus, Bayesian analysis amounts to an update of prior beliefs in the light of new evidence, using Bayes’s theorem (Muthén & Asparouhov, 2012), whereas the traditional testing of the null hypothesis is neutral on the prior evidence and only relies on the evidence of the data observed in the study at hand. The positive point about Bayesian inference is that it combines new evidence with prior knowledge and opinions. This can be done repeatedly with, possibly, an accumulative shift in the probability of a certain state of affairs.
Bayesian analysis comes with its own weak points, of which the estimation of the prior probability of a state of affairs being true is the most important. The shift in probability as a consequence of new evidence partly depends on the prior probability. For example, if you do not believe in extrasensory perception, you will assign a prior probability of near zero to an expected positive outcome, and only very strong evidence can lead to accepting a posteriori that extrasensory perception does exist. However, if you have good reason to believe that individualism-collectivism is a major distinction between East Asians and European Americans (which corresponds to a very high a priori probability), a single study reporting negative findings will have a limited effect; in other words, a prior belief is not immediately reversed by a single set of incompatible data. The subjectivity of choosing the prior probability is a significant concern and needs careful attention by researchers. However, in the long run, traditional NHST and Bayesian analysis should lead to very similar outcomes.
Meta-analysis and replication research, geared toward validating existing results, are important ways for consolidating findings. At the same time, some of the issues that we discussed are also of concern with meta-analysis. For example, there is a risk that a cultural bias component in some variable is shared by various studies and remains embedded in the consolidated findings. An advantage of meta-analysis is that, if there are multiple operationalizations of the same construct, it is possible to specifically test the influence of such method variables. Furthermore, meta-analysis can be used to test other theoretically meaningful moderators of effect sizes (e.g., economic, social, or cultural processes that influence means or correlations). With replication, a controversial point is whether exactly the same method and procedure have to be followed as in the original study (literal replication) or whether only essential features should be retained (constructive replication; Lykken, 1968). In cross-cultural psychology, this strategy comes with the additional difficulty that literal replication of a study may not be possible. If measures are not equivalent (see the next section) and need to be adapted for a new target group, this can be said to amount to nonliteral replication. If anything, such concerns underline the importance of replication as a strategy in cross-cultural research; it is being addressed in more detail by Milfont and Klein (2018).
Differentiating Between Valid Differences and Data Bias
Psychological traits and processes are intrinsically difficult to measure. At one end, we have quasi-experiments, in which we may observe overt behavior or measure brain activity and physiological states in samples from distinct populations. At the other end, we may directly question individuals about their feelings and attitudes. With both types of data, the challenges for interpreting the results as evidence for or against postulated cultural differences can be formidable. There are two key issues. The first concerns confounding, an issue that applies to many fields of inquiry and one of the main concerns for quasi-experimental methods (see above). A confounding variable is a third variable, not included in a study, that correlates with both the dependent variable and the independent variable and can account for an observed relationship between them. Applied to cross-cultural psychology, the number of years of school education and response styles are two examples of confounds that are likely to be present, but there are numerous other variables. For example, Bender and Chasiotis (2011) found that number of children can be a factor in accounting for differences in autobiographical memory between samples from independent and interdependent populations. The other key question is whether measures (behavioral, physiological, or verbal) are equivalent. The cross-cultural equivalence and bias framework has been developed focusing primarily on verbal measurements (e.g., self-reports in surveys), although it is relevant also for observations and physiological measures.
Valid comparison of scores of respondents with a diverse background requires that the instrument can be taken as a common standard of whatever it is supposed to assess. For example, the comparison at face value of scores on a questionnaire for extraversion only makes sense if (a) the hypothetical construct of extraversion is the same across the groups studied and (b) the questionnaire represents the construct on an identical measurement scale; that is, the relationship of the observed score variable with the underlying construct should be identical. Hence, the quality of any psychological mapping depends on an adequate analysis of the psychometric properties of the collected data across the samples. Van de Vijver and Tanzer (2004) have listed sources of bias that may affect cross-cultural measurement. They distinguished three broad categories: construct bias (e.g., incomplete overlap in definitions of constructs across groups, differential appropriateness of the behavior sampled for a scale), method bias (e.g., differential effects of response styles, reference group effects, differential representativeness of samples), and item bias (e.g., poor translation, differential construct validity).
Findings of bias should not be seen as the end of the road, but as a starting point for further analysis into the causes of bias. Processes of translation and adaptation of instruments are closely linked to what Werner and Campbell (1970) have called “decentering.” Because an instrument is developed within a particular context, it will contain features characteristic of that context, which have little to do with the construct or domain that is being assessed and should be avoided.
The article by Boer et al. (2018) provides an overview of the extent to which analysis of equivalence is implemented and reported in current cross-cultural research. The news they bring about the present state of our field is not good—few studies actually try to rule out effects of these biases. However, researchers are far from helpless here. Boer et al. (2018) provide an overview of these issues and outline resources for addressing psychometric equivalence in cross-cultural research.
Visionary projects of past decades largely started off as the work of a single researcher or small team, with others joining in later or repeating a study elsewhere. This means that methods and instruments, as a rule, have been developed within a particular context; issues of transfer and equivalence were addressed (or not addressed) later on when a project expanded. Most well-known intelligence batteries, questionnaires, and social survey scales have been translated into multiple languages, but none was developed from the start by a broad team with representatives from a range of countries. The future of cross-cultural psychology may well lie in large projects set up by multiple researchers. Large-scale projects for international assessment of quality of education, such as the Program for International Student Assessment (PISA) and Trends in International Mathematics and Science Study (TIMMS), currently provide the highest standards of excellence (for an overview, see Cresswell, Schwantner, & Waters, 2015), even though equivalence of the data is not perfect (e.g., Asil & Brown, 2016). The most important feature of such projects is that conceptual frameworks and questionnaire scales are designed by groups of experts who represent the entire set of countries involved. Items for performance tests and questionnaires are contributed by a number of national assessment centers, and are then translated, discussed, field trialed, and edited, rather than adapted from existing instruments that were developed in one country.
Multilevel Models
Recent advances in multilevel models allow analysis of data at each level as well as analysis of the relationships between levels (e.g., Hox, 2010; Muthén & Muthén, 1998-2017) with statistical tests of interactions between higher level variables (mostly countries) and lower level variables (individuals nested within these countries). In multilevel studies, there are either separate data for each level (e.g., GDP per capita at country level and a questionnaire score on income or personal wealth at individual level) or, alternatively, information for one level can be derived from the other level through aggregation or disaggregation of scores (Van de Vijver, van Hemert, and Poortinga (2008). Aggregation is common in cross-cultural survey–type research, for example, on values or personality dimensions; a culture-level score for a variable can be obtained through aggregation of individual scores in a sample, usually by calculating the sample mean (sometimes called a “citizen mean”; Leung & Bond, 2004; Smith et al., 2013). Disaggregation occurs when individuals within a sample are attributed some characteristic on the basis of population-level information (e.g., persons from the United States should be individualistic as they live in an individualist society).
Avoiding theoretical debates for the moment, it is important to note that individual- and country-level scores based on individual-level data are based on statistically independent information (Dansereau, Alutto, & Yammarino, 1984). This implies that there can be, but need not be, a shift in meaning between levels. For example, two variables such as “feeling stressed” and “following rules at work” can be uncorrelated or negatively correlated when examining the scores of individuals within separate countries. Yet, when the same data are aggregated to the nation level, a positive correlation may appear (Hofstede, 1980). In statistical terms, if there is a monotonic function describing the relationship between scores at the two levels, the same structure applies across levels and it then makes sense to use the same concept. In the case of nonmonotonic relationships, it is misleading to use the same concepts at the two levels. The former instance has been referred to as an “isomorphic relationship,” the latter as a nonisomorphic relationship. It is possible to test more complex statistical models that investigate the conditions when the individual and aggregate correlations diverge (see Van de Vijver et al., 2008). These techniques also have important applications in the identification of measurement bias (see Boer et al., 2018).
The conceptualization of psychological phenomena at the population level is a thorny issue that has a long history in psychology, including notions such as “Volksseele” (“soul of the people,” see, for example, Wundt, 1913) and social representations (Moscovici, 1984; see also Jahoda, 1988). Evidence of isomorphism between individual level and country level has been reported, for example, for the Big Five personality factors of the five-factor model (McCrae & Terracciano, 2008). This suggests that we can use the same concepts at both levels and refer meaningfully to an extravert person as well as to a national population with high extraversion.
For values, a cornerstone of cross-cultural psychology, the evidence is more mixed. Hofstede (1980) found an interpretable factorial structure at the population level, but no clear structure at the individual level. He accepted nonisomorphism (unequal factorial structures) and insisted that his four dimensions were country-level dimensions; applying them to individuals would amount to what is called an “ecological fallacy” (Robinson, 1950). 3 Triandis, Leung, Villareal, and Clark (1985) proposed to use two sets of terms: individualism and collectivism at the level of countries, and idiocentrism and allocentrism at the individual level. These proposals have not had much following, and most studies use the terms individualism and collectivism for both populations and individuals; in much of the literature, the inhabitants of individualist countries are taken to be individualists, and in collectivist countries, the citizens are collectivists. 4
Schwartz, (1992, 1994) proposed two different structural arrangements for values at individual and country level. However, analysis of a large data set across levels suggested a two-dimensional structure at both the individual level and the country level (Fischer, Vauclair, Fontaine, & Schwartz, 2010). Despite this similarity of structural relationships between value items at the levels of individuals and countries, Schwartz (2014) has insisted for theoretical reasons on a three-dimensional solution at the country level that differs from the two-dimensional solution at the individual level. This suggests that theoretical arguments prevail over psychometric findings, or at least, that there is a theoretical incompatibility between the two levels. Offering a technical solution, Fontaine (2011) has suggested that nonisomorphism mathematically implies that measures are not equivalent, opening up a path for understanding how meanings around items and scales are culturally constructed. The question whether, and if so how, individual psychological attributes, such as values, can be nonisomorphic at this stage is difficult to answer and requires further theoretical and empirical development.
Defining Culture and Cultures
So far, we have dealt with more technical aspects of design and data collection. Yet, the central concept of “culture” also requires methodological consideration. Scientific researchers derive “interpretations” or “inferences” from their data. From a psychometric perspective, interpretations in psychology can be seen as “generalizations” from behavior measurements or observations to some psychological construct (i.e., some trait, process or behavior domain) and to some population of individuals (Cronbach, Gleser, Nanda, & Rajaratnam, 1972). In most culture-comparative studies, generalizations are from data obtained on individuals nested in samples to some postulated construct that refers to conditions in the external environment (such as climate or mode of economic subsistence) or in the psychosocial context (such as values).
“Culture” or cultural variables can be modeled as antecedent variables, as outcome variables, and as mediator or moderator variables. Given the innumerable ways in which the concept of culture is being used (Baldwin, Faulkner, & Hecht, 2006), either the term should be avoided (Poortinga, 2015) or its intended meaning in a specific research setting should be clarified. For example, the emerging field of geographical psychology avoids the concept of culture completely and, rather, focuses on the spatial distribution of psychological traits (e.g., Rentfrow & Jokela, 2016), without making reference to cultural influences.
Interpretations of cross-cultural data differ in the degree of coherence in psychological functioning that they presume (e.g., Berry, Poortinga, Breugelmans, Chasiotis, & Sam, 2011; Poortinga & Van Hemert, 2001). Insofar as differences on separate variables hang together and have the same roots, they can be interpreted in common terms; otherwise, separate explanations are required. Much culture-comparative research focuses on broad dimensions that facilitate parsimonious interpretation of differences in very diverse behaviors. Such dimensions include interdependent versus independent self-construal (Markus & Kitayama, 1991), individualism-collectivism (Hofstede, 1980; Tönnies, 1887; Triandis, 1989), tightness-looseness (Gelfand et al., 2011), analytic versus holistic or dialectic thinking styles (Nisbett, 2003). Moreover, these dimensions are often reduced to a dichotomy (e.g., between individualist and collectivist countries or regions). Dichotomies simplify research designs and facilitate communication of research findings to large audiences (e.g., collectivists are X, dialectic thinking is caused by Y), but are very hard to validate. Critical voices have argued that such dimensions as mentioned are too broad and not clearly demarcated as to which aspects of behavior they do and do not pertain to (cf. Jahoda, 2011; Segall, 1996; Smith et al., 2013). In other words, they are beyond critical testing and falsifiability. Moreover, if a study is designed with only two categories (i.e., cultures or cultural regions) on the independent variable, any difference that emerges on some dependent variable can be interpreted as support for the validity of the categorization. Obviously, this adds considerably to the probability of finding “false positives,” mentioned above.
Not all interpretations in cross-cultural psychology are in terms of rather unconstrained (and imprecise) dimensions that are nearly impossible to invalidate (falsify). In general, a construct or domain should be open to critical cross-cultural analysis insofar as it can be determined which items or stimuli do belong to it and which do not. Discussions of functional equivalence (Fontaine, 2005) need to take a more central role in our research endeavors. Constructs such as personality traits and cognitive abilities are open to critical examination across cultures, but analysis is not simple. For example, considerable effort is needed to examine the construct equivalence cross-culturally of personality dimensions as postulated in the five-factor model of personality. Focusing on the locally relevant environment within which a person is operating, other important personality dimensions might come to the fore. In fact, with local personality inventories developed in China and South Africa, it was found that the Big Five dimensions underrepresent social-relational aspects of personality (Cheung et al., 2001; Nel et al., 2012). 5
Even more open to critical research are generalizations to limited domains, such as “skills,” “rules,” “symbols,” “practices,” or “cultural conventions.” A measurement scale for such a domain can be constructed with a representative sample of the elements that belong to it; in other words, the researchers have information about content validity, or absence of it, across cultures. 6 Such a classification can be further refined through item bias analysis (Sireci, 2011). Obviously, there is a trade-off, less inclusive interpretations tend to be less parsimonious, but ultimately, we see little future for concepts that cannot be refuted.
The idea of generalization also applies when moving from samples of subjects or participants to the populations they represent (Cronbach et al., 1972). Cross-cultural research often implicitly or explicitly aims at inferences that hold for all human groups. Ideally, a representative set of samples needs to be randomly drawn from all such groups. When “cultures” are identified with countries, this means that a sample of countries has to be selected from all countries in the world (see, for example, Boehnke, Lietz, Schreier, & Wilhelm, 2011). The sample of individuals selected within each country, in turn, has to be representative of all inhabitants. A few (often affluent) countries and samples of students from a few universities are likely to be a poor representation of the population of generalization. As discussed by Boer et al. (2018), a large number of cross-cultural studies involve only two samples and rely on students. It is difficult to justify broader claims with such samples.
In traditional ethnography, the notion of “a culture” was applied to an often isolated group of people with a particular way of life, which differed from that of other groups. Each population was thought to be characterized by homogeneity (behavior variance is minimal within groups), differentiation (behavior variance with other groups), and permanence (continuity over time; Berry et al., 2011). However, upon closer examination, even in these isolated conditions, it is often hard to demarcate exact cultural boundaries, with cultural artifacts, beliefs, and norms often showing fluid transitions between groups (d’Andrade, 2001). Due to the proliferation of the application of the term “culture” (e.g., “urban culture” and “youth culture”) and ever more extensive interactions between members of all kinds of groups, cultural anthropologists have abandoned the classical characteristics.
These processes of interactions between members of different communities leading to changes over time have been the focal point of acculturation, multiculturalism, and migration research. Homogeneity, differentiation, and permanence are almost per definition problematic in this area of research within cross-cultural psychology. Diversification is the hallmark of the study of “superdiversity” in locations where dynamic communities with large variation in ethnic and national origin and belongingness live together (e.g., Van de Vijver, Blommaert, Gkoumasi, & Stogianni, 2015). Morris, Chiu, and Liu (2015) have used the term “polycultural psychology” to refer to multiple influences on individuals. Polyculturalism implies that “individuals take influences from multiple cultures and thereby become conduits through which cultures can affect each other” (p. 631). In the end, each person can be seen as a unique composition of influences and developments, conceptually represented by the term identity or social identity. Even in the case of more identifiable cultural groups, there is a tendency in the literature to accept that boundaries between them are fuzzy and variable. In view of the broad array of definitions of culture, reflecting the “extraordinary malleability of the construct” (Jahoda, 2012, p. 299), it has been suggested that it may be futile to define “culture,” and researchers should be able to follow an eclectic approach, in which they can are free to choose their own position (e.g., Baldwin, Faulkner, Hecht, & Lindsley, 2006; Soudijn, Hutschemaekers, & Van de Vijver, 1990). Yet, even in eclectic approaches, there is still a reference to assumed cultural ideals or standards that now influence individuals (see the definition of polycultural psychology above).
It seems inescapable to us that any culture-comparative study requires acceptance of some “essentialism” in culture, that is, the notion that a cultural group has specific and distinctive characteristics more shared among its members than among outsiders (e.g., Fuchs, 2001; Meyer & Geschiere, 1999). This does not need to imply that culture is a natural kind. However, within the context of a comparative study, it has to be treated as such, because the focus of our research is on whether a group of individuals as defined by a researcher differs from another group of individuals on some trait or characteristic of interest to that researcher. We cannot escape categorization processes of either samples or events in comparative research. Even the most relativistic approach assumes an empirical distinctiveness that deserves explanation. Stated another way, our research modus operandi requires “reification,” that is, dealing with abstractions (cultures, cultural groups) as if they have an existence and object-like properties. The cardinal point of this argument is that researchers should be expected to provide a precise account of their research both in terms of the constructs or domains addressed and in terms of the (groups of) participants to which their findings pertain.
Anchoring Cross-Cultural Research
In the previous two sections, we have moved from statistical and psychometric concerns to conceptual issues related to methodology. In this section, we further elaborate how behavior-in-context can be anchored, not only in terms of sociocultural variables but also physiologically and biologically. We briefly address two points: (a) the contribution (realized and potential) of cultural neuroscience as a new direction of research for resolving the methodological traps in research with overt behavioral data and (b) the scope for a more integrative approach, and consequently, more solid interpretation, by addressing various perspectives on a target pattern simultaneously.
It is convenient to distinguish two subfields in cultural neuroscience, research on psychophysiological variables (often focusing on instantiations in the brain), and research on genetic variation between groups. In psychophysiological studies, functional magnetic resonance imaging (fMRI) is currently the favorite method. The typical design involves contrasting two (due to costs and set-up constraints often small) samples of participants from East Asian and European American descent under a few stimulus conditions. Such a study has two outcomes: differences between stimulus conditions and differences between samples, where validity of the former differences is a necessary condition, but not a sufficient condition for supporting the latter differences. Given the predominance of individualism-collectivism as a framework, whatever cross-cultural differences in blood flow variation between conditions are being observed across experimental conditions tend to be interpreted in terms of an individualism-collectivism dichotomy (e.g., Chiao et al., 2009; Hedden, Ketay, Aron, Markus, & Gabriel, 2008; Immordino-Yang, Yang, & Damasio, 2014). However, correlating brain data to selected self-report measures does not provide strong proof. It might be argued that across studies, there is accumulation of evidence in the sense that in numerous small studies, significant findings have been reported (e.g., Han & Northoff, 2008; Han, Northoff, Vogeley, Wexler, & Kitayama, 2013). For reasons highlighted repeatedly by others before, it should be clear that the way these studies are conducted makes them vulnerable to false-positive outcomes (Button et al., 2013; Simmons et al., 2011). The most widely circulated critique by Vul, Harris, Winkielman, and Pashler (2009) has shown that the presence of at least some statistically significant results is almost inescapable given nonindependent analysis of voxels (small brain loci) and the large data sets in fMRI recordings. In the wake of replication failures, it is also important to consider issues of power—functional neuroscience studies due to their complexity (in terms of costs, data collection procedures, data preparation and analysis) are often underpowered. It is noteworthy (and a positive development) that recent reviews of social neuroscience have become more cautious in their interpretation of functional brain imaging studies (Allen & DeYoung, 2017).
Cultural neuroscience is a new and emerging field, and strong theoretically driven predictions of differences in specified narrowly defined brain region or an a priori defined neurophysiological pathway are still lacking. Although null hypotheses tend to be poorly protected against false-positive outcomes, developments in cultural neuroscience with psychophysiological variables are widely seen as exciting, and we agree. The dependent variables, including blood flow patterns and event-related brain potentials (ERPs) in classic paradigms using EEG (e.g., Sonke, Van Boxtel, Griesel, & Poortinga, 2008) may not be fully equivalent (e.g., due to differences in prior exposure to stimuli). Yet, the strength is that many sources of bias present in other cross-cultural research do not apply, and psychophysiological measures are much closer to a common “standard of comparison” (e.g., Poortinga, 1989) than questionnaires and other psychometric scales. We just plead for stronger research designs with larger and more diverse samples, a more theory-driven definition of target variables (brain nuclei and pathways), precise predictions of outcomes, and the inclusion of stimulus conditions where no cross-cultural differences are expected.
The second subfield in cultural neuroscience (reviewed by Chen & Moyzis, 2018) addresses differences in distributions of genetic polymorphisms (alleles) across populations, especially in neurotransmitters and hormones (e.g., Fischer, 2013; Kim & Sasaki, 2014). A first point to note is that polymorphisms potentially can provide plausible accounts of broad cross-cultural differences, such as East–West dichotomies, because the expression of any gene may well extend over a wide range of behaviors and situations. The main point of concern is that associations are explored between two kinds of phenomena, genetic variation and behavioral variation, without a clear understanding of the causal pathways from genetic structure via epigenetic expression in proteins to overt behavior (e.g., Dick et al., 2015). Any interpretation based on findings derived from only a few samples are open to alternative interpretations, especially if no molecular pathways are explicitly studied. However, extension of studies that include samples from numerous societies are becoming rapidly more feasible as genetic data banks are expanding. For a review of this field, that in our view holds great promise for the future, we refer to the article by Chen and Moyzis (2018)
These critical points are a word of caution, but we need to reemphasize that these endeavors provide a much-needed balance in cross-cultural research that has been dominated by questionnaire data. For a better anchoring of findings and to address the balance between invariance and variations in behavior, cross-cultural psychology may take guidance from classical ethology. In our view, cross-cultural psychology will gain in relevance if studies ask multiple questions simultaneously, as suggested half a century ago by Tinbergen (1963). Tinbergen was working in the field of ethology, that is, the biological study of behavior, which at the time was based mainly on field observations of behavior patterns. 7 Tinbergen argued that (inductive) analysis of behavior patterns in ethology has to address four questions: (a) causation (internal and external), (b) survival value or function, (c) ontogeny, and (d) evolution.
Two of these questions are discussed more extensively in the manuscript by Liebal and Haun (2018), who show how a broader approach may lead to a more integrated perspective on human variation and invariance. With Liebal and Haun, we see the idea of multiple questions as a scaffold for framing research in cross-cultural psychology and find it convenient to take Tinbergen’s scheme as a lead. The four questions of Tinbergen continue to be mentioned prominently in textbooks on animal behavior (Manning & Dawkins, 2012; Wynne & Udell, 2013), although their formulation has been the subject of much discussion (see, for example, Bolhuis & Verhulst, 2009), and they may not all be equally relevant or accessible in a particular cross-cultural study.
In addition to these four classic questions, we can mention another question for our field, concerning the origin and development of behavior patterns in historical time in specific context. Cole (e.g., Cole & Hatano, 2007; Cole & Packer, 2011) has described time scales at various levels, from physical time through phylogenesis, cultural-historical genesis (development in historical time), and ontogenesis, down to microgenesis (representing time from moment to moment). “Culture,” as the context for human behavior and ontogenetic development, changes over historical time. The further elaboration of questions is a matter of future analysis, but this does not detract from the principle that we should try to address research questions from more than one perspective simultaneously, even though they are addressing overlapping concerns (Poortinga, 2011).
Making Culture-Comparative Research More Robust
We outlined a number of challenges and criticisms, in no particular order. In this section, we briefly summarize recommendations on how the cross-cultural research community can attempt to meet the challenges outlined before. In doing so, we try to distill key insights from the current set of articles and highlight, where in the research process, each issue appears particularly important. Given the diverse nature of cross-cultural studies spanning the whole of psychology, our points by necessity will be relatively broad and some steps may change in order of importance, depending on the specific research domain. However, we hope to provide some guidance to aspiring students of cross-cultural research to improve on our past mistakes and errors. It is also important to emphasize that each individual step in any single project requires careful attention. Although when reading research reports, we tend to be impressed with the strongest feature of a research project (e.g., sophisticated design, elegant analysis, large samples), we should rather think of the proverbial chain: It is the weakest link that determines the strength of a project. 8
The Planning Stages
Defining the process of interest. Given the multilevel nature of cultural phenomena, it is a good practice to start thinking about the outcome variable (Kozlowski & Klein, 2000), and whether it applies to individuals, groups, or countries. Alternatively, there may be a more specific question or set of questions (e.g., what is the effect of X on Y), possibly derived from a theory or based on observations that need to be more carefully tested.
For a long time, a major distinction in the field was between universality and cultural specificity, widely referred to with the terms “etic” and “emic” (e.g., Berry, 1989) and associated with quantitative and qualitative research methods. These now tend to be seen not as mutually exclusive, but as complementary (e.g., Reichardt & Rallis, 1994). There is an increasing advocacy for mixed methods (e.g., Creswell, 2009) and “consilience” (Leung & Van de Vijver, 2008) as a strategy to strengthen the validity of cross-cultural inferences. This implies that findings are more convincing when they are based on diverse sources of data and different research methods, provided the research is designed with a view to explicit refutation of alternative interpretations (Berry et al., 2011).
2. Composition of the research team. Individuals or teams need to be identified that (a) have adequate local expertise (languages, customs, etc.) on all populations that might be included in the study and (b) can bring relevant theoretical and methodological expertise to the project. As a study progresses, the research team may need to be extended or changed. All the steps during the planning stages are not intended to be linear—more likely planning will cycle through various iterations of these steps before starting the study. Although, often, the initiator retains a leading role in a project, the ownership ideally is shared among the members of the research team. Obviously, shared ownership comes with shared responsibility of all owners and their commitment at all stages of a project.
3. Defining the theoretical network for the target psychological variables and/or process. What is the psychological trait, construct, process (PsycTCP) that is the target? It has to be specified why a culture-comparative approach is important for the project. Is the PsycTCP assumed to be universal or culturally specific (see Fontaine, 2012; Lonner, 2011; Norenzayan & Heine, 2005)? At this moment, the first step needs to be revisited and the functionality of a target variable and the level at which it operates (group level, individual level) have to be specified. How is the PsycTCP anchored developmentally, ontogenetically, historically, and perhaps phylogenetically (e.g., what are the underlying principles that bring about this PsycTCP)? Functional equivalence and construct bias issues have to be considered. It is only within the context of a theory that differences between cultural groups on some dependent variable can be predicted from their position on an independent variable.
4. Specifying which context variables (sociocultural, economic, etc.) are relevant and how they are related to the PsycTCP. The nomological network of a variable has to be considered, including mediators, moderators, boundary conditions, and discriminant validity. Are there intermediate processes (mediators); do relationships change depending on some third variable (moderators); where should we expect to find these relationships and where not (boundary conditions, discriminant validity)?
5. Ruling out alternative explanations and confounding variables. At this moment, a good theoretical model should be in place. Now it is time to play devil’s advocate: What variables would challenge the expected findings or provide an alternative explanation? What other variables/processes would imply similar or different results? How can these be ruled out theoretically and empirically? In other words, how can the mechanism and processes proposed by the research team be distinguished from alternative interpretations? The scope for spurious relationships has to be recognized. A classic case is that if samples from only two groups differ in score distributions on each of two (or more) variables, then there is a correlational relationship between the variables in the combined samples (ϕ > 0). By adding elements to the design of a study and expanding the nomological network, threats to a preferred interpretation can be addressed.
Specifically, the following are addressed:
a. Building controls for confounds into the design, that is, factors that are not a target of a study, but may point to an alternative interpretation of findings, if they are included in the design. It is impossible to control for all possible confounds, those that are most relevant theoretically and methodologically need to be targeted. One notorious example is response styles in questionnaire research.
b. Focusing on process-oriented research and testing plausible if–then relationships. If a difference in means on a single variable between two or more groups is expected, potential antecedents should be identified and included. The more precisely the theoretical processes can be described that lead to the predicted difference (specific hypotheses), the less likely it is that alternative explanations can account for the same phenomenon.
c. Tracing patterns of change in variables over time. Longitudinal studies can strengthen claims about causality and provide insights into the temporal ordering of processes.
Operationalizing Theoretical Predictions
6a. Selecting the specific populations to be sampled and deciding on the sampling strategy. In the early stages of the planning, cultural groups of interest were identified. Now, it needs to be specified what samples are required to test the hypotheses while controlling for alternative and confounding variables. There should be a clear prior reason for selecting populations on the basis of their position on the independent (causal) variable.
6b. Selecting/adapting/constructing the measures. Tests, questionnaires, and/or other data collection procedures have to be developed, translated, and adapted. It makes sense to pretest all measures in each population and to do initial checks on (item) equivalence. These steps have been extensively discussed in other sources (Van de Vijver & Leung, 1997; Van de Vijver & Tanzer, 2004). If surveys are to be used, the theoretical process should be specified through which responses are generated (Bagozzi, 2011). For example, is it assumed that there is a latent variable such as neuroticism that “causes” an individual to experience various emotionally instable experiences or are emotional reactions driven by largely independent events in a person’s life (death of a loved one, an accident, serious illness, etc.)? These different perspectives have implications for how the data have to be analyzed later. If the use of multilevel analyses with potential aggregation or disaggregation is anticipated, “composition models” (Chan, 1998) may be considered. These models refer to the functional relations that link variables across levels, including how items should be phrased, and what statistical information is required to aggregate data from a lower level to a higher level. All these questions should follow the definition and conceptualization of the theoretical variables, as specified in Steps 1 to 4.
6c. Conduct a power analysis. How large a sample is required, at the culture level and at the individual level, to guarantee adequate power, that is, the probability of rejecting the null hypothesis when the alternative hypothesis is true? Low power means not only a low probability of detecting a genuine effect but also a low probability that an observed statistically significant effect is genuine (i.e., replicable and not due to random fluctuations). The power of a study is a function of sample size and expected effect size. There are multiple free online options available (see, for example, Faul, Erdfelder, Buchner, & Lang, 2009, 2007)
7. Writing out the study protocol and preregistering the study. All the steps that have been taken up to this point are the basis for a study protocol that clearly specifies the hypotheses, the operationalization of the key variables in the hypotheses, information on all items, and procedures that are necessary for somebody else to replicate the study (including the protocol for scoring, and the protocol for analysis). We have discussed experimenter degrees of freedom above—this step is most important for addressing shortcomings. The information mentioned is common for most research. For cross-cultural research, an important additional step is the identification and analysis of bias. What steps are being taken to identify bias, how are various forms of bias going to be dealt with (in particular method and item bias), and what are the criteria for excluding conditions, stimuli/items, and individuals. Ideally, literal replication of a study should be possible on the basis of the information available to researchers (including supplementary information posted on the Internet). We strongly encourage preregistration. There are multiple formats and platforms available (such as https://osf.io; https://aspredicted.org/). This typically involves depositing the study protocol and documentation on some public Internet site, to enable later checks of the final report against the research plan and to enable other researchers to replicate the study.
Data Collection and Analysis
8. Collecting the data. The quality of data collection should be monitored; for qualitative research methods, this is essential. Any deviation from the study protocol should be documented.
9. Data checking and equivalence testing. As with all research, there will have to be a check for missing data. Further issues that arise in cross-cultural research we discussed at some length above are as follows: Are the samples (a) representative for the intended population of generalization and (b) comparable across the populations in the study (e.g., in terms of demographics, education level)? Additional data collection may be needed, but this has to be weighed against violation of the specified sampling strategy. Equivalence/invariance of the data will have to be established, excluding stimuli, items, and/or individuals (in line with the preregistered study protocol). If multilevel analyses are conducted that require aggregation or disaggregation, a check is needed whether aggregation is justified and whether the construct(s) show isomorphism.
10. Analyzing the data. Once researchers are satisfied with the quality of the data, they can proceed to test the hypotheses. Model fit, effect sizes, and confidence intervals are to be reported in accordance with the submitted study specification. Considering effect sizes and confidence can help moving beyond the problems with null-hypothesis testing that we took charge with above. Interpretation of the data has to be in line with the hypotheses and predictions; exploratory analyses should be clearly labeled as such. Unexpected patterns or results from these exploratory results may well be informative and can be important to report!
Writing the Report and Submission
With research reports, the primary objective is to have them disseminated through publication. Beyond this, data and protocols should also be submitted to some established data sharing platform so that other researchers can (re)analyze the data and possibly use them for additional purposes.
Conclusion
There can be no doubt that since the beginning of cross-cultural psychology as an established field of research in the mid-20th century, important insights and knowledge have been gained. However, using Reichenbach’s (1938) distinction between discovery and justification—or between exploration and verification—we would argue that our field has been stronger on the former than on the latter aspect of the research process. If cross-cultural psychology wishes to avoid a backlash with the relevance of findings being dismissed, standards for design and analysis need to be shored up. In this article, we have tried to set the stage for future achievements, that is, what has to be done to maintain, and perhaps expand, cross-cultural psychology as a relevant field of research. Our argument is that the field will become increasingly less relevant if, methodologically, we continue to accumulate statistically significant differences with studies based on few, often ad hoc, samples from a few populations (countries, regions), with these differences being interpreted in terms of broad but theoretically fuzzy variables. We do not argue for methodological constraints on initial exploratory studies of some new idea, but for validation and the accurate estimation of the size of cross-cultural differences, research designs are needed that allow the rejection of plausible alternative interpretations, including interpretation in terms of artifacts (lack of equivalence, bias) and confounding variables.
Are we asking for too much in this article? In field research outside the well-defined and controlled laboratory conditions, researchers will frequently have to compromise ideal experimental standards. As with any research, in the current financial and political environment, researchers face constraints of resources in term of money, time, and staff; funding for cross-cultural research may even be more precious than in other fields. In addition, scientists tend to be evaluated in terms of the number of their publications, even though it is widely acknowledged that volume of output is a rather poor index of quality. However, as cross-cultural psychology aspires to be a field of science, we will be judged ultimately by the validity of our findings. As with any scientific inquiry, if serious methodology flaws invalidate findings or do not allow ruling out plausible alternative explanations, then unfortunately, the researchers have to identify and address these issues. This may, at times, require abandoning a project (at least until more finding becomes available or a better design can be implemented), but as scientists, we have to lower the risk of making incorrect inferences about our data.
Footnotes
Acknowledgements
We would like to dedicate this article to the memory of Gustav Jahoda.
This article has profited from two different symposia that were held at the of the International Association for Cross-Cultural Psychology (IACCP) conferences in Reims, France, and Nagoya, Japan. We would like to thank the participants and presenters at these symposia for providing critical discussion and input.
Authors’ Note
The order of the authors is alphabetical.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
