Abstract
Structural equation modelers judge multi-item constructs against three requirements: (a) multiple items converge in a single dimension; (b) individual-level patterns of item convergence are invariant across countries; (c) aggregate-level patterns of item convergence replicate those at the individual level. This approach involves two premises: Measurement validity hinges solely on a construct’s internal convergence, and convergence patterns at the individual level have priority over those at the aggregate level. We question both premises (a) because convergence patterns at the aggregate-level exist in their own right and (b) because only a construct’s external linkages reveal its reality outreach. In support of these claims, we use the example of “emancipative values” to show that constructs can entirely lack convergence at the individual level and nevertheless exhibit powerful and important linkages at the aggregate level. Consequently, we advocate a paradigm shift from internal convergence toward external linkage as the prime criterion of validity.
Introduction
In a recent issue of this journal, Aléman and Woods (2015) question well-known measures of human values from the World Values Surveys (WVS). Specifically, the authors criticize Inglehart and Welzel’s (2005, pp. 48-56) “traditional vs. secular-rational values” and “survival vs. self-expression values” as well as the updates of these constructs labeled “secular values” and “emancipative values” (Welzel, 2013, pp. 59-73).
All four of these constructs are multi-item combinations. They have been used in dozens of publications 1 for a number of meaningful purposes, such as (a) mapping cultural differences across the major nations of the globe; (b) tracing change in human values over time and across generations; (c) linking trajectories of cultural change to socioeconomic transformations, including rising life expectancies, expanding education and informational connectedness, advancing globalization and urbanization, technological progress and improving living standards; (d) explaining differences in national policy outcomes, from gender equality to sexual liberation to peace and sustainability; and (e) modeling regime dynamics, such as where democracy emerges and flourishes and where impartial governance and other aspects of institutional quality prevail and where all of this is not the case. 2
Aléman and Woods raise doubts about this research by a sweeping claim: None of the involved value constructs is valid. The authors derive this postulate from a simple finding that they take a long way to demonstrate: The individual-level coherence among the items that constitute our value constructs varies across countries and is weak in many of them. 3 Aléman and Woods conclude from this presumably groundbreaking discovery that our value constructs are incomparable across cultures. Any cross-cultural pattern found with these constructs, such as the Inglehart–Welzel cultural map of the world, is hence unreal.
Aléman and Woods’ critique resonates with a growing consensus on the meaning of measurement equivalence. In a nutshell, this consensus centers on the following axiom: When one uses individual-level data to create a multi-item index and then calculates group-level averages on this index, these averages are comparable only if they represent in each group the same pattern of inter-item convergence (cf. Davidov, Duelmer, Shlueter, Schmidt, & Meulemann, 2012; Stegmueller, 2011; van de Vijver, 2011). As we will argue, this notion of equivalence is fundamentally flawed and needs to be replaced by a decidedly “nomological” view: Numerically similar averages of a construct are equivalent in a substantive sense whenever they map in corresponding fashion on the construct’s supposed antecedents, outcomes, or correlates. This nomological criterion is entirely about a construct’s external linkages, which—as we will show—require no internal convergence at all.
We do not debate Aléman and Woods’ evidence that our value constructs show variable and often weak inter-item convergence at the individual level within countries. In fact, we ourselves have pointed out this phenomenon a number of times (Inglehart & Welzel, 2003, 2005, pp. 231-244; Welzel, 2013, pp. 74-79, pp. 110-112).
We do, however, reject Aléman and Woods’ conclusion that a construct’s country-level scores are incomparable merely because they emerge from variable individual-level convergence patterns. Our rejection is informed by a proper understanding of the “micro-macro puzzle,” which allows for only one conclusion in this matter: A construct’s measurement features at the individual level provide no information whatsoever about the same construct’s validity at the country level. As we will outline, this insight rectifies a widespread misunderstanding of the “ecological fallacy.”
Our second contribution is to explain why inter-item convergence is a misleading standard that has created a false sense of “measurement equivalence”—a false sense especially when one evaluates constructs based on a combinatory instead of a dimensional logic. Many scholars are not aware of this distinction, which is known in the methods literature as the difference between “formative” and “reflective” constructs (Coltman, Divenney, Midgley, & Venaik, 2008). It is, hence, worthwhile to explain this distinction and to point out its implications for measurement equivalence.
Our third contribution is to demonstrate that cross-national differences in a construct’s inter-item convergence are not an infallible sign of incomparability. On the contrary, if these differences map closely on a “third” standard of reference, then these differences embody a common meaning—the very basis of comparability. We will evidence this point for the construct of emancipative values: The individual-level convergence among this construct’s items increases systematically with a major aspect of modernization at the country level: cognitive mobilization. As a construct designed to capture key features of modernization, emancipative values are not invalidated by the fact that certain properties of their measurement, including inter-item convergence, are themselves shaped by modernization. The exact opposite is the case.
In conclusion, we maintain that our value constructs are as valid, meaningful, and real as most of the literature treats them. We nevertheless appreciate Aléman and Woods’ criticism because it provides an overdue opportunity to highlight a widespread misconception of measurement equivalence.
Understanding the Micro–Macro Paradox
Although the data we use to create our value constructs are collected from individuals, the purpose of these constructs is not to measure internally convergent personality traits. Instead, we aim to capture value configurations that emerge first and foremost—and sometimes solely—at the group level, the place where culture takes shape. Value configurations at the group level, most notably countries, describe prevalence features in collective mentalities. By definition, these prevalence features represent a culture-type phenomenon that only surfaces in the aggregate and, hence, does not exist at the individual level. As Aléman and Woods examine our value constructs exclusively at the individual level, they miss the whole point of our approach.
Against this backdrop, it is helpful to remember what Kaase (1986) termed the “micro-macro-puzzle in the social sciences.” Referring to Converse’s (1964) discovery of “ideological incoherence” among individuals, the micro–macro puzzle denotes the well-known fact that incoherence between attitudes at the individual level contrasts sharply with coherence among aggregate measures of the same attitudes. With respect to public opinion, which is “public” only in the aggregate, Page and Shapiro (1992) and Stimson, MacKuen, and Erikson (2002) provide ample evidence for this pattern. It is visible in the fact that a given pair of attitudes regularly correlates much more weakly at the individual level within countries than aggregate measures of the same attitudes correlate across countries. Furthermore, individual-level correlations often vary considerably from country to country but this within-country variability does in no way diminish the between-country correlation among aggregate measures of the respective attitudes.
For instance, the “choice” and “voice” components of emancipative values 4 correlate at R = .22 at the individual level. 5 By contrast, aggregate measures of these components correlate at an R of .62 between countries. 6 Within countries, the individual-level correlation varies from −.12 to + .41, but the strong between-country correlation exists despite this variability. 7 With regard to multi-item constructs, the micro–macro puzzle means that these constructs always reach stronger and more robust convergence between than within countries. An obvious conclusion from this observation is that weak and variable inter-item convergence within countries is (a) the norm and (b) irrelevant for convergence patterns that exist at the aggregate level between countries.
Inglehart and Welzel (2003, 2005, pp. 231-244) propose three causes of the micro–macro puzzle. First, individual-level survey data are inflicted with large amounts of measurement error, known as “background noise,” that partially obscures true correlations. 8 However, much of the error is distributed randomly, so aggregation cancels most of it out–reducing a source of pattern disguise. 9 Second, pooling aggregated data across diverse sets of countries widens variability in correlated variables beyond the “clouds” within which correlations remain murky (and much of the within-country variation is inside the clouds 10 ). Third, countries are powerful entities of socialization whose cultural reproduction over subsequent generations creates inert trajectories: This implies pattern consistency in the aggregate. Culturally ingrained value orientations are no exception: The components of emancipative values cohere much more closely at the aggregate than at the individual level. 11
It follows from all this that the individual level of analysis is unsuited to reveal some of the most striking convergence patterns in values. Many convergence patterns—in particular those that generate culture—surface only in the aggregate.
Moreover, the same value orientations mean something different at different levels of analysis (Leung & Bond, 2007). If free from measurement error, scores in emancipative values at the individual level indicate how much a person prioritizes this set of values. At the aggregate level, the average in emancipative values indicates how prevalent these values are in a country. 12 In other words, the same variable measures personal preference strength at the individual level but social norm strength at the aggregate level. These are two different things, of which only the latter says something about culture. Accordingly, the correlation among the same pair of orientations also captures something different at the individual and aggregate levels.
To illustrate this point, let us consider the correlation between out-group trust 13 and emancipative values. At the individual level within countries, the correlation between these two orientations tells us to what extent a person who trusts more than most others in her country is also more emancipatory in her orientations than most others in this country. The correlation between country-level averages of the same two variables tells us something else: To what extent the prevalence of trust associates with the prevalence of emancipatory orientations. The strength of this aggregate-level association is in no way invalidated by its weakness at the individual level. 14 Actually, the association could be absent at the individual level or even show a reversed sign, and yet this evidence would be entirely inconclusive for the aggregate relationship. Since patterns at different levels of analysis mean something else, there is a wall of non-inference between them.
The Nature of Ecological Effects
There is nothing strange about the fact that strong convergence among aggregate measures of a set of orientations coexists with variable and weak convergence among the same orientations at the individual level. Actually, the social nature of human existence makes this pattern quite common—through “ecological” effects. Ecological effects are manifestations of social influence. Social influence operates in such a way that a given orientation shapes people’s other psychological and behavioral traits through the prevalence of the orientation in question, irrespective of whether a person herself embraces the respective orientation. This is a frequent but largely overlooked regularity that Welzel (2013, pp. 110-112, Box 3.1) conceptualizes as an “elevator phenomenon.”
An example is the association between emancipative values and nonviolent protests. The association exists at the individual-level within countries: Persons with stronger emancipative values participate in nonviolent protests more frequently, and this is so in every country. However, the correlation is only modestly strong, and another pattern in the association between emancipative values and nonviolent protests is much more powerful 15 : When emancipative values become more prevalent in a country, everyone’s protest activity “elevates,” regardless of whether the person in question herself endorses emancipative values (Welzel, 2013, p. 230).
Elevator effects of this kind are manifestations of social influence. They shape the fabric of societies by determining which psychological and behavioral traits become prevalent in a population. But elevator effects are inherently ecological in character and, hence, not mirrored in corresponding individual-level associations. Accordingly, there exist meaningful and consequential convergence patterns in the aggregate, which have no equivalent at the individual level. But the latter does not render unreal the former.
In light of this insight, Przeworski and Teune’s (1970, p. 73) famous dictum that an ecological correlation is spurious, if it is not reflected in the same way at the individual level within each aggregate unit, is profoundly flawed. Scholars continue to recite this quote as a warning against the “ecological fallacy.” But the irony is that this very statement is itself a flagrant illustration of a fallacy in the opposite direction of inference, known as the “individualistic fallacy”: inferring the validity of an aggregate-level relation from its existence at the individual level.
The ecological principle that shapes much of the social reality also shapes multi-item measures of values, including emancipative values: Because of ecological effects, the inter-item convergence of these values powerfully surfaces at the aggregate level between countries, while remaining variable and weak at the individual level within countries (Welzel, 2013, pp. 74-79). Accordingly, one cannot assess the equivalence of country averages in emancipative values by examining these values’ convergence at the individual level within countries. Doing so is to ignore the wall of non-inference.
Yet, judging the equivalence of country-level averages from convergence patterns at the individual level within countries is exactly what multi-group confirmatory factor analysis does—the new booming industry in cross-national survey research (cf. Davidov et al., 2012; Stegmueller, 2011). 16 No doubt, multi-group confirmatory factor analysis (MGCFA) is an excellent tool to examine item sets for dimensional unity at different levels of analyses. But dimensional unity is no criterion for combinatory constructs, which allow for (a) multi-dimensionality as well as (b) variability in the dimensionality pattern.
Dimensional and Combinatory Logic
MGCFA follows a “latent variable” logic that dominates the field of structural equation modeling (SEM). From the viewpoint of SEM, multi-item constructs are valid if—and only if—all included items converge in a single dimension and show no variability in this feature across countries. This approach involves two problematic premises.
The first premise is that multi-item constructs only make sense when their constituent items represent inseparable manifestations of a single dimension. In this logic, divergent variance among constituent items is just measurement error. Consequently, validity hinges entirely on item convergence. Constructs whose constituents do not strongly converge are “unreal” in this view.
The dimensional logic can be a useful guide of construct formation for some purposes. But there is a powerful alternative logic that informs many of the most well-known multi-dimensional constructs. These constructs can be as meaningful as one-dimensional ones and often show stronger external linkages than those (as a result of the “bandwith-fidelity” dilemma, see fn. 22). Alexander and Welzel’s (2011) two-dimensional construct of “effective democracy” is a case in point. 17 Effective democracy consists of “democratic rights” as the base component and “law enforcement” as the factor that makes the base effective. The two constituents are not supposed to be perfectly convergent. On the contrary, they are combined because they tap distinct properties of effective democracy, which is theoretically predefined as the interaction of these properties. 18 Hence, to measure effective democracy in its predefined meaning, one must measure the combined presence of these properties—no matter how strongly they correlate. Contrary to dimensional logic, this combinatory logic actually requires constituent components to be at least partly divergent. 19 The reason is that a combination of components can only make a difference, relative to what each single component does, when these components cover partly separate things. Their combination is justified merely by the fact that they represent mutually complementary qualities under an overarching idea.
In combinatory logic, divergent variance among constituent components is not considered as measurement error but as complementary reality coverage. The methods literature characterizes combinatory constructs as “formative” and juxtaposes them to the dimensional logic of “reflective” constructs (Coltman et al., 2008). Goertz (2006, pp. 10-11) addresses the same distinction by the terms latent versus ontological constructs.
At any rate, it should be clear that item convergence is an altogether inadequate criterion when the logic of construct formation is combinatory. This is important to note because Welzel’s (2013, p. 60) measure of emancipative values is introduced explicitly as a combinatory construct, not a dimensional one. Specifically, the 12 items over which emancipative values are measured are portrayed as additive qualities under the definition of emancipation. Thus, to measure a subject’s overall response to matters of emancipation, partial responses must be added up over all relevant items, in deliberate disregard of how consistent the partial responses appear throughout the entire item set. To measure an overall response on a defined field, such as emancipation, consistency among the partial responses is simply no requirement. 20 Accordingly, the combinatory logic assumes compositional substitutability among partial responses (Goertz, 2006, pp. 10-13). And substitutability is an accurate assumption when variability in the composition of partial responses does not affect how an overall response relates to its expected antecedents and consequences.
The quality criteria for combinatory constructs are twofold. Theoretically, the combination must make sense such that the components meaningfully complement each other under an overarching idea. Empirically, the combination must make a difference in that it maps closer on its expected antecedents or consequences than does each of its components. Consequently, the yardstick to judge a combinatory construct is its predictedness and predictiveness relative to other aspects of reality. Datler, Jagodzinski, and Schmitt (2013) call this criterion external validity, 21 in contrast to internal consistency. From an epistemological point of view, external validity is the preferable criterion 22 : When we have predictive power, we also have interventionist potential to change things in a desirable direction—the ultimate purpose of science.
If Aléman and Woods had assessed the value constructs of Inglehart and Welzel under external validity, their conclusions had to be radically altered. Inglehart and Welzel and their co-authors have shown in scores of publications that their value constructs associate at exceptional strength and in meaningful ways with several dozen key indicators of (1) socioeconomic development, (2) cultural legacies, and (3) institutional performance—which are some of the most fundamental aspects of societal existence (cf. Inglehart & Welzel, 2010). The correspondence of the value constructs with these aspects of social reality ranges from 60% to 80%, across almost a 100 countries representing more than 90% of the world population. Whatever the causality behind these associations might be, they are so pervasive that there can be only one conclusion: These value constructs tap something real.
The same is true of the widely cited Inglehart–Welzel cultural map. The pattern behind this map is so robust in the aggregate that it re-occurred in almost identical shape throughout six consecutive waves of the WVS, despite the fact that the country composition has been considerably changing from wave to wave. Moreover, the two dimensions on this map correlate strongly with other measures of cultural differences, taken from different data under the guidance of different concepts. For instance, the constructs of “individualism/collectivism” and “autonomy/embeddedness” share almost 80% variation with Inglehart and Welzel’s (2005, p. 137) value constructs. Strong and meaningful correlations also exist with a society’s geo-climatic, pathogenic, linguistic, colonial, and other historic features—all of which testify to the validity of the value constructs from the WVS (Welzel, 2013, p. 122). 23
Misjudgments of Measurement Equivalence
The second questionable premise is that cross-national variability in a construct’s item convergence is an infallible indication of incomparability (cf. Davidov et al., 2012; Stegmueller, 2011; van Deth, 2013; van de Vijver 2011). From this premise, one had to conclude that the same overall scores in emancipative values mean something different when they emerge from different compositions of partial scores.
However, this conclusion overlooks the possibility of compositional substitutability: The same overall performances across a thematic field map similarly on this field’s expected correlates, no matter how different the mixture of partial performances is that generate the same overall performance. Whenever this pattern exists, it is the overall performance across the field, not the composition of its partial performances, that matters—a clear case of compositional substitutability. Of course, the theoretical challenge is to identify thematic fields of such obvious relevance.
Let us consider the field of emancipative values under these auspices. Welzel (2013, pp. 84-86) shows that individual-level distributions over the item set of emancipative values are strongly mean-centered and single-peaked in each country, giving the term central tendency real meaning. Welzel (2013, pp. 74-79) also demonstrates that the constituent components of emancipative values cohere in widely different strength in different countries. But he does not conclude from this finding that the same country scores in emancipative values are incomparable.
The reason to not jump to this conclusion is straightforward: Two numerically similar scores in a given measure are comparable, if their similarity maps closely on a “third” standard of reference, a so-called tertium comparationis. Thus, comparability properly understood boils down to external linkages, not internal convergence. And external linkages is entirely a matter of a construct’s association with its expected correlates, whether these correlates operate as antecedents, consequences, or concomitants of the construct in question. In other words, the strength of a construct’s associations with its supposed correlates reveals how well this construct maps on “third” standards of reference. External linkage, in this sense, is actually the foremost measure of a construct’s reality coverage.
Comparability Testing Beyond Coherence
Cross-national variability in item convergence is too premature a finding to jump to the conclusion of incomparability. To come to a valid judgment concerning this issue, three questions need close examination.
The first question is how much variability in the strength of inter-item coherence exists independent of the country means in emancipative values. Only if a given coherence strength does not tie country means into a limited range, could one judge the same country means as in-equivalent, at least as concerns their underlying coherence. But insofar as a given coherence strength ties country means into a limited range, in-equivalence between means can at most be partial, not complete. And that partiality might cover only a minor section of the total variation.
The second, and more important, question is whether the existing coherence variability actually matters. It only would if its existence obscures the external linkages that we theoretically expect from country means in emancipative values. Only if this is the case, could one infer incomparability from variability in coherence.
The third question is to what extent coherence variability is erratic or systematic. Only if the coherence variability is erratic, can it be classified as noise that undermines comparability. If, however, this variability maps in systematic fashion on other aspects of reality, differences in coherence strength unfold on a common substance base—the essence of comparability.
Let us address these issues point by point. Figure 1 plots country means in emancipative values against the differential coherence of these values per country (as indicated by the Cronbach’s alpha). 24 It is obvious that country means are higher when the coherence of emancipative values is stronger. Hence, the same country mean in emancipative values can represent different coherence strengths, yet these differences are tied into a limited range. In other words, in-equivalence with respect to the same country means’ coherence is a minor phenomenon. To be precise, country means are to 63% equivalent as concerns their coherence strength, which is evident from the R2 in Figure 1, indicating the overlapping variance between mean scores and their coherence strength. Consequently, country means are for the most part equivalent with respect to coherence.

Country means and coherence strength in emancipative values.
The next consideration is whether the country means in emancipative values continue to associate with their expected correlates, even if we take into account the limited variation in coherence strength. To test this possibility, we regress an expected correlate of emancipative values on the country means in these values, under control of the cross-national variability in these means’ coherence.
The expected correlate of our choice here is the “effective democracy index” (EDI). The EDI is a refined measure of democracy that downgrades Freedom House’s “civil liberties” and “political rights” ratings for deficiencies in law enforcement that these ratings do not cover but which are tapped by the World Bank’s “rule of law” and “control of corruption” scores. 25 Our theory posits that country means in emancipative values predict the countries’ EDI scores fairly well. Indeed, Model 1 in Table 1 shows that country means in emancipative values explain 68% of the cross-national variance in effective democracy, across a global sample of 100 countries that represent more than 90% of the world population. Now, the question is whether this effect exists independent of variability in these values’ coherence and, accordingly, persists under control of this variability. As Model 3 in Table 1 shows, this is beyond doubt the case. In fact, the inclusion of the coherence variability does not add much to the explained variance in effective democracy.
Regressing Effective Democracy on Country Means in Emancipative Values and their Coherence.
Note. Entries are unstandardized regression coefficients with partial correlations in parentheses. Test statistics for heteroskedasticy (White test), collinearity (variance inflation factors), and influential cases (DFFITs) indicate no violation of OLS assumptions. The dependent variable is Alexander, Inglehart, and Welzel’s (2012) Effective Democracy Index, updated for 2012 and rescaled so that the theoretical minimum is 0 and the maximum is 1, with fractions for intermediate positions. Independent variables are taken from Waves 4 to 6 of the WVS, using the latest survey for each country. Thus, temporal coverage varies from 2000 to 2012. EDI = Effective Democracy Index; EVI = Emancipative Values Index; WVS = World Values Survey.
EVI stands for Emancipative Values Index, as defined by Welzel (2013, pp. 69-74). “EVI: Mean” measures per country the arithmetic population mean in these values. b. “EVI: Coherence” measures per country the Cronbach’s alpha with respect to the four constituent sub-indices of emancipative values. c. “Mean × Coherence” is a multiplicative interaction term between “EVI: Mean” and “EVI: Coherence.” To build this term, “EVI: Mean” and “EVI: Coherence” have been centered on their arithmetic means and have been introduced in this rescaled format in Model 4.
p < .050. **p < .010. ***p < .001.
Similar results are obtained for other correlates of emancipative values. 26 Consequently, overall scores on emancipative values are largely equivalent as concerns their association with other aspects of reality, despite the fact that the same overall scores can emerge from somewhat different compositions of partial scores—clear evidence for compositional substitutability. 27
The third consideration addresses the forces that induce coherence into emancipative values. If we can identify such forces and show that their presence at the country level systematically strengthens the individual-level coherence in emancipative values, then variability in this coherence can, again, not be taken as an indication of incomparability. For the variability maps on a common reference standard—the very basis of comparability.
Coherence-Inducing Forces
Due to the “general theory of emancipation,” cognitive mobilization should operate as a coherence-inducing force (Welzel, 2013, pp. 74-79). Cognitive mobilization is a pivotal aspect of modernization; it advances through expanding education, skill specialization, widening access to information, technological progress, and greater intellectual stimuli in people’s daily activities. All these are aspects of rising knowledge societies, which have increased people’s cognitive capacities. This is obvious in rising IQ levels—the so-called “Flynn effect” (Trahan, Stuebing, Hiscock, & Fletcher, 2014). Flynn (2012) himself interprets this effect as indicating a general rise in intellectual capacities, triggered by the cognitive impulses of emerging knowledge societies. Resonating with Pinker’s (2011) “escalator of reason,” Flynn speculates that cognitive mobilization also enhances people’s moral judgment capacities, improving their understanding of universal ethical principles, such as those related to human emancipation.
Confirming this assumption, cognitive mobilization explains fully 75% of the cross-national differences in emancipative values, as Figure 2 illustrates: In societies that are more advanced in cognitive mobilization, emancipative values tend to be more prevalent.

Cognitive mobilization and emancipative values.
Equally important, cognitive mobilization also strengthens the coherence in people’s moral judgment. This is obvious from Figure 3, which plots the individual-level coherence in emancipative values per country against each country’s cognitive mobilization. We see a strongly linear distribution, suggesting that emancipative values become more coherent as a country’s cognitive mobilization advances. 28

Cognitive mobilization as a coherence-inducing force in emancipative values.
Interestingly, cognitive mobilization eliminates the impact that cultural traditions seem to exert before we take cognitive mobilization into account. This is evident from Table 2 where we regress the coherence of emancipative values per country on the countries’ cognitive mobilization and their democratic tradition, using Gerring, Bond, Barndt, and Moreno’s (2005) “democracy stock” variable. 29
Regressing the Coherence of Emancipative Values on Cognitive Mobilization, Democratic Traditions, and Global Linkages.
Note. Entries are unstandardized regression coefficients with partial correlations in parentheses. Test statistics for heteroskedasticy (White test), collinearity (variance inflation factors), and influential cases (DFFITs) indicate no violation of OLS assumptions. The dependent variable measures per country the Cronbach’s alpha with respect to the four constituent sub-indices of emancipative values, as defined by Welzel (2013, pp. 69-73). Cognitive mobilization is measured using a rescaled version of the World Bank’s “Knowledge Index (KI).” Democratic Traditions are measured using a rescaled version of Gerring et al.’s (2005) “democracy stock” variable. Global linkages are measured using Dreher, Gaston, and Martens’s (2008) scores of a country’s integration into global economic, social, cultural, and political exchange. All predictors have a scale range from minimum 0 to maximum 1, with fractions of 1 indicating intermediate positions.**p < .010. ***p < .001.
The democratic tradition is a first-rate measure of cultural traditions more generally speaking: 72% of the variance in this measure across 188 countries is due to differences between the 10 culture zones defined by Welzel (2013, p. 89). This evidence makes sense because democracy is the signature feature of Western culture. Accordingly, the very endurance of democracy indicates how early and deeply cultures around the world have been “infiltrated” with Western values. From the viewpoint of institutional learning, it is plausible that persistent democratic socialization over many generations induces coherence into emancipative values: The emphasis of these values on freedom of choice and equality of opportunities addresses some of democracy’s most fundamental principles.
It is equally plausible that a country’s exposure to “global justice scripts” induces coherence into emancipative values because these scripts often invoke emancipatory ideals, such as human rights and anti-discrimination norms. As a proxy for such exposure, we use Dreher et al.’s (2008) measures of global linkages in Table 2.
The regressions in Table 2 resolve these issues quite clearly: (a) Western heritage, manifest in the strength of the democratic tradition, shows no more effect on the coherence of emancipative values, once we take cognitive mobilization into account; (b) the same is true for global linkages: no effect after controlling cognitive mobilization; (c) the latter, by contrast, powerfully induces coherence into emancipative values—fully irrespective of the democratic tradition and global linkages. Accordingly, coherence is induced into emancipative values by a key aspect of modernization. This finding underlines the validity of emancipative values as a measure of modernization’s manifestation in collective mentalities.
Nevertheless, there are reasons to remain skeptical. Perhaps, many people’s responses to survey questions on matters of emancipation are meaningless because these people do not understand what emancipation is about. 30 All the more, this might be the case when the respective issues are not controversial in a society, reflecting a tacit consensus on traditional morality.
As plausible as this suspicion might appear at first glance, it is untenable on closer examination. The WVS does not ask respondents for their position on emancipation in an abstract sense. Instead, the WVS directly addresses such down-to-earth topics as male dominance, child obedience, and heterosexual norms. It is hard to believe that people have no first-hand experience with such fundamental realities of everyday life: These issues are integral parts of what might be described as “evolutionary normality” in moral systems.
For centuries and millennia, male dominance, child obedience, and strict heterosexuality have been the norm throughout most human societies. Taking the opposite—emancipatory—positions on these issues is an “evolutionary novelty” that signals a breakup of traditional limitations on human morality (Alexander, Inglehart, & Welzel, 2015). Where societal conditions have not matured to this point, respondents will naturally stick to the evolutionary norm and take traditional positions on questions of emancipation. If so, we will inevitably obtain a low score in emancipative values, which tells us something real: How little appeal emancipatory ideals have in a society.
Whether the majority of a society wishes to be measured against the standards of emancipation is a different question. But this question should not concern researchers when a society’s performance on these standards has predictable consequences for such important things as human rights, democracy, peace, and sustainability.
The Flaw of Within-Group Fixation
MGCFA has become the chief tool in assessing measurement equivalence. Unfortunately, the standard practice of MGCFA is flawed in its fixation on within-group configurations. The model in Figure 4 illustrates this point for the hypothetical relationship between life satisfaction and perceived freedom. 31 In the eyes of MGCFA, the key question here would be whether life satisfaction and perceived freedom represent a single latent variable, something like a higher-ordered subjective well-being factor. To answer this question, MGCFA examines whether the two variables relate to each other in the same way within each group. Thus, the group mean becomes the standard of reference while differences between groups are ignored. Doing so assumes that the location of group means does not matter when in fact this might make a big difference, especially if differences between groups dwarf those within groups—a common pattern when the groups are countries. Thus, MGCFA eliminates a major source of variation and reduces the question of uni-dimensionality to small-scale variation within groups.

Hypothetical relationship between life satisfaction and perceived freedom.
In our illustration, MGCFA would draw the conclusion that life satisfaction and perceived freedom do not reflect a common well-being dimension because they do not co-vary everywhere in the same way relative to the given group mean. But this conclusion ignores that the two components behave quite similar if one uses one-and-the-same reference standard, namely, the global mean. Judged against the global mean, people who are relatively satisfied with their life also believe to be relatively free in our illustration. If this were indeed so, it would be the more important information. Yet MGCFA blinds out precisely that part of reality.
In light of this illustration, the problem of MGCFA is its wholesale fixation on within-group configurations when literally everything that defines culture takes place at the group level, shaping configurations that exist between groups but not necessarily within them.
Conclusion
Advocates of structural equation modeling judge multi-item constructs against two standards: (a) multiple items converge in a single dimension; (b) within-group patterns of item convergence are invariant.
These requirements involve two far-reaching premises: Measurement validity hinges solely on a construct’s internal convergence, and configurations within groups have priority over those between them.
We have argued that both premises are profoundly flawed for some clear reasons. To begin with, configurations between groups—especially countries—exist in their own right, have real consequences, and are more clearly structured than configurations at the individual level within groups—the level where measurement error is abundant. Also, societies are powerfully shaped by ecological effects for which no individual-level equivalents within groups exist. Consequently, the premise of ontological primacy of within-group configurations over those between groups is mistaken. In fact, when we deal with culture, the exact opposite ontological order applies: Between-group configurations are far more important than those within groups.
Next, internal convergence can be a point of consideration to examine multi-item constructs. But external linkage is another aspect and one that tells us more about a construct’s outreach into reality, including its predictedness, predictiveness, and explanatory power. And while internal convergence is only a criterion for dimensional constructs but not for combinatory ones, external linkage is a criterion for both. Thus, the priority of internal convergence is a fallacious premise too.
As both premises are flawed, the SEM-approach has no explanation for an overlooked but frequently occurring phenomenon: A construct shows weak and variable internal convergence but nevertheless exhibits powerful external linkages. Specifically, the same overall scores on a multi-item construct can emerge from differently composed partial scores, and yet these overall scores map in similar fashion on a construct’s expected antecedents and consequences. This phenomenon indicates compositional substitutability: Variable compositions are mutually substitutable as long as they produce the same overall score. Whenever we encounter compositional substitutability, we have discovered a thematic field on which the overall performance matters more than the composition of partial performances. Since compositional substitutability is the anti-thesis of internal convergence, it is beyond the comprehension of the SEM-approach.
Aléman and Woods’ criticism is based on the premises of the SEM-approach and mistaken for this reason. If our theories of modernization and emancipation involved value constructs supposed to cohere inside individuals, Aleman and Woods’ critique would have a point. But that very explicitly never was our goal. Instead, our constructs intend to capture value configurations that emerge first and foremost, and at times solely, in the aggregate. Only value configurations that exist in the aggregate can have an impact on other systemic phenomena of some relevance, such as human rights, democracy, peace, and sustainability. As long as we are interested in such outcomes, we should continue to study aggregate value configurations. At the same time, we should stop judging these aggregate configurations by whether they are replicated at the individual level. From a societal point of view, the aggregate is a reality level in its own right.
In conclusion, we advocate a decidedly “nomological” view: Constructs should be judged valid whenever the same overall scores map in corresponding fashion on expected antecedents or consequences—fully irrespective of internal convergence. Scholars should consider something as real when it shows up as real in its preconditions and outcomes.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The preparation of this article was supported by a subsidy granted to the Higher School of Economics (HSE) by the Government of the Russian Federation for the implementation of the Global Competitiveness Program. Data (“misconcept.dta”), syntax (“misconcept.spv”), and documentation (“Misconcept_ReadMe.pdf”) of our analyses are available for download at CPS’s Dataverse site.
Notes
Author Biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
