Abstract
The possible dependency of criterion validity on item formulation in a multicomponent measuring instrument is examined. The discussion is concerned with evaluation of the differences in criterion validity between two or more groups (populations/subpopulations) that have been administered instruments with items having differently formulated item stems. The case of complex item stems involving two stimuli description sentences (double-barreled questions) is thereby compared with the setting where items contained a single sentence. Using empirical data, the latent criterion validity differences are evaluated across three groups that are randomly assigned to conditions characterized by item stems with differing number of stimuli. The results indicate that validity of an instrument can be influenced by the specific way item stem is formulated. Implications for empirical educational, behavioral, and social science research are discussed.
Keywords
Measuring instruments consisting of multiple components, such as tests, inventories, testlets, scales, self-reports, questionnaires, surveys, and so on (referred to as “instruments” or “scales” for short below) are very often utilized in the educational, behavioral, and social sciences. Such instruments are highly popular in these and related disciplines in part due to their theoretically and empirically appealing property of yielding converging pieces of information about latent constructs of main interest (e.g., Raykov & Marcoulides, 2011). The validity of these scales may well be considered the bottom line of measurement in those and related sciences (e.g., McDonald, 1999). As discussed in the literature (e.g., Campbell & Fiske, 1959; Messick, 1995), multiple evidences based on substantive interpretation of scale scores are an integral part of conclusions in support of construct validity. Such support can also come from variable correlations when in agreement with predictions based on prior theory and/or previous empirical research. This general approach to validity assessment seems to be broadly accepted in the social sciences. Less commonly used appears to be however the process of identifying sources of invalidity, an aspect that is a primary scale validation concern as well (e.g., Messick, 1995). Invalidity can be closely related to construct-irrelevant variance, and thus the possibility for the latter becoming part of the observed variance on a multicomponent measuring instrument under consideration. In particular, a potential cause of construct-irrelevant variance may be construct-irrelevant difficulty, for instance that resulting from differential item functioning (e.g., Holland & Wainer, 1993; see also Raykov & Marcoulides, 2018). More specifically, complex item formulation may be a primary source of construct-irrelevant difficulty leading to compromised scale validity, which can apply to both construct and criterion validity.
This source of possible instrument validity loss is the main concern of the present article. The article deals with the potential effect that complexity of item formulation could have on scale validity. Complexity means in what follows ambiguous and complex linguistic and grammar structures, and can be described by (a) the use of vague and imprecise terms (e.g., Graesser, 2006; Lenzner et al., 2010), (b) the number of clauses (Yan & Tourangeau 2008), (c) the number of words and syllables (Le Payne, 1951), or (d) more complex grammatical and lexical structures. The special focus of this article is in particular item stem formulation complexity when using more than one object or stimulus attribute that respondents have to deal with in a single item or question in a test, scale, or a measuring instrument more generally. Such items have been referred to as double-barreled questions (DBQs; Le Payne, 1951; Olson, 2008; Oppenheim, 1992; Suchman, 1950). For example, an inventory aimed at measuring the Big Five personality constructs (BFI-S; Hahn et al., 2012), includes items such as “I see myself as someone who is original, comes up with new ideas.” The construct-irrelevant difficulty in items like this may be a consequence of the complication experienced by respondents while trying to relate to the first stimulus “original,” then to the second stimulus “comes up with new ideas,” and finally to provide what is in effect a single response to both stimuli. Although some scholars recommended avoiding DBQs (Le Payne, 1951; Oppenheim, 1992; Schaeffer & Presser, 2003; Suchman, 1950), the latter have been commonly used over the past several decades in measuring instruments employed in the behavioral and social sciences. Another example that is relevant to the study underlying this article is taken from a popular measure of human values, the so-called Portrait Values Questionnaire (PVQ; Schwartz, 2003). In the PVQ, respondents are asked to compare themselves with an imagined person in the context of the question “How much are you like this person?” Respondents are then required to use a set of items composed of two sentences, such as “Preserving the environment is important to them. They strongly believe that people should care about nature.” This item is a particular example of a DBQ, since subjects have to evaluate simultaneously two stimuli contained in it, namely, related to (a) the importance of preserving the environment and (b) the strong belief that people should care about nature.
In this article, we expect that having to evaluate such complex questions, besides the problem of their length, increases the difficulty of the respondents’ task to respond to two stimuli with a single answer. In the following discussion, we examine the potential impact of item wording, and specifically the usage of more than one stimulus in an item, on latent criterion validity. We utilize below a subset of the items of the PVQ as an example, for which we generated items of differing complexity and subsequently examine latent criterion validity. We use the PVQ because of the pronounced complexity of the original items in it, which could be reduced by splitting them into two separate parts. We also decided to use the PVQ due to its relevance as a personality measure in research on individual differences in motivation and behavior. As such, the PVQ has been used by numerous researchers (e.g., Bardi & Schwartz, 2003; Lee et al., 2011; Marcus et al., 2017, to mention but a few), and its items are included in large-scale international and national surveys, such as the European Social Survey (ESS, 2018) or the GESIS-Panel study (GESIS Panel, 2017).
Hence, the focus in the remainder is not on evaluation of the validity of the PVQ, which has been carried out elsewhere (Schwartz, 2003; Schwartz et al., 2015). This article is instead concerned with examining the possibility of validity loss due to construct-irrelevant variance associated with the use of DBQs. To this end, the following discussion is focused on the question how DBQs affect latent criterion validity as reflected in latent correlations between an overall measure and a relevant criterion variable (Raykov et al., 2018). The rest of the article addresses specifically the possibility that respondent answers to a DBQ item could be explained not only by its content but also by the complexity of item formulation, including in particular item verbalization, wording, and presentation. This is a potential source of invalidity, as indicated above, which is due to construct-irrelevant variance stemming from construct-irrelevant item difficulty.
An additional aim of the article is also to discuss a methodology for comparing latent criterion validity coefficients across independent groups. To accomplish this aim, we use the popular latent variable modeling (LVM) approach (B. O. Muthén, 2002). The remainder of the article demonstrates the possible relationship between latent criterion validity and individual item stem formulation, and adds to the literature on the relation between measurement quality and item (survey question) presentation form (Menold & Raykov, 2016).
Double-Barreled Questions: Theory and Past Research
A reason why DBQs can be associated with construct-irrelevant difficulty, as pointed out above, is the fact that study participants engage simultaneously in assessing the given two (or more) stimuli in the item stem but are supposed to provide a single response considering all of them (Menold, 2020). The cognitive response process involved thereby has been described in the extant literature as consisting of four cognitive tasks (e.g., Tourangeau et al., 2000): (a) understanding the meaning of the question, (b) remembering relevant information for responding to the question, (c) using the retrieved information in arriving at an item response, and (d) selecting the response from the provided item response option set. Including double barrels in an item can complicate comprehension, as more information should be assessed by the respondents in order to understand the meaning of the item. In addition, retrieval of relevant information and reasoning on the item when assessing a single response to its two parts is complicated as well, because respondents have to evaluate simultaneously two stimuli, compare them, and assess their judgment on both. A particular problem can emerge if respondents differently assess each of the stimuli, thus, making it more difficult to arrive at a single response on the item under consideration. What could thereby result is a shortcut in the information evaluation process, as respondents might mentally modify the items by ignoring one of the stimuli while responding to the other (e.g., Krosnick, 1991). Respondents might also replace a difficult question with a different, but easier one to respond to (Kahneman, 2012). This increases the risk that respondents consider a content that differs from that stated in a DBQ, a process that would be beyond the researchers’ control and may well entail an increase of construct-irrelevant variance. These cognitive mechanisms may result in validity loss, as the DBQ inventories would not measure the initially anticipated construct.
An argument in favor of the practice of implementing DBQs in behavioral measuring instruments would be the assumption that the used stimuli are very close in their lexical meaning and possess nearly the same substantive content. From the point of view of constructing measuring instruments, the stimuli included in a DBQ could be therefore assumed to be functioning as parallel (e.g., Lord et al., 2008), that is, exchangeable when measuring a concept (construct) of concern. Even if the stimuli are seemingly parallel in their meaning, it is important to realize that respondents have (a) to engage in relevant cognitive activities for understanding each of them, (b) to assess if their meaning is similar, and (c) to arrive at the conclusion of their similarity. As a consequence, respondents could experience a cognitive burden when assessing two or more apparently similar aspects of a considered question at the same time as well.
Although the need for avoiding DBQs has been a repeatedly raised issue for decades, and one can find a substantial body of literature that reiterate this wisdom, many psychological and educational scales have continued to employ DBQs. As an example in point, we would like to mention the Big Five Personality scales by Hahn et al. (2012); other examples can be found in the State-Trait Anxiety Inventory, Trait version by Bieling et al. (1998), or in the modification of the Rosenberg Self-Esteem Scale by Zimprich et al. (2005). In addition, while further issues of question wording, such as negative formulation and reverse-keying, have been empirically investigated (e.g., Cole et al., 2019; Gu et al., 2017; Miller & Cleary, 1993; Pastor et al., 2020), the impact of DBQs on the validity of measurement has rarely been addressed in past empirical research.
Some empirical scale construction studies available in the literature provide support for the idea that removing or revising DBQs can lead to higher measurement quality (see Menold, 2020). A main part of those studies show that the measurement quality of multicomponent measuring instruments that included DBQs and other problematic questions (negatives, double negations, etc.) was improved following such revisions (e.g., J. C. Campbell et al., 2009; Fowler, 1992; Gemenis, 2013; Stafford, 2011). However, the studies implemented many additional revisions beyond dealing with DBQs and their potential adverse effect cannot be separated from that of handling other problems. Indications that DBQs increase the difficulty level of the respondents’ cognitive response process come from qualitative cognitive pretesting studies (e.g., Velloza et al., 2020; Yorkston et al., 2008), but there is a lack of systematic experimental comparison of questions with and without DBQs.
After an extensive literature search (involving use of Google Scholar, PsychInfo, ISI Web of science, and Scopus), the authors of the present article were able to identify two experimental studies that addressed the effect of DBQs. From these two studies, Grant Levy (2019) and Menold (2020) demonstrated that different stimuli in a DBQ did not complement each other, when evaluated separately. Furthermore, Bassili and Scott (1996) showed that in a telephone survey, participants reported more difficulties with the original DBQs than with their revised versions that contained just one stimulus.
To summarize, empirical evidence on the negative impact of DBQs is rather limited. In addition, little is known about how question complexity and particularly the use of DBQs impacts latent variable measurement in terms of its validity. Yet demonstrating an effect of DBQs on validity, which is the goal of this article, holds a strong promise for helping educational, behavioral, and social science researchers to enhance the quality of used measuring instruments through an optimized formulation, verbalization, wording, and presentation of their items or questions. To achieve the aims of this article, we first introduce in the next section a needed method for evaluation of latent criterion validity within the framework of LVM.
Criterion Validity Differences Across Item Stem Presentation Modes
Group differences in measurement quality coefficients such as reliability or validity have been of special interest to methodologists and substantive scholars over the past several decades. As one of the pioneers in this area, Feldt (1969) demonstrated possible discrepancies in consistency of measurement when a multi-item instrument was presented to different populations (groups, subpopulations). More recently, Menold and Raykov (2016) showed that item presentation mode could be associated with group differences in composite reliability. Following that previous research, this article addresses also the question of how to evaluate potential differences in latent criterion validity across groups when using different formulation of item stems. As indicated above, we hypothesize that latent criterion validity may decrease in instruments including more than one stimulus in DBQs, as compared with corresponding simpler (parallel) items that contain a single stimulus and are not double barreled.
To examine this conjecture, we utilize the LVM-based approach for studying latent criterion validity described in Raykov et al. (2018) (see also Raykov & Marcoulides, 2015; Menold & Raykov, 2016). In the unidimensional case assumed here, the underlying single-factor model of relevance is
where y is the p×1 vector of instrument elements (components, questions, items, tasks, or problems; p > 1 in general), η is the construct evaluated by them, Λ is the factor loading matrix, and ε is the p×1 vector of unique factors comprising observed measure specificity and “pure” measurement error that is assumed uncorrelated with the construct (with this model presumed identified, if need be, using additional parameter constraints; e.g., Raykov & Marcoulides, 2011). In the remainder of this article, we consider the case of a single criterion measure, denoted Z, which case is extended straightforwardly along the lines of the following discussion to that of multiple criterion measures that need not be only observed variables. We adopt thereby the definition of the latent criterion validity coefficient (LCVC) of an instrument consisting of the above y measures as the correlation coefficient ν = Corr(Z, η), where Corr(.,.) denotes correlation (Raykov et al., 2018).
In the simplest case of g = 2 groups (and known respondent group membership), the group difference in the criterion validity coefficients, denoted Δν(u, w), is
where the groups are formally indexed by “u” and “w.” This setting can be straightforwardly generalized to the situation with more than two groups, as utilized in the empirical application involving comparison of DBQs with single stimulus versions (SSQs) that is discussed later in the article.
As pointed out earlier, we are interested in the setting where the studied groups result from random assignment. Hence, the variance of the observed criterion, Z, can be assumed to be group-invariant (as group assignment could also be thought of as having occurred after measuring the criterion Z; this assumption is directly testable using an appropriate F-test, say; e.g., Raykov & Marcoulides, 2012). 1 Therefore, the key quantity in Equation (2), Δν(u, w), is re-expressed now as follows:
where Var(.) and Cov(.,.) denote variance and covariance, respectively.
With the preceding discussion in mind, we will be concerned next with point and interval evaluation of the LCVC group difference in Equation (3) in an empirical investigation.
An Empirical Study of Latent Criterion Validity Differences Between Double-Barreled Questions and Questions With Reduced Number of Stimuli
In this section, we investigate the effect of DBQs on latent criterion validity. To be in a position to compare DBQs and SSQs, we employ different versions of an already existing questionnaire. Specifically, we utilize four items of the PVQ (Schwartz, 2003) using the German questionnaire of the European Social Survey 2 (ESS, 2018; Schwartz et al., 2015). The PVQ contains 21 questions that measure 10 different basic human values (mostly by two items per value construct), which are grouped into four second-order values (Schwartz, 2003). We use four items of the PVQ that measure two strongly related basic value constructs, “Tradition” and “Conformity.” In particular, two of the items measure “Tradition” (named Trad1 and Trad2), and another pair of items measure the value construct “Conformity” (Conf1 and Conf2). The full text of the items is presented in Table 1 (with their indicator variable names given within parentheses). The value construct Tradition describes respect for and acceptance of traditional cultural or religious customs, whereas the construct Conformity is defined as “a restraint from actions, inclinations and impulses likely to upset or harm others and violate social expectations or norms” (Schwartz, 2003, p. 268). Both Tradition and Conformity are considered indicative of the second-order value “Conservation,” and are strongly correlated (as evinced by an estimated latent correlation of .92 in the publication by Schwartz et al., 2015). With this in mind, a testable unidimensional structure could be assumed for these four items, even though they represent two distinct basic value constructs. The response options used with the items are also included in Table 1.
Items of the PVQ Instrument (Translated Into English).
Note. Source for translated items: Schwartz (2003). Source for the items used in the experiment: German European Social Survey (2018) Questionnaire. Items were written in third form for the “person” and gender-specific form was not used.
The proportion of respondents selecting the “Do not know” answers was lower than 1% in each of the random groups, and we did not include these individuals in the reported analyses.
As indicated earlier, and seen from Table 1, each item of the PVQ consists of two different sentences and is therefore a DBQ. Two SSQ versions were generated from the initial DBQ version. The first version, referred to as “Stimulus 1” contained only the first sentence of a corresponding DBQ question. The second, “Stimulus 2” version contained the second sentence. For example, consider the original PVQ “Trad1” item “It is important to them to be humble and modest. They try not to draw attention to themselves.” The Stimulus 1 version of this item was “It is important to them to be humble and modest.” The Stimulus 2 version of the item was “They try not to draw attention to themselves.” The same scheme was applied to the other items considered.
We use the items of the PVQ because of their initially complex structure that could be reduced by separating the barrels. For the LVM process, we prefer to select four items (indicators) that represent a latent variable. This is possible to do with the four items evaluating the construct “Conservation.” We avoid here the inclusion of additional items, since this would likely imply more than a single latent variable, while the aim of this study is not to evaluate validity of complex structures or that of the PVQ.
To accomplish our goals in this article, we conducted an experimental study and collected data using three item formulation forms (versions) consisting of (a) the original PVQ, (b) the Stimulus 1 version, and (c) the Stimulus 2 version. Data were collected in 2016 by means of a commercial online access panel of German-speaking adults living in Germany and aged 18 years or older. 3 The quota sample defined by gender, age, and education had the aim of approximately resembling this population. From the n = 435 participants in the experiment, 47.8% were men. With regard to education, 29.4% had a university degree and 30.6% a high school degree. With regards to age, 26.2% were younger than 40 years, 38% were between 40 and 60 years old, and 35.9% were 60 years old or older.
Following random assignment, n1 = 134 respondents answered the original PVQ items, n2 = 161 respondents answered the Stimulus 1 version, and n3 = 137 respondents answered the Stimulus 2 version. None of the experimental groups differed significantly with respect to gender, age, and education (associated p values, p > .10); the pertinent hypotheses were tested with χ2 tests (e.g., Raykov & Marcoulides, 2012). Participants in all groups were offered and used the original instruction of PVQ and original response options, which did not vary across the groups (see Table 1).
As a criterion measure, we used an item from the “Aggression” subdimension of the Short Scale of Authoritarianism (KSA-3) by Beierlein et al. (2014). This item had the following text: “One must counter societal outsiders and nothing-doers with strong actions.” The item was presented with the response categories 1 = do not agree at all, 2 = slightly agree, 3 = moderately agree, 4 = strongly agree, and 5 = fully agree, as provided by Beierlein et al. (2014). Neither the question wording of the criterion item nor its offered response categories varied between the three randomly formed groups. Respondents answered this “Aggression” item first, and only then administered their group’s version of the PVQ items (either original PVQ or Stimulus 1, or Stimulus 2 version); for this reason, the group assignment could not affect the scores on the criterion measure. We expected a latent correlation between the “Aggression” item and the DBQ version not to be lower than .30, which is the size of the manifest correlation between “Aggression” and the subdimensions of “Conservation” reported by Beierlein et al. (2014). We also anticipated this correlation as being higher than .30 in both SSQ versions in the present study, as validity should be expected to increase due to the reduction of complexity and construct-irrelevant variance.
In order to point and interval estimate the possible LCVC differences across these three groups (see Equation 3), which are of special interest in this article and section, we begin by fitting a multigroup model in Equation (1) for the observed measures in the experimental groups; we correlate thereby the underlying construct η with the criterion measure within each of the three groups. (For completeness of this article, we provide in the appendix the Mplus input file used for this purpose, which is a minor modification of the corresponding source code in Raykov et al., 2018; cf. Muthén & Muthén, 2020.) For the particular method illustration purpose in this section, we proceed with the unidimensional model in Equation (1), since this simplest structure is plausible based on substantive considerations (e.g., Schwartz et al., 2015, see also above). Given that each item of the PVQ has six possible response options in each group (see Table 1), we use thereby the robust maximum likelihood method for model fitting and parameter estimation (e.g., Raykov & Marcoulides, 2011). This model is found to be associated with fit indices that are tenable: χ2 = 21.379, degrees of freedom (df) = 17, p = .210, root mean square error of approximation (RMSEA) = .042, with a 90% confidence interval (CI) of [.000, .091], and comparative fit index (CFI) = .982. The associated parameter estimates, with standard errors, are presented in Table 2.
Parameter Estimates, Standard Errors, t Values, p Values, and Confidence Intervals (Further Below) Associated With the Multigroup Model (Software Output Format).
Note. Trad1 to Aggr1—items of used instruments (see main text); DBQ = double-barreled questions; SSQ = single stimulus questions; F1 = latent construct evaluated by instrument; SE = standard error; Est./SE = t value for the test of the null hypothesis of pertinent parameter being 0 in the population; new/additional parameters = latent criterion validity coefficient in each group (LCVC); D1 = difference of LCVC between Original PVQ and Stimulus 1 groups; D2 = difference of LCVC between Original PVQ and Stimulus 2 groups; D3 = difference of LCVC between Stimulus 1 and Stimulus 2 groups.
As seen from Table 2, the LCVC is highest in absolute value in the group with the Stimulus 1 version, being estimated there at −.57; while in the Original PVQ group it is estimated at −.34, in line with the expectation stated before. Further, in the Stimulus 2 group the LCVC approaches zero and has there the estimated value of .03. Thereby, the estimated LCVC difference between the Stimulus 1 and Original PVQ groups is
Based on the preceding discussion, we can interpret the results of this empirical study as an example where differences in item formulation are associated with notable differences in instrument criterion validity, which can also depend on the specific stimuli included in the item stems.
Conclusion
This article was concerned with the possibility that criterion validity of a multicomponent measuring instrument may not only depend on the substantive content of its items but also on item formulation, leading to different cognitive burden levels for respondents when answering pertinent item versions. Specifically, we addressed the question of potentially increased construct-irrelevant variance if more than one object, subject, or attribute is to be evaluated in individual items, as in DBQs. Using an empirical data set collected for the purpose of the study, we examined the discrepancy in latent criterion validity as a function of the number of stimuli included in the items of the instrument. The article adds to the validity-related literature by providing empirical evidence that items consisting of more than one stimulus that respondents have to respond to, may entail decreased instrument criterion validity. The results reported demonstrate that if the number of stimuli included in the stem of a given item(s) is reduced so that the stimulus left in the item stem is construct relevant, criterion validity can be enhanced. Alternatively, if the stimulus left in the item stem(s) is construct irrelevant, a considerable loss in criterion validity may ensue. Our findings can also be interpreted as suggesting that stimuli involved in a DBQ can be of differential relevance with respect to the measured criterion. This relationship can explain their lower criterion validity as compared with SSQs that contain a criterion-relevant stimulus only. Since our empirical example was not based on fairly large samples in the groups used, we would like to point out that further studies are needed before one can place more trust in the above results and their offered interpretation.
The preceding discussion in this article also hints to potentially higher construct-irrelevant difficulty associated with a DBQ. The empirical example used and reported findings demonstrate that the assumption of parallelism of the stimuli included in a DBQ, which authors of inventories are frequently making, may well be questionable, as has been also recently shown by Grant Levy (2019) for a different kind of DBQs. Menold (2020) found lacking measurement invariance between the single stimuli of different DBQ questions. In the specific example of the subscales of PVQ used in the present article, the two modified versions of original questions, were found to be differentially associated with a criterion and therefore of different relevance to it.
The findings of the last section suggest that one should be very cautious when using or revising DBQs. With respect to this activity, an often found recommendation in the literature is to remove one of the stimuli from a DBQ in order to reduce complexity (e.g., Olson, 2008). The empirical example in our illustration section shows however that doing this in a random fashion would be beneficial only if both stimuli, to begin with, are construct- or criterion-relevant to essentially the same degree. If a DBQ involves stimuli with notably different relevance for the criterion of interest or underlying construct, then removal of a randomly selected stimulus would be a risky activity, possibly resulting in single-stimulus item stems and instruments containing such that are characterized by low criterion validity (like the Stimulus 2 version in the illustration section example). The results in the preceding section did not show therefore that avoidance of double stimuli in a question would necessarily result in improved measurement validity. Hence, this article provides explanation of the finding that sometimes validity cannot be increased through decreasing item complexity. When such revisions are undertaken, special attention should thus be paid to the potential relevance of a corresponding item’s part with respect to the construct or criterions of interest.
At the same time, we would like to emphasize that our article did not aim to suggest that in every educational, behavioral, or social science study the validity of an instrument is necessarily reduced by the inclusion of a DBQ, or that a DBQ is always associated with a considerable degree of construct-irrelevant variance. Rather, our goal was merely to show using an empirical example that when groups from the same population are administered scales with varying question wording, the criterion validity of an instrument could be markedly related to the specific manner in which its items are formulated.
This article further suggests that additional attention needs to be routinely paid not only to the issue of how many stimuli are included in an item stem but also to the degree to which each stimulus is relevant for the construct under investigation. That is, what may at first seem like an inconsequential change in item formulation (verbalization, wording, or presentation mode), could in fact have a more profound effect on the quality of measurement of the intended construct to study, and in particular potentially alter its association with a criterion variable. An instrument (item) modification, such as removing part of a question (item stem), could lead to changes in the criterion validity and limit in this sense subsequently the utility of the data collected with the revised instrument. If this modification is to be pursued for substantive reasons, however, new studies need to be first conducted before one could claim in a more trustworthy way what the validity of the revised instrument would be. Using a single rather than multiple stimuli in all items during the item and instrument construction process can help avoid the problem of construct-irrelevant difficulty and differential stimulus relevance for a studied construct. Last but not least, the present article has also exemplified how one can readily examine criterion validity differences (in particular, in latent criterion validity coefficients) across various measuring instrument modifications or revisions in empirical educational, behavioral, and social science research.
Footnotes
Appendix
Authors’ Note
This research was in part carried out while T. Raykov was visiting the GESIS—Leibniz Institute for the Social Sciences, Mannheim, Germany, and the Institute of Sociology of the Technische Universität Dresden, Germany.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
