Abstract
In a recently published article, Van de Calseyde and Efendić (2022) argue that inner-crowd wisdom (i.e., the reduction in error afforded by aggregating two estimates from a given person relative to a single initial estimate from that person) is enhanced when people are instructed to adopt the perspective of someone with whom they disagree prior to making a second estimate. Here, I present a reanalysis of Van de Calseyde and Efendić’s data and argue that evidence supporting their primary claim spuriously arises from anticonservative multilevel models. Specifically, Van de Calseyde and Efendić assess their data via random-intercept models and fail to account for item-level effects of experimental condition. Such an approach generally allows analysts to reap the enhanced statistical power of multilevel models without implementing appropriate checks on that power; in this case, underestimation of item-level variance appears to have driven an illusory benefit of perspective taking.
Van de Calseyde and Efendić (2022) recently published an article in which they argue that taking the perspective of someone with whom one disagrees improves inner-crowd wisdom (Vul & Pashler, 2008), or the reduction in error that is afforded by averaging over multiple guesses from a given individual relative to an initial, single guess. The primary aim of this commentary is to argue that that conclusion is not supported by the data and, rather, is an illusory finding that was yielded by an unfortunately common misuse of multilevel models for hypothesis testing: failing to incorporate a maximal random-effects structure (Barr et al., 2013). This commentary will proceed as follows: First, I will discuss the evidence provided by Van de Calseyde and Efendić; second, I will provide a high-level introduction to multilevel models and the importance of incorporating maximal random-effects structures; third, I will present a reanalysis of the authors’ data with Bayesian multilevel models that incorporate maximal random effects.
Van de Calseyde and Efendić (2022) present data from five experiments. Their experimental paradigm entailed participants estimating a series of numeric quantities twice across two rounds. In Experiment 1a, these quantities were weights, in pounds, of pictured objects (e.g., “What is the weight of an adult male polar bear?”); in Experiments 1b, 2, 3, and 4, these quantities were percentages (e.g., “What percent of the world’s airports are in the United States?”). The key between-participant manipulation was implemented after the first round of estimates was gathered: Before making a second guess, participants in one group were told to make another guess that was different from their first one, a second group was told to make a guess from the perspective of a friend “whose views and opinions are very different from yours,” and a third group (in Experiments 2, 3, and 4) was told to make a guess from the perspective of a friend “whose views and opinions are very similar to yours.” These three conditions will be referred to as the self, disagree, and agree conditions. (Note that participants in Experiment 4 also responded to additional items that were not expected to yield inner-crowd wisdom, but those stimuli will not be discussed.)
The outcome variable of interest was differential mean-squared error (MSE; note that for unaggregated data, which are leveraged for multilevel models, this value reduces to squared error), that is, the difference in MSE between a first guess and the average of the first and second guesses. Larger differential MSE indicates a larger benefit of inner-crowd wisdom. According to Van de Calseyde and Efendić (2022), the data from each experiment support the claim that inner-crowd wisdom is improved by the disagree condition relative to the self condition; additionally, two of the three experiments that featured the agree condition (as well as the aggregated data from those experiments) supported the claim that the disagree condition is more beneficial than the agree condition (see Table 1 for a summary of the reported p values).
Reported p Values and 95% Credible Intervals (CIs) From a Reanalysis of Data From Van de Calseyde and Efendić (2022)
Note: Columns, from left to right, indicate experiment, the conditions being compared, the reported p value in the article, the resulting 95% CI from a random-intercept model in the brms package, and the resulting 95% CI from a maximal-random-effects model in the brms package. Bold text indicates either p < .05 or else a 95% CI outside of zero.
The p value for the self-versus-agree comparison in Experiment 3 is .06 when all the data from that experiment are modeled in the lme4 package.
The authors report only results for the disagree-versus-agree comparison in the original article.
Van de Calseyde and Efendić (2022) analyzed their data with multilevel models (which are also sometimes called mixed-effects or hierarchical models; see Singmann & Kellen, 2019, for a recent overview). Broadly speaking, multilevel models allow analysts to account for effects pertaining to grouping variables in a data set. A grouping variable is any discrete variable that has multiple observations corresponding to each of its levels (e.g., participants and items). Analysts may estimate unique overall means (i.e., random, or group-level, intercepts) or, if applicable, unique effects of a predictor (i.e., random, or group-level, slopes) for the different levels of a grouping variable. If each level of a grouping variable corresponds to only one value of a predictor (e.g., if each participant is exposed to only a single experimental condition), then random intercepts are appropriate; if each level of a grouping variable corresponds to multiple values of a predictor (e.g., if each item is evaluated in each experimental condition), then random slopes are appropriate.
Statement of Relevance
Multilevel models are increasingly relied upon for statistical analysis in experimental psychology. In the context of hypothesis testing, multilevel models are appealing because they allow analysts to control for effects at lower levels of analysis (e.g., participants and items) and therefore draw more generalizable conclusions at the population level. However, this generalizability hinges on a thorough accounting of lower-level effects—failure to fully account for variance at lower levels may lead to false-positive findings at the population level and consequently warp any subsequent theorizing based on those findings. Such nonmaximal random-effects structures are common in experimental psychology; they often arise not because analysts are unaware of the importance of maximal random-effects structure but, rather, because widely used frequentist estimation packages fail to converge when all random effects of interest are included. I demonstrate here the usefulness of Bayesian estimation when maximal models fail to converge with frequentist methods.
In the context of hypothesis testing, multilevel models are appealing because they allow analysts to control for, say, specific stimuli and participants in a particular experiment and draw conclusions that are more likely to generalize to novel stimuli and participants. But this generalization hinges on a thorough reckoning of random effects. If only random intercepts are estimated when random slopes are also appropriate, then that impoverished random-effects structure will raise the probability of a Type I error (Barr et al., 2013; Oberauer, 2022). This underestimation of group-level variance remains a persistent issue in experimental psychology (Yarkoni, 2022) despite calls to better account for such variance first appearing in the literature nearly 60 years ago (Coleman, 1964).
That said, it is important to note that maximal random effects are not a panacea; if random effects are roughly uniform across levels of a grouping variable, then a model with maximal random effects will be overparameterized and consequently render hypothesis tests too conservative (Matuschek et al., 2017; Oberauer, 2022). Accordingly, Oberauer (2022) recommends that nonmaximal models be used for inference only if they are decisively favored over a maximal model in a formal comparison. In the present case, however, effects of experimental condition are rather volatile across items (note the strongly varying distance between conditions across items in Figure S1 in the Supplemental Material), and so random-intercept models do not provide a more parsimonious account of the data relative to maximal models (see Table S1 in the Supplemental Material for model comparisons).
Given that random-intercept models are generally inappropriate when random slopes may also be estimated, Barr and colleagues (2013) and Oberauer (2022) recommend that analysts implement a maximal random-effects structure when estimating multilevel models. However, this recommendation is frequently thwarted by issues with model convergence in widely available statistical software. Fortunately, recent advances in the accessibility of Bayesian estimation have given analysts a user-friendly recourse when convergence fails for frequentist approaches to estimation. I will next present results of a Bayesian reanalysis of the data in question that employs maximal random-effects structures.
Results
Scripts for replicating all original analyses reported next are located at the Open Science Framework at https://osf.io/pjvfw.
Reanalysis of presented models
Van de Calseyde and Efendić (2022) present findings from multilevel models that include condition as a lone predictor and random intercepts for participants and items. Recall that condition was manipulated between participants, and so random participant-level intercepts are appropriate in that case. However, effects of condition are present across items, and so maximal random effects should include item-level random slopes for condition. For these data, models with item-level slopes trigger singularity warnings in the lme4 package (Bates et al., 2015) in R statistical software, which was used by Van de Calseyde and Efendić. However, such models converge without issue in the Bayesian brms package (Bürkner, 2017), which allows users to specify models in lme4 syntax and subsequently estimate parameters via Markov chain Monte Carlo algorithms (most notably, the No-U-Turn Sampler [NUTS]; Hoffman & Gelman, 2014) in Stan probabilistic language (Stan Development Team, 2022). Users may additionally declare custom priors and incorporate other elements into their models that are not possible in lme4 (see the next subsection on the omnibus analysis), but my primary aim is to demonstrate the usefulness of Bayesian estimation even in the absence of those additional specifications. (In a departure from lme4 syntax, I did have to adjust the “iter” argument, for analyses pertaining to Experiment 1a, and “adapt_delta” argument, for analyses pertaining to the remaining experiments, in brms to facilitate model convergence and prevent divergent transitions [Betancourt, 2016] for the models in Table 1.)
Table 1 presents the results of a reanalysis of the Van de Calseyde and Efendić (2022) data with Bayesian models that contain random intercepts or else a maximal random-effects structure (i.e., random intercepts for participants and random intercepts and slopes for items; note that maximal random effects also entail estimates of pairwise correlations between the intercept and slopes corresponding to a given grouping variable). My reanalyses entailed all models that evaluated the key claim that adopting a disagreeing perspective enhances inner-crowd wisdom relative to the self and agree conditions. These models were estimated for each individual experiment as well as an aggregation of data from Experiments 2, 3, and 4 (see footnote 5 in the original article).
The third column of Table 1 contains p values from the random-intercept models that are reported by Van de Calseyde and Efendić (2022). The fourth column displays 95% credible intervals (CIs) from random-intercept models estimated in Stan via the brms package. Note that the results reported by Van de Calseyde and Efendić are taken from models that are fit to subsets of the data that are relevant to each comparison of interest rather than to all the data from a given experiment. In contrast, my reanalyses estimated all the effects of interest simultaneously. (Estimating effects from all the data is the more appropriate choice because the full data set is necessary for estimating item-level intercepts and slopes.) This difference in data sets yields only one inconsistent outcome, which is for the agree versus self condition in Experiment 3. However, a model fit to all the data from Experiment 3 in the lme4 package replicates the pattern of results from my brms model (see notes for Table 1). Thus, the random-intercept models that are estimated in brms all yield the same patterns of findings that are found with lme4.
The key outcomes of interest are displayed in the final column of Table 1, which contains 95% CIs from models with maximal random-effects structures. Again, the difference between these models and the random-intercept models is the addition of a random slope of condition for items. The use of maximal random effects removes every positive result that is reported by Van de Calseyde and Efendić (2022)—that is, for every analysis that found a benefit of the disagree condition relative to either control condition, that benefit is no longer present.
Why does the addition of random slopes so strongly impact the inferences drawn from these models? It is important to note that the uncertainty of the estimates increases substantially in the maximal models (see Tables S2 and S3 in the Supplemental Material for posterior means and standard deviations). This increased uncertainty arises from the fact that the effect of experimental condition at the item level is quite inconsistent, as can be seen in Figure S1 in the Supplemental Material, and the inconsistency of this effect is captured only by the maximal models. This sensitivity of the maximal models to the inconsistency of the experimental manipulation consequently introduces substantially more uncertainty into the population-level estimates relative to those from the random-intercept models.
Omnibus analysis
As a final analysis, I combined the data from all the experiments to evaluate whether the disagree condition might generate superior inner-crowd wisdom relative to the self condition when all the relevant data are analyzed simultaneously. This analysis proceeded differently than for the ones that are summarized in Table 1. First, I used mean-squared proportional error (MSPE) as my outcome of interest to account for the different scales of the items (i.e., deviations between guesses and a true value are divided by the true value prior to squaring; Fiechter & Kornell, 2021; Lorenz et al., 2011; Rauhut & Lorenz, 2011; van Dolder & van den Assem, 2018). Second, I modeled MSPE across participants’ two guesses (see preceding citations as well as Vul & Pashler, 2008) rather than differential MSPE. This model therefore had three predictors: guess, experimental condition, and the guess-by-condition interaction (this last predictor gauged the differential reduction in error over guesses and was of primary interest). The choice to model MSPE across guesses arose from the fact that differential squared error is not described particularly well by any readily implemented linear or generalized linear model. In contrast, MSPE is better described by a hurdle-gamma model (i.e., a gamma model that allows zeros in the response variable; Heilbron, 1994) than is differential MSE by an ordinary linear model (R2 = .29 for my omnibus models versus .04 to .06 for the models in Table 1; R2 is calculated according to Gelman et al., 2019).
Third, in addition to the group-level effects estimated for participants and items, I also estimated (a) random intercepts and slopes for each experiment as well as (b) separate participant-level effects for people assigned to either experimental condition. These choices were both motivated by implementing maximal random effects. Fourth, I evaluated a set of three Gaussian priors (µ = 0; σ = {0.5, 1, 2}) on all population-level effects across three different models. These weakly informative priors allowed me to calculate Bayes factors via Savage-Dickey ratios (Wagenmakers et al., 2010) for each effect of interest. One benefit of Bayes factors over 95% CIs is that they allow analysts to quantify evidence in favor of either the alternative or null hypothesis. I will present Bayes factors as a ratio of evidence in favor of the null relative to the alternative (i.e., BF01). Following recommendations from Gallistel (2009), I evaluated whether BF01 ever suggests support for the alternative hypothesis across these three reasonable priors (see Figure S3 in the Supplemental Material for how these priors correspond to variance in participant means); to consistently find evidence in favor of the null (i.e., BF01 > 1) suggests convincing support for that hypothesis.
All estimates from these hurdle-gamma models are presented in Table S4 in the Supplemental Material. For present purposes, it is sufficient to note that every Bayes factor for the guess-by-condition interaction supports the null hypothesis (from narrowest to widest prior, BF01 = {2.95, 6.16, 11.44}; also see Figure S2 in the Supplemental Material). Thus, after aggregating every relevant data point from Van de Calseyde and Efendić (2022), I am still unable to find any generalizable evidence to support the claim that adopting a disagreeing perspective enhances inner-crowd wisdom relative to making a second guess without perspective taking. In fact, a set of Bayes factors that is calculated across three reasonable priors supports the conclusion that there is no difference between those conditions.
Discussion
The results of my reanalyses and new omnibus analysis are perfectly consistent: The data from Van de Calseyde and Efendić (2022) do not offer any evidence for any generalizable effects of perspective taking on inner-crowd wisdom (item-level effects are touched on in the next paragraph). This pattern of findings is in stark contrast to claims made in the original article. However, those claims were based on multilevel models that used nonmaximal random-effects structures and therefore underestimated item-level variance across experimental conditions; those models consequently yielded misleadingly precise population-level estimates.
Three omnibus hurdle-gamma models suggested that taking a disagreeing perspective is not generally helpful. However, certain items appear to have benefited from perspective taking. Specifically, three items in Experiments 1b through 4 (i.e., Items 12, 15, and 16 in Figure S1) have corresponding 95% CIs outside of zero for the guess-by-condition interaction (that is, the disagree condition generated greater inner-crowd wisdom relative to the self condition) in all three omnibus models. Of the 16 items used across all experiments, those items yielded the fifth, second, and most MSPE following the initial guess (also see Figure S4 in the Supplemental Material). This pattern suggests that taking a disagreeing perspective might be beneficial in scenarios for which people are initially highly inaccurate. Future work should aim to characterize scenarios that generate inaccurate initial assessments (e.g., when participants are novices on a task; Fiechter & Kornell, 2021) and so may subsequently stand to benefit from adopting a disagreeing perspective.
As multilevel models become more prevalent in experimental psychology, it is critical that researchers take steps to appropriately check the statistical power that is afforded by those models. Here, I have shown that user-friendly Bayesian estimation packages provide one way for analysts to implement maximal random-effects structures when convergence fails for frequentist approaches to estimation. Bayesian estimation is an extremely useful tool for drawing generalizable inferences from multilevel models.
Supplemental Material
sj-docx-1-pss-10.1177_09567976241245411 – Supplemental material for Drawing Generalizable Conclusions From Multilevel Models: Commentary on Van de Calseyde and Efendić (2022)
Supplemental material, sj-docx-1-pss-10.1177_09567976241245411 for Drawing Generalizable Conclusions From Multilevel Models: Commentary on Van de Calseyde and Efendić (2022) by Joshua L. Fiechter in Psychological Science
Footnotes
Acknowledgements
J. L. Fiechter is now at the Air Force Research Laboratory. I thank Philippe P. F. M. Van de Calseyde and Emir Efendić for sharing their data, materials, and analysis code, without which this commentary would not be possible. I also thank Dale Barr, Henrik Singmann, and an anonymous reviewer for their helpful comments on previous drafts of the manuscript.
Transparency
Action Editor: Patricia J. Bauer
Editor: Patricia J. Bauer
Author Contributions
J. L. Fiechter is the sole author of this article and is responsible for its content.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
