Abstract
The main advantage of experimental research lies in the possibility of systematically investigating the causal relation between the variables of interest. The well-known advantages result from (a) the possibility to manipulate the independent variable, (b) random assignment of participants to the experimental conditions, and (c) the experimenter’s control over the operationalization of the variables and the general experimental setting. We argue that it is exactly these elements that constitute core advantages of experimental research but that are—at the same time—associated with side effects, which are often out of focus when researchers derive theoretical conclusions from their experimental findings. We discuss potential restrictions linked to these core elements of experimental research. Implications for both theory development and research design are discussed.
When (social) 1 psychologists look for methodological avenues to investigate their research questions, they usually first consider whether experimental approaches are feasible. The main reason for this preference over other methodological approaches is that experimental procedures allow for an investigation of causal relations between the variables of interest. For example, by experimentally manipulating the number of individuals watching a person in need, researchers can assess whether or not the number of bystanders influences helping behavior. Unless the variables of interest cannot be manipulated (e.g., personality traits, socioeconomic background), researchers in social psychology will usually rely on experiments. It is interesting that even when it does not seem possible to manipulate the crucial variable, researchers search for options to experimentally manipulate some proxy of this variable (e.g., by manipulating relative socioeconomic status). A look at social psychology textbooks and journals leaves little doubt that the experimental approach is the “silver bullet” for investigating social psychological research questions—and that social psychologists rely on it whenever possible (for a recent review see, e.g., Wilson, Aronson, & Carlsmith, 2010).
We very much subscribe to the experimental approach and its advantages, and we usually rely on this approach in our own work. However, in this article, we discuss that the experimental approach comes along with side effects that have implications for the theoretical conclusions derived from experimental findings. It is important to note that most of these side effects are linked inherently to the nature of the experimental procedures. For the present purpose, we emphasize three core features of experiments that are displayed in the first column of Table 1.
Overview of Concerns Associated With Experimental Features and Their Implications
First, the experimental approach holds that the assumed causal variable is manipulated to yield at least two different values and that this manipulation is under the experimenter’s control. In other words, the quantitative or qualitative specification of a potential causal variable needs to be changed in at least one experimental condition. Second, participants are randomly assigned to the different experimental conditions, thereby excluding the possibility that participants self-select the condition they are in. Third, experimenters are in control of the experimental setting. On the basis of their theorizing, experimenters select and determine the operationalization of the independent variable as well as how and when the dependent variable is assessed. Moreover, the experimenters try to keep as many situational aspects as possible constant within and across conditions (e.g., time, location, instructions, etc.; see Wilson et al., 2010). On the one hand, these three experimental features, which are manipulation/change of the assumed causal variable, random assignment, and experimenters’ control, have clear, well-known, and highly important benefits. On the other hand, however, we argue that these features may have implications that often are overlooked when researchers relate their experimental findings to the theories from which they originally derived their hypotheses. Our discussion of these possible implications is structured as follows: For each feature, we first outline the potential problem and related side effects (summarized in column 2 of Table 1). We then turn to the implications and discuss advantages and disadvantages of different possibilities to tackle the problem (summarized in column 3). Going through Table 1 row by row, we address the three core features of experimental approach in turn. In doing so, we elaborate most on the first feature (manipulation/change aspect). It is important that we emphasize that the goal of this discussion is not at all to discredit experimental research in social psychology but rather to sensitize researchers to some implicit consequences of our important “silver bullet.”
Experimental Feature 1: Manipulation of Independent Variable
Researchers interested in a particular concept are well aware that the dependent variable of interest can be influenced by numerous variables. For example, individuals’ helping behavior has been demonstrated to be affected by the number of bystanders (Darley & Latané, 1968), whether individuals are thinking of money (Vohs, Mead, & Goode, 2006), or whether they are in a happy or sad mood (Carlson, Charlin, & Miller, 1988). The common approach is to focus on one potential causal factor (or on a rather limited number), for example, on mood to then manipulate this potential causal factor and, finally, to assess whether this manipulation causes differences in the dependent variable while assuring that the causal variable in focus is not confounded with other variables. To manipulate an independent variable, researchers are necessarily required to change the specification of this variable in at least one experimental condition. For example, by experimentally manipulating mood, researchers change participants’ mood either toward the more positive or the more negative side, depending on the experimental condition. When this variation results in differences in the dependent variable, a causal relation is inferred. With respect to the example on the relation between mood and helping behavior, readers of the corresponding discussion section could expect a statement such as “happy mood increases helping behavior.” However, this conclusion may sometimes deserve a second look because it may actually not be warranted on the basis of the empirical evidence at hand.
In the aforementioned example, one might wonder whether it is (a) the absolute level of negative versus positive mood or (b) the change of the mood that drives the observed differences. The discussion section statement “happy mood increases helping behavior” seems to suggest the first alternative, whereas the empirical evidence is inherently linked to the change aspect. To illustrate this point, assume that we had assessed individuals’ mood before and after the mood manipulation. Also assume that before the manipulation, both experimental conditions had a mean mood score of 5 on a 9-point rating scale. Further assume that the mood manipulation was successful by symmetrically improving the score of all participants in the positive condition by 2 points and reducing the score of all participants in the negative condition by 2 points such that the average score of the two experimental conditions differed by 4 points (3 vs. 7) on the mood ratings assessed after the manipulation. Now suppose that we found significant differences in helping behavior between the two experimental conditions. What would be a proper conclusion? Would the observed difference in helping behavior reflect the difference in the absolute average mood scores of the two conditions (3 vs. 7), or would it reflect the type of mood change in the two experimental conditions (−2 vs. 2)? We cannot be sure by solely looking at the experimental findings whether the observed influence of the manipulation is because of higher versus lower levels or because of the change. The discussion section statement “positive mood increases helping behavior” points to the first alternative. However, it is important to realize that the change aspect is inherently linked to the empirical findings that are based on experimental manipulations. 2
It is interesting to note that the change aspect that is an inherent part of our silver bullet often is neglected in the theoretical conclusions derived from experimental research. For example, with respect to our own work, we have investigated how mood influences cognitive processing strategies. In none of the empirical articles have we discussed whether the observed findings are because of differences in average mood levels or changes in mood states. Neither reviewers nor editors have ever raised this issue. With respect to feelings, Schwarz (2012) has explicitly stated in his theorizing that “changes in one’s feelings are more informative than states” (p. 294). It is interesting to note, however, that this important theoretical element of feelings-as-information theory has almost no counterpart on the empirical side. We argue that the absence of a discussion of the level versus change issue is not at all specific to mood and cognition research but rather general. Other fields of research, where the change-level distinction might have important theoretical implications are studies on the influences of power, self-esteem, social status, need for affiliation, stereotype threat, or uncertainty, to name just a few examples. Are differences in perceived power or social status psychologically equivalent to perceived increases or reductions in power or social status? Do increases in self-esteem exert similar or different influences on outgroup derogation as high levels of self-esteem? Although explicit statements of the change issue are sometimes provided, as by Schwarz (2012) on the role of feelings, we question whether this important issue is adequately reflected in theorizing (and empirical work).
In light of the potential confound between change and level as an inherent aspect of experimental research, several issues become worthy of consideration. First, one needs to discuss whether experimental effects that result from the change versus the level component may have different implications with respect to whether adjustment or accumulation are likely consequences of a manipulation. Second, turning from potential consequences in the form of adjustment versus accumulation to assumptions about underlying mechanisms, it might be argued that the change component is more closely linked to salience and accessibility influences than the level aspect and that for this reason it may play a key role for the emergence of experimental effects. In the following, we will first discuss these two aspects and then turn to suggestions on how to address these issues on a theoretical and on an empirical level.
Adjustment and accumulation as consequences of change versus level
It might be argued that the outlined confound between level and change refers to a conceptual subtlety that plays only a minor role when it comes to actually understanding the psychological function of the variable in question. In contrary, however, we argue that identifying the source of an experimental effect (either the change or the level component) holds important implications, perhaps most directly for the consequences of the variable in question over time.
Research has demonstrated that individuals are especially vigilant for changes (e.g., Olson & Janes, 2002), and in line with this vigilance, changes—even rather subtle ones—often have a particular influence. It is important to note that the influence of such changes often disappears after some time, even if the level differences (compared to the starting point) persist. For example, research by Brickman, Coates, and Janoff-Bulman (1978) has suggested that changes in individuals’ life circumstances influence their life satisfaction, but that individuals adjust to new situations so that the initial consequences related to the change are attenuated (for related evidence, see Lucas, Clark, Georgellis, & Diener, 2003). Presumably, individuals often adjust to changes in the environment (e.g., Frederick & Loewenstein, 1999; Lyubomirsky, 2011)—even if the “manipulation” is still effective. Perhaps most prominently, this notion is captured in prospect theory (Kahneman & Tversky, 1979), which holds that utilities are perceived and evaluated relative to the status quo. Given that the status quo is adjusted to after a change, changes have a particularly strong effect relative to the influences of the absolute level. Such an adaptation and adjustment perspective suggests that effects that are observed because of experimentally induced change may often be rather limited in their duration. Looking at experimental procedures, it can be noticed easily that in their empirical work researchers seem to be aware of the fading effect of experimental manipulations because, in many cases, the dependent variable is assessed within a very short delay of the experimental manipulation. Because of this immediate assessment, adjustment processes are unlikely to have taken effect.
It is interesting to note that one may just as well present opposing arguments. The effect of a particular variable may not attenuate but increase over time, if the variable in question is continuously present. For example, (a) individuals will suffer more from stress if stress is ongoing for an extended time; (b) the immediate effect of mild unpleasant states is usually weaker than the effect of intense unpleasant states but can last longer and be stronger in the long run (Gilbert, Lieberman, Morewedge, & Wilson, 2004); or (c) repeated exposures to a stereotype should increase, rather than decrease, corresponding stereotype threat effects in performance situations.
Given that for many domains, experimental evidence pertaining to such repeated or enduring manipulations and delayed measurements is rare, our knowledge on such potential effects is rather limited (we readily admit that this argument may not hold for all domains). Because change is inherent to the experimental approach, and because both adjustment and accumulation is possible, theorizing on whether a particular variable of interest falls in the adjustment or in the accumulation category seems necessary. In a related vein, Van Lange (2013, p. 46) recently has pointed out the often rather “brief time horizon” of social psychological research, and he argued that because of an emphasis on these “short term influences” (p. 46), there is often rather little theoretical advancement in this respect. Presumably, this short-time horizon is to a substantial degree because of researchers’ relying on experimental approaches. In line with Van Lange’s observation, we suggest that it would promote theoretical and subsequent empirical work if researchers would at least discuss this issue—even if their experimental data did not allow for direct conclusions to this question.
Salience and accessibility as consequences of change
The different consequences of the change versus the level component discussed earlier are presumably to a substantial degree a result of the different relation of the two components to the notion of salience and accessibility. By changing a psychological variable, the salience and the accessibility of that variable are likely to increase, because “changes and differences are much more accessible than absolute levels of stimulation” (Kahneman, 2002, p. 456). For example, by manipulating individuals’ mood, the salience of this variable is presumably increased. Given that individuals focus on information that is most accessible or salient, a variable that has been changed is more likely to be influential relative to the situation when no recent change of this variable has occurred. Thus, we cannot really tease apart whether a particular effect of some variable (here mood) is solely due to the induced relative differences between experimental conditions or whether increased accessibility is a necessary ingredient for the effect to emerge. It might be argued, “Isn’t that what the underlying theory says?” Indeed, sometimes the underlying theories explicitly incorporate and address the notion of accessibility. We argue, however, that in many cases, if not most, an explicit discussion is missing. Put differently, our theories often fail to address explicitly a mechanism that might contribute substantially to the observed effects. With respect to the role of accessibility in experimental paradigms, at least three questions seem worth being considered: (a) How do checks of the effectiveness of an experimental manipulation influence accessibility? (b) Does increased accessibility always lead to stronger influences? and (c) How is the salience of a change related to the base level of the independent variable? We address these three questions before turning to possible solutions of the change versus level confound.
Manipulation checks and accessibility
An increased accessibility of some variable may not only stem from the change that is related to the experimental manipulation but also from the manipulation check that usually is applied in experimental settings. In most cases, assessing a variable increases its accessibility. This mechanism holds for both the experimental and for the control group, and its influence is particularly strong because the manipulation check is frequently applied immediately before the dependent variable is assessed. Thus, the manipulation check often required by reviewers and editors may influence the effect of the variable in question (see, e.g., Kassam & Mendes, 2013; Kühnen, 2010; Sigall & Mills, 1998). It seems worthwhile to consider not having a manipulation check before the assessment of the dependent variable—at least when the effectiveness of a particular manipulation in a series of studies has already been demonstrated. 3 It is interesting to note that responses to the manipulation check often are entered into mediation analyses. The current perspective points to the potential risk that the independent variable has an effect on the dependent variable because its accessibility was increased by the manipulation check, and without this assessment, the effect of the independent variable would be less pronounced (see also Fiedler, Schott, & Meiser, 2011).
Increased accessibility and the effectiveness of an experimental manipulation
Whereas the combination of change and presence of manipulation check may increase accessibility and, in turn, the strength of an effect, one could also argue that too much accessibility might reduce potential effects. For example, research on priming has demonstrated that priming a particular concept too blatantly may eliminate, or even reverse, the effects that are observed with more subtle forms of priming (Lombardi, Higgins, & Bargh, 1987; Martin, 1986; Strack, Schwarz, Bless, Kübler, & Wänke, 1993). Accessibility may contribute to individuals perceiving unwanted contaminations of their judgments, and this can elicit correction processes that counteract the initial effect (for a discussion, see Bless & Schwarz, 2010). Thus, stronger manipulations, going along with more pronounced changes and an increased salience of the variable of interest, may at least under certain conditions have weaker effects than more subtle manipulations. From this perspective, subtle manipulations, which are often highly evaluated because they are viewed as stronger tests of the underlying theory, may sometimes be more likely to affect the dependent variable than more blatant manipulations. These considerations need to be qualified by the notion that it is presumably not the increased accessibility per se that drives these effects but rather the salience of the experimental procedure that renders the influence of the variable as unwanted and elicits correction processes (see Greifeneder, Bless, & Pham, 2011, for a discussion of this issue with respect to mood).
The relation between change and baseline
When discussing the salience and the consequences of change, we may speculate whether the Weber–Fechner law applies to experimental research. According to the Weber–Fechner law, perceptions of differences between stimuli (i.e., perceptions of change) depend on the magnitude of the stimulus. A change on a given variable is more likely to be perceived when the baseline level of the variable is low rather than high. For example, when holding 1 kg, the addition of 100 g is more likely to be detected than when already holding 10 kg. If the Weber–Fechner law applies to experimental manipulations, then it ought to be expected that the consequences of a particular manipulation do not only depend on the strength of the manipulation but also on the starting level of the variable of interest. Small changes would be more likely to show an effect if the a priori level is low rather than high. For example, giving some additional power to individuals who usually have little power would have more psychological consequences compared with giving some additional power to individuals who usually have higher power (note that the opposite hypothesis could be derived for other reasons, which nicely illustrates the potential and the importance of such considerations). Similarly, improving the affective state of individuals who are already quite happy to begin with should have less consequences than the same improvement in individuals who are rather unhappy to begin with. Again, theorizing on this topic, which is related closely to the experimental approach, is rather rare (see e.g., Maglio, Trope, & Liberman, 2013, for an exception). Presumably, addressing such aspects would contribute to a better understanding of the investigated variable.
Theoretical and empirical implications of the change-level confound
The considerations on the potential confound between the presence of change and of differences in absolute levels are linked to several implications on both the theoretical and the empirical level. With respect to theorizing, we believe that much benefit results from explicitly addressing the change issue. The underlying theories could outline whether the predicted consequences of a variable are because of (a) the absolute level, (b) the change, or (c) both components. This discussion would require theoretical arguments about whether the effect of induced change in a certain variable declines after some time 4 or the induced change is likely to have an enduring effect that might even increase over time. Such theorizing would, in turn, promote the differentiation between variables that are likely to elicit adjustment processes and those that will result in accumulative effects. In addition to contributing to refined and more precise theorizing, addressing these issues would also be very helpful for transferring obtained experimental findings to applied settings. We argue that, in many cases, such transfers essentially require assumptions about whether effects are contingent on changes in accessibility and about whether adjustment or accumulation processes are likely to emerge. If so, the remarkable silence of many of our theories on these issues constitutes a potential obstacle for deriving theoretically founded applications from experimental social psychology findings (see also Van Lange, 2013, on the relation between applicability and theory evaluation). For example, individuals have been demonstrated to eat less unhealthy food from a red rather than a blue plate (Genschow, Reutner, & Wänke, 2012; Reutner, Genschow, & Wänke, 2015). Assuming that food is not served on red plates in most situations, the experimental manipulation can be considered as a change. Thus, it is unclear whether the robust experimental effect (see Bruno, Martani, Corsini, & Oleari, 2013) would still be observed if individuals persistently ate their food from red (blue) plates and such an approach would contribute to a successful diet. In fact, Wänke and colleagues themselves (Wänke, personal communication, July 2, 2015) would agree that it remains unclear whether or not red plates are likely to have an enduring effect on individuals’ tendency to eat healthier food.
In addition to advocating for an explicit discussion of the change/accessibility aspect in theorizing, we propose several suggestions that would address—though not solve—parts of the problem empirically. In turn, we discuss (a) prior assessment of the variable of interest, (b) the potential of correlational evidence, and (c) repeated within-subjects manipulations.
Prior assessment of the variable of interest
One possible empirical approach rests on assessing the crucial variable before the experimental manipulation. For example, one could assess mood prior to a subsequent mood manipulation. 5 Then researchers could not only test for the influence of the experimental manipulation but also try to disentangle the level and the change aspect. The difference (change) between the pre- and postmanipulation measure could be computed, and how this difference is related to the dependent variable could be tested. Is there a relation? Does this relation hold if the mean-level differences induced by the experimental conditions are controlled for? How does the size of this relation compare with the effect of the experimental manipulations? We readily admit that such an approach may come along with its own problems regarding interpretation. However, we also believe that it may provide some useful clues to the relation between the change and the absolute-level aspect. In particular, when combined with the usual procedure (no prior assessment), such evidence would substantially advance theorizing.
Considering correlational evidence
Another possibility to address the outlined issue rests on the consideration of correlational evidence. Admittedly, correlational evidence may entail many interpretation problems, as discussed in almost every social psychology textbook. Yet, correlational evidence could bypass the change issue because the variables of interest are simply assessed (and not manipulated or changed). Moreover, the accessibility aspect could be addressed by assessing the dependent variable before the independent variable or preferably in both potential orders (potential order effects would further inform about the role of accessibility). Such complementary correlational evidence could provide strong empirical support and allow for additional, more refined theorizing. Note that such a perspective departs from the often perceived tendency to ascribe a minor and inferior role to correlational evidence (see, e.g., Wilson et al., 2010).
Repeated manipulations/within-subjects designs
A further possibility to address potential adjustment and accumulation processes lies in repeated manipulations or repeated measurements. Perhaps most directly, adjustment and accumulation processes can be tested by examining the potential effects over time. For this purpose, repeated measurements of the dependent variable over time are required. Time series analyses (for an overview, see Velicer & Fava, 2003) then allow for an understanding of causality and the development of effects over time. An alternative to solely assessing the dependent variable repeatedly rests on repeated manipulations. When the same participants are exposed repeatedly to the same manipulation, one can test for the persistence, incline, or decline of the initially observed effect.
In this respect, one can compare situations in which participants are repeatedly exposed to the same treatment with situations in which participants are exposed to different levels of the independent variable (for a discussion of statistical, methodological, and theoretical aspects of within- versus between-subjects design, see Greenwald, 1976; Keren, 1993). On the one hand, the repetition of a manipulation is likely to allow for adjustment processes. On the other hand, repeated employment of different levels of the independent variable is likely to increase the salience/accessibility aspect. We argue that, if possible, it is worthwhile to compare empirically the influence of the two variants because this comparison, in turn, would allow for theoretical conclusions.
For example, in research on how perceptual fluency influences judgments of truth, individuals are repeatedly presented with statements that are either high or low in fluency (e.g., Dechêne, Stahl, Hansen, & Wänke, 2010; Hansen, Dechêne, & Wänke, 2008; Reber & Schwarz, 1999). Wänke and colleagues (for a review, see Wänke & Hansen, 2015) investigated the influence of fluency on truth judgments when fluency was manipulated between subjects (constant fluency across all trials) relative to when fluency was manipulated within subjects (fluency varied across trials). In keeping with the aforementioned line of reasoning, the within-subjects design relatively increases the salience of fluency, whereas the between-subjects design is more likely to allow for adjustment processes as the number of repetitions increases. Hansen and colleagues (2008) observed the well-documented effect of fluency on judgments of truth in the within-subjects design. However, when fluency was manipulated between participants, smaller or no effects of fluency on judgments of truth emerged. Wänke and Hansen (2015) argued that “changes in an internal state are often more noticeable and perhaps more relevant as a diagnostic cue regarding the environment than the absolute level of that state” (p. 195; see also Shen, Jiang, & Advaval, 2010). On the basis of this conclusion, Wänke and Hansen, in turn, provided a refined theorizing on how fluency influences judgments and decisions, which demonstrates very well that much can be learned from comparing repeated manipulations in between- versus within-subjects designs (for additional findings with respect to between- versus within-subjects designs, see Crawford, Kay, & Duke, 2015; or Hsee, Blount, Loewenstein, & Bazerman, 1999, for a conceptual framework on joint versus separate evaluations).
The within- versus between-subjects designs comparison is indirectly related to the possibility that perceived change may have different sources. In this respect, Wänke and Hansen (2015) have argued that change may result from prior states (e.g., the last item was easy to read) but also from what is usually expected (e.g., individuals expect that advertisement statements are easy to read). The latter issue highlights the notion that to understand and interpret participants’ behavior in experiments, researchers should take into account how their experimental manipulations differ from what is expected. In many cases, however, little is known unfortunately about how the variable in question is distributed outside the lab. Are our participants usually in a happy or a sad mood, are participants experiencing power in their everyday life, is food usually served on red plates, are we more likely to be confronted with loss-frame situations than we are with gain-frame situations, and so forth? Depending on participants’ implicit or explicit expectancies derived from their knowledge outside the lab or from their naïve theories about what happens in psychological experiments, the very same experimental manipulations may exert quite different effects.
We readily admit that the change/adjustment issue may not pertain to all experimental manipulations to the same degree. We propose that the change issue is particularly relevant when experimental manipulations are, in a broader sense, linked to the activation of concepts (e.g., when activating social norms, stereotypes, aggression concepts, episodic memories, etc.) or when experimental manipulations are designed to change psychological experiences (e.g., mood, fluency, uncertainty, regulatory focus, etc.). The issue may be less relevant when environmental features that individuals frequently encounter in considerable variance are manipulated. For example, when manipulating the physical appearance of target individuals and subsequently assessing perceived communal and agentic traits, it seems that the change aspect may be less relevant (at least in between-subjects designs).
The discussion in this section suggests that attending to the change/adjustment issue is possible on both a theoretical and an empirical level. It seems, however, as if this issue is frequently underestimated or neglected and that more attention would be very beneficial for theorizing in social psychology as well as other fields of psychological research.
Experimental Feature 2: Random Assignment
Random assignment of participants to experimental conditions is a core element of experimental procedures. The random assignment (if successful) ensures that the experimental conditions differ only with respect to the manipulation and that the participants assigned to the different experimental conditions do not systematically differ on other variables. Experimenters can then assess the situational causation by comparing the experimental conditions while potential dispositional influences are controlled because they are distributed evenly across conditions (i.e., the dispositional effects provide the basis for computing the error term). In other words, random assignment prevents that differences between the experimental conditions on the dependent variable can be attributed to variables other than the manipulated independent variables.
Though researchers rely on random assignment to eliminate unwanted person effects, they at the same time are aware that individuals’ behavior is a function of the person and the situation (Lewin, 1935, 1951). Over the years, excellent discussions on the Person × Situation interaction have been presented (e.g., Bowers, 1973; Buss, 1979; Funder, 2008; Kihlstrom, 2013) that discuss the shortcomings of purely situational or purely dispositional explanations. Though the exact nature of this interaction (e.g., dynamic vs. static) may differ depending on the underlying model, it is obvious that methodological approaches that eliminate or ignore one of the two components are unable to capture their combined interactive influences. On the basis of this general assumption, numerous social psychology experiments incorporated person factors and revealed that individuals’ dispositions moderate the size or even the direction of almost any situational influence.
Independent of the Person × Situation debate, it seems worthwhile to speculate about the psychological consequences of random assignment. To do so, we look at why individuals might face particular situations. At least three different reasons can be distinguished. First, individuals may end up in a situation accidently (perhaps even against their own intention)—a situation that matches the random assignment procedure. For example, in experiments on the consequences of power, individuals are randomly assigned to a situation in which they do or do not have power. Second, individuals may have different options and may be able to choose which situation they want or do not want to be in. For example, if offered the choice, some individuals may opt for the power situation and other individuals for the nonpower situation. Third, in many everyday situations, individuals are neither assigned to a situation nor can they simply choose to be in a particular situation. Often individuals need to strive to end up in a situation and have to engage in activities that lead to this situation. Experimental research is usually remarkably silent on how random assignment (i.e., compared with active striving) may influence the interpretations of the obtained findings. This is interesting because being in a situation may sometimes have opposite effects to striving to be in a situation. For example, experimental research on the consequences of power suggests that high power leads to more positive affect than low power when individuals are randomly assigned to high- or low-power situations (Keltner, Gruenfeld, & Anderson, 2003; Kifer, Heller, Perunovic, & Galinsky, 2013). However, by random assignment, individuals are granted power that is independent of their striving for power. Other research, however, has suggested that striving for power is associated with negative affect (e.g., Emmons, 1991). If we combine the two sets of findings, it seems as if being there versus getting there might make a big difference.
Theoretical and empirical implications
Many of our theories are again silent on whether a particular variable exerts its influence independent of whether individuals have self-selected the respective situation or not. Given that most evidence is based on experimental research with its random assignment component, the implicit assumptions seems to be that random assignment (vs. self-selection) and the corresponding psychological consequences do not influence the effects of a particular variable. Although we readily agree that there are many situations for which the distinction between random assignment versus self-selection does not matter, we speculate that there are also many situations for which this distinction may become important. If so, it seems fruitful to address this issue in theorizing.
Let us assume that the theoretical analyses would suggest that for a given variable of interest, random assignment versus self-selection could play a crucial role. If so, it would be interesting to compare findings of the two settings (admittedly, there are settings that hardly allow for self-selection because the respective “experimental” conditions may seem so unattractive that participants would not choose them voluntarily). For example, Gaines and Kuklinski (2011) elaborated on the challenges and possibilities of such an approach. In essence, they suggested to randomly assign participants either to the random assignment or to self-selection condition. Within the random assignment conditions, the experimenter randomly assigns participants to one of the experimental conditions. Within the self-selection condition, participants can choose which condition they prefer. Gaines and Kuklinski applied their approach to the question how negative political campaigning influences the evaluation of politicians. It is interesting that negative campaigning exerted quite different effects relative to a control group, depending on whether participants were assigned to negative campaigning or they themselves chose to be informed about the content of negative campaigning. We argue that results from such comparisons would strongly contribute to the application of social psychology to situations in which the ingredients of experimental settings (here random assignment) are not given. In other words, not relying on purely experimental procedures may sometimes advance our knowledge and theorizing better than could be achieved by adhering to the experimental approach the entire way through.
Experimental Feature 3: Researcher’s Control Over the Experimental Setting
A further core element of experimental research is that researchers have complete control over the experimental setting. Researchers select how to manipulate independent variables and how to assess dependent variables. They determine the delay between the experimental manipulation and the assessment of potential effects, and they choose a cover story that holds everything together. Moreover, they can eliminate, or at least reduce, the influence of variables that might otherwise co-occur with the influences of the independent variable.
Ideally, the selection of the operationalization is driven by how representative the respective manipulations and assessments are for the underlying theories. However, it may be speculated that in practice, there are systematic biases in the sampling of the general setting and the operationalizations (see Fiedler, 2011). In fact, this problem has been identified repeatedly over the last few decades. McGuire (1973) described how researchers who are convinced of their theories may go through several rounds of optimizing their operationalizations (setting, independent variable, dependent variable, etc.) until the expected effects are finally observed (for a discussion, see also Greenwald, 1975). More recently, Fiedler (2011) outlined that this selection or sampling bias pertains to several aspects, such as design of the study, variables and measures, and performed analyses. If experimental operationalizations often represent a biased sample of the underlying theories, then it becomes a crucial issue as to how to deal with biased sampling, both on a theoretical and an empirical level.
The biased sampling becomes evident, when we take a look at prototypic aspects of the experimental setting. For example, it seems as though researchers prefer novel judgmental targets with a substantial malleability and as though they frequently rely on hypothetical scenarios and judgments or decisions that are linked to little or no personal consequences. As described in the section on change, there is usually a rather short delay between the manipulation and the assessment of its effect. Frequently, the manipulation check precedes the dependent variable, thus, increasing the accessibility of the manipulated construct. In many settings, information about social situations is provided semantically, and, although dealing with social situations, no real other person is present. Moreover, by signing up for an (laboratory) experiment, participants usually agree to cooperate with the experimenter for a fixed amount of time (with cooperate, we explicitly do not refer to demand effects; for a discussion see Bless, Strack, & Schwarz, 1993). In turn, experimental settings often reduce or eliminate otherwise active motivations and goals. Thus, the independent variable can exert its influence in a sort of motivational vacuum. Admittedly, there are many studies that differ from the above descriptions; for example, studies in which individuals interact with other persons or studies in which individuals perceive personal consequences. However, we propose that in general, the above characteristics are quite frequent and dominant elements of experimental social psychological research.
Theoretical and empirical implications
It is quite obvious that many of the above aspects potentially contribute to the effect of the independent variable. Again, we do not at all argue that the experimental approach leads to false interpretations. However, we propose that in subscribing to the power of experiments, we often overlook the implications of our biased sampling for the conclusions at the general, theoretical level. In fact, the rejection of the null hypothesis usually leads to inductive inferences that “X influences Y” (for a general discussion of the logic of the null hypothesis testing, see Krueger, 2001). Frequently, potential moderating variables are addressed, but these moderators are rarely linked to the general problem of how the experimenter selected and created the experimental setting—in other words, a discussion of the selective sampling of the design and the operationalizations is often missing.
By not making the biased sampling explicit, we may fail to see limitations and boundary conditions of the underlying theory. Addressing this issue would increase the precision of the underlying theory—a key element for theory evaluation. Admittedly, specifying such aspects would on the one hand reduce parsimony and generality (for a discussion of these issues, see Gawronski & Bodenhausen, 2015). Moreover, addressing potential limitations and boundary conditions for every experiment would be a cumbersome endeavor, and we do not advocate an extensive discussion of the sampling of the design and the operationalizations for each and every experiment. However, in general, more attention should be allocated to this aspect, in particular when theories are presented in a review format (e.g., in series such as Advances in Experimental Social Psychology, Psychological Review, European Review of Social Psychology, or in other journals that provide the platform for review papers). We argue that regular readers of such review series find such theoretical considerations less often than would be desirable.
Empirically, it seems straightforward to broaden the empirical support by applying diverse methods, designs, and operationalizations that tackle the same research questions (see Wells & Windschitl, 1999; Westfall, Judd, & Kenny, 2015, for discussions of the importance of stimulus sampling). Given the above perspective, it becomes obvious that in the long run, conceptual replications can provide very fruitful answers because they address the question of whether the initially observed effects are potentially caused by some perhaps unknown aspects of the experimental procedure (for a discussion of conceptual versus direct replications, see e.g., Stroebe & Strack, 2014; see also Brandt et al., 2014; Cesario, 2014; Lykken, 1968; Schwarz & Strack, 2014). Whereas conceptual replications are adequate solutions for broadening the sample of situations (for examples, see Stroebe & Strack, 2014), the present perspective, in addition, emphasizes that it is important that the different conceptual replications do not share too much overlap in general aspects of the experiment (see also Schwartz, 2015, advocating for conceptual replications). Taking a somewhat closer look at the experimental research reported in the leading journals of our field, it seems as if some aspects are rather constant across many experimental settings, even if they fall into the category of conceptual replications.
Note that acknowledging the experimenter’s control over the materials, the setting and the sampling of participants may also have implications for the interpretation of effect sizes. Effects size can be increased, for example, by using a particularly strong manipulation. For instance, in persuasion research, the effect of argument quality is a direct function of which arguments are selected. Moreover, effect size will be increased the more control the experimenter exerts, such as by assuring that there are no fluctuations in noise from outside or by assuring that the sample is homogeneous rather than heterogeneous (e.g., in age, in a similar background, etc.). Thus, effect sizes are also a function of how experimenters sample materials, settings, and participants. It seems that when comparing effect sizes across studies, the main interpretation usually rests on conceptual differences between different variables of interest, whereas the selection processes in operationalizations are less frequently discussed. Moreover, for aforementioned reasons discussed, effect sizes obtained in laboratory experiments are often not informative about the strength of effects in applied settings. In most cases, laboratory experiments are designed to provide evidence for the existence of a causal relation between variables, but they do not necessarily provide a good sense of the strength of the relationships out in the world.
Again, we do not at all argue that the experimental approach leads to false interpretations. However, we propose that in subscribing to the power of experiments, we often overlook the implications of our biased sampling for the conclusions at the general, theoretical level.
General Conclusions
In this article, we have discussed manipulation, random assignment, and experimenters’ control over the experimental setting as core elements of the experimental approach. We have argued that these elements that contribute to the numerous advantages of experimental procedures are potentially linked to side effects that are often overlooked. It is important that our discussion focused on side effects that follow from essential elements of the experimental procedure as such instead of side effects that might emerge from contextual details of the concrete proceedings of running experiments (e.g., demand effects, experimenter expectancy effects) that have been discussed elsewhere in the literature (see Nichols & Edlund, 2015, for an overview). In this respect, Klein et al. (2012), for example, emphasize the importance of considering the social context of experiments when interpreting experimental findings. In our concluding discussion, we broaden the scope and discuss the implications of the points raised by us in this article for theorizing, empirical research, and structural aspects of the publication system.
There is no question that any empirical research has limitations and constraints. Any study is conducted under constraints because it is based on selected materials and selected participants assessed at a selected time. One of our main concerns is that these constraints are rarely considered when researchers take the step back from their obtained findings to theoretical conclusions. In part, because of the inductive approach, it is assumed that constraints in how the data were obtained only require discussion if there are plausible assumptions that a particular constraint may influence the obtained findings. For example, most laboratory studies do not take place on Sundays; however, it seems rather unlikely that the day of the week influences how affective states regulate cognitive processes. We propose that unlike the day of the week example, core elements of the experimental approach may often require some additional discussion. Is change an important or necessary ingredient for observing the effect? Does the effect require an increased accessibility of the underlying construct? Is the effect likely to be observed when participants can self-select the respective situation? How long would the effects of the experimental manipulations last? Should we expect accumulation or adjustment in cases of repeated or endured confrontation with the manipulation? Given that these issues are closely, and in parts inevitably, linked to the experimental approach, and given the dominance of the experimental approach in many research domains, the myopia of our theories on these questions is rather surprising. We suggest that addressing these questions on a theoretical level would promote a better understanding of how experimental data relate to the underlying theories. Making implicit assumptions explicit and testable, in turn, leads to refinements and sharpening of theories—which Van Lange (2013) has discussed as a key criterion for the evaluation of psychological theories.
In our discussion, we focused on the relation between the experiment and the underlying theory, which is usually referred to as internal validity. Nevertheless, it is obvious that the implications of our analyses also pertain to external validity and to how findings can be generalized and applied to settings outside the laboratory. Potential applications are confronted by circumstances where our theories are remarkably silent about the constraints of their empirical support. Consequently, possible applications need to solve theoretical questions. For example, applications need to address and solve the (theoretical) question of whether a particular effect is likely to attenuate or accumulate when the same manipulation is repeated within the same participants. Whereas application is not the core focus of the current discussion, we strongly believe that improving our theories by incorporating limiting (or nonlimiting) conditions would contribute to better applications of social psychological research (see also Van Lange, 2013, on applicability as a benchmark for theory evaluation).
So do we not know this already? Do we not know that we are testing our theories under constrained conditions? Yes, of course we know this, but we often fail to explicitly state how these constraints are linked to our theorizing. The present perspective holds that there is rather little discussion of such aspects in empirical research articles and little systematic discussion in review papers. Moreover, a systematic debate of these issues is not available in presentations of the experimental method (e.g., Wilson et al., 2010). It is important to note that in the present perspective we do not suggest dropping the silver bullet (i.e., the experimental approach), but instead we emphasize the need for more sensitivity to potential side effects of the dominant research approach.
Our discussion has several implications for empirical research that we have outlined. In general, it seems worthwhile to test empirically whether the constraints of the experimental approach influence the obtained findings. Conceptual replications are central and important for demonstrating that obtained research findings are not because of unwanted or intentional selective sampling of the experimental material and settings (see Fiedler, 2011; Greenwald, 1975; Schwartz, 2015). Note that conceptual replications are meaningful, independent of whether or not the obtained findings support or contradict the initial effect. Whereas successful conceptual replications increase the trust in the underlying theory, nonsuccessful conceptual replications point to variables that moderate the hypothesized effect—which, in turn, can lead to a refined and improved underlying theory. Given the importance of conceptual replications, our field requires a good balance between introducing new theories and testing the limitations and potentials of older theories (for related arguments, see also Fiedler, 2004; Kruglanski, 2004).
The above analyses suggest that it would be helpful if some conceptual replications were conducted outside purely experimental settings. Ironically, evidence that is usually considered as inferior may solve some of the problems. For example, at first glance, it may seem better to have three experimental studies rather than two experimental and one correlational demonstration of the hypothesized effect. Our discussion suggests, however, that the correlational studies can highlight issues that cannot be addressed within an experimental approach.
Currently, our field seems plight with many things, such as potential fraud, direct replicability, or p hacking. All of these issues share a similar concern regarding whether or not reported findings are conclusive with respect to the underlying theories. In light of the prominent current discussion on questionable research practices in our field, the present perspective holds that it may be also very important to take a closer look at the interpretation problems that arise from adequately performed experiments.
Footnotes
Acknowledgements
We thank Michaela Wänke, Rainer Greifeneder, Klaus Fiedler, Jennifer Eck, and three anonymous reviewers for helpful comments on earlier drafts of this article.
Declaration of Conflicting Interests
The authors declared that they had no conflicts of interest with respect to their authorship or the publication of this article.
Funding
This research was supported by a grant from the Deutsche Forschungsgemeinschaft to the first author (BL 289/16-1 and 16-2).
