Abstract
Laboratory experiments have a long history within sociology, with their ability to test causality and their utility for directly observing behavior providing key advantages. One influential social psychological field, status characteristics and expectation states theory, has almost exclusively used laboratory experiments to test the theory. Unfortunately, laboratory experiments are resource intensive, requiring a research pool, laboratory space, and considerable amounts of time. For these and other reasons, social scientists are increasingly exploring the possibility of moving experiments from the lab to an online platform. Despite the advantages of the online setting, the transition from the lab is challenging, especially when studying behavior. In this project, we develop methods to translate the traditional status characteristics experimental setting from the laboratory to online. We conducted parallel laboratory and online behavioral experiments using three tasks from the status literature, comparing each task’s ability to differentiate on the basis of status distinctions. The tasks produce equivalent results in the online and laboratory environment; however, not all tasks are equally sensitive to status differences. Finally, we provide more general guidance on how to move vital aspects of laboratory studies, such as debriefing, suspicion checks, and scope condition checks, to the online setting.
More than 50 years ago, the sociologist Morris Zelditch Jr. (1969) posed the question, “Can you really study an army in the laboratory?” Zelditch made the convincing argument that although, of course, you cannot bring an entire army into the laboratory for an experimental study, you can successfully test theoretical factors that guide a soldier’s behavior in the laboratory. For example, in an army, soldiers are treated differently by rank; in a laboratory, theoretically similar status distinctions can be reproduced while controlling for extraneous factors, thus isolating the causal effect of rank (Zelditch 1969). Insights such as these (e.g., how a person’s rank, or status, influences interactions) allowed status characteristics and expectation states theory (SC-EST) to become one of the most influential sociological social psychological theories of the past 60 years (Sell 2018; Wagner and Berger 1993).
In the 50 years since Zelditch (1969) posed his provocative question, numerous laboratory experiments have been conducted confirming SC-EST predictions, with individuals higher in status viewed as more competent and capable by others, and in turn, given more influence in interactions than people with lower status. Key to the success of the SC-EST research paradigm is the standardized experimental setting developed to test the original theory and its extensions (Berger 2014; Berger et al. 1977; Correll and Ridgeway 2003; Webster and Sobieszek 1974). As originally developed and subsequently used, the standardized experimental setting takes place in a controlled laboratory environment. To study status processes, researchers typically bring participants into the laboratory and manipulate participants’ status, the status of those with whom the participant is interacting, and/or the conditions under which these interactions are occurring (Wagner and Berger 1993). Researchers then examine group dynamics on a task, examining how participants’ status affects their own and others’ behavior.
Experiments have been called the gold standard for studying behavior and causality, and they have long been used to provide causal tests of social science theory (Lucas 2003b; Webster and Sell 2014). This traditional role and strength of experiments—to provide causal tests of theoretical hypotheses—is increasingly complemented by the ability of experiments to reach broad swaths of the population, such as in population-based survey experiments (Mutz 2011). Because lab experiments are time and resource intensive, and because researchers sometimes wish to reach a specific population, use of online samples has increased rapidly. However, researchers tend to use online experiments to study attitudes and opinions, as studying behavior online remains difficult. In this study, we develop and test methods for adapting SC-EST’s standardized experimental setting to the online environment. In doing so, we join a growing movement across the social sciences of adapting traditional behavioral laboratory experimental methods to the online setting (Arechar, Gächter, and Molleman 2018; Giamattei and Lambsdorff 2019; Horton, Rand, and Zeckhauser 2011). The online setting brings both new opportunities and novel difficulties for behavioral experiments, which we discuss in depth.
To examine the potential for SC-EST experiments to be conducted online, we compare findings from a behavioral study conducted in the online setting with those of a comparable study conducted in SC-EST’s traditional laboratory environment. We also compare three different tasks commonly used to test SC-EST (meaning insight, contrast sensitivity, and decision making). In doing so, we provide a contemporary test of the sensitivity of the three tasks to status distinctions and examine whether the different tasks are equally able to transition from the lab to the online setting. Our results provide clear evidence in favor of some tasks over others.
In what follows, we first summarize the advantages and disadvantages of the laboratory and online mediums. We then provide an overview of SC-EST and the tasks that are typically used to test this theoretical paradigm. We follow with a description of our methodological strategies for translating experiments from the laboratory to the online setting. Finally, we conclude with tests of (1) the extent of suspicion the tasks elicit in the laboratory and online settings, (2) the tasks’ ability to meet SC-EST’s scope conditions in the online setting, (3) the tasks’ sensitivity to status differences, and (4) the differences (or lack thereof) in findings between the laboratory and online settings. All of our tools, including experimental mediums, protocols, and data cleaning/analysis files are available at GitHub (https://github.com/biancamanago/SCTOnlineTasks).
Background
Across the social sciences, a considerable amount of experimental work is moving online. This is partially due to the convenience of the online setting. Some things researchers are able to do in labs, however, such as (simulated) live interaction or deception, are more difficult to recreate in the online environment. Below, we review the advantages and disadvantages of laboratory and online experiments. Because past research has discussed these considerations at length (Arechar et al. 2018; Gillan and Daw 2016; Horton et al. 2011; Nosek, Banaji, and Greenwald 2002; Palan and Schitter 2018; Reips 2000), we pay particular attention to the factors most relevant for SC-EST, such as resources (material and immaterial), characteristics of the subject pool, data quality, experimenter control, and behavioral outcomes.
Benefits and Potential Drawbacks of Laboratory and Online Experiments
Resources
Compared with online studies, laboratory studies are resource and time intensive (Kraut et al. 2004; Litman, Robinson, and Abberbock 2017; Suri and Watts 2011). Laboratory experiments require physical space and research assistants, resources that may be difficult or impossible for some researchers to obtain. Additionally, laboratory research takes considerably longer to conduct than online studies. This is partly because laboratory-based experimental studies, like other forms of in-person research (e.g., interviews or focus groups), are put on hold when in-person interactions are difficult or impossible to arrange, because of circumstances as benign as university breaks or as serious as a global health pandemic (Kraut et al. 2004; Litman et al. 2017; Suri and Watts 2011). Online research is generally able to proceed despite such factors, thereby offering an important benefit for researchers.
In the laboratory setting, the construction and maintenance of participant pools takes considerable amounts of time and energy; in contrast, for online experiments, participant pools already exist through well-established platforms (e.g., Amazon Mechanical Turk [MTurk], Qualtrics, Prolific Academic, NORC’s AmeriSpeak, Ipsos’s KnowledgePanel). A readily available subject pool is not only a boon for the efficient dissemination of research, but it is also theoretically relevant. Because status differentials are by no means permanent, it is preferable to conduct a study over a short period of time, reducing the chance that acute events will quickly change societal views and resulting behaviors. In summary, compared with laboratory-based experiments, Internet-based experiments are less time intensive, thereby providing practical and theoretical advantages.
Subject Pool
In terms of the size and ability to customize a subject pool, the online environment has advantages over the laboratory environment. Depending on a university’s size and location, student or community-based subject pools may be limited in both size and heterogeneity (Henrich, Heine, and Norenzayan 2010; Tompkins and Swift 2019). In contrast, for online experiments, all researchers have access to a broad participant sample, regardless of their physical location (Casler, Bickel, and Hackett 2013; Kraut et al. 2004; Miller et al. 2017; Nosek et al. 2002; Pontin 2007; Tompkins and Swift 2019). Additionally, online subject pools are generally more heterogeneous than undergraduate samples (Mullinix et al. 2015). Heterogeneity of a subject pool has pros and cons. On the one hand, if researchers are interested in studying whether participants’ characteristics (e.g., their race/ethnicity, political attitudes, socioeconomic status) are associated with reactions to experimental manipulations, relatively homogeneous undergraduate subject pools provide little variation. On the other hand, if researchers are interested in isolating theoretical mechanisms, homogeneous samples can effectively control for extraneous variability (Lucas 2003a). Importantly, online samples offer the ability to customize; researchers can make the samples as hetero- or homogeneous as the research question and theoretical argument requires.
Data Quality
Online subject pools are more easily accessible than laboratory samples, but researchers have questioned the quality of data from online samples. For example, some researchers have worried that online participants will be less attentive, will be less honest, and will have higher rates of attrition than laboratory participants (Chandler et al. 2015; Paolacci and Chandler 2014; Zhou and Fishbach 2016). Despite these concerns, online participants are continuously found to be a source of high-quality data, with similar or better measures of attention than laboratory participants (Casler et al. 2013; Hauser and Schwarz 2016; Horton et al. 2011; Miller et al. 2017; Palan and Schitter 2018; Paolacci and Chandler 2014; Paolacci, Chandler, and Ipeirotis 2010).
One outstanding concern is the non-naiveté of online participants. Some research suggests participants who have experienced deception in a previous study respond similarly to naive participants in subsequent studies (Barrera and Simpson 2012). Other research suggests effect sizes are attenuated by “career participants,” i.e., respondents who participate in many academic studies (Chandler, Mueller, and Paolacci 2014; Chandler et al. 2015). The evidence is thus mixed in terms of how past experience in research affects subsequent participation. In our empirical studies, we examine the effect of non-naiveté on data quality, specifically on rates of suspicion when studies use deception, as is the case with many SC-EST experiments.
Experimenter Control
Another possible drawback to the online environment is the lack of experimenter control and, potentially, threats to the internal validity of the study (Crump, McDonnell, and Gureckis 2013; Reips 2000). In contrast to lab studies in which all participants experience roughly the same conditions, the online environment is rife with possibility. Unless researchers control for these factors, participants can take the study at any time of day or night, on any device, and in any setting. This convenience for participants is unparalleled, but these external factors introduce the potential for extraneous variation, unrelated to the experimental manipulations, to affect the data (Reips 2000). Indeed, past research has examined how even the smallest changes to experimental procedures can affect SC-EST experimental findings (Kalkhoff and Thye 2006; Kalkhoff, Younts, and Troyer 2008; Troyer 2001). Additionally, despite software designed to deter this behavior, participants can still look up information during the study on another device; this possibility is especially concerning for studies that involve deception and simulated interaction with a partner, such as SC-EST. That said, trials are usually timed, which presents very little opportunity for participants to look up information about the tasks at hand. Nonetheless, such possibilities necessitate careful experimental and debriefing designs that are able to determine participants’ mind-sets during online studies.
Studying Behavior
A key benefit of laboratory studies is the ability to directly observe behavior (Bales et al. 1951; Benard and Mize 2016). Compared to self-reported attitudes or opinions, examining behavior has some key benefits. First, people are not always accurate at guessing how they would feel or behave in a situation, as they are sometimes asked to do in vignette-style studies, a popular type of survey experiment (Collett and Childs 2011; Mutz 2011). Second, although meta-analyses suggest that people’s attitudes and behaviors tend to be highly correlated for most issues, there can be a disconnect when studying some issues, such as socially desirable topics like race and gender, where participants behave notably differently than their expressed attitudes would suggest (Pager and Quillian 2005; Quadlin 2018; Vaisey 2014). A focus on behavioral outcomes has long been a strength of the SC-EST literature, which shows that status distinctions lead to both different assumptions about and behavior toward others (Ridgeway 2019). Increasingly, researchers across a range of fields, such as behavioral economics and social psychology, are demonstrating that some behavioral outcomes can be reproduced in the online environment (Arechar et al. 2018; Horton et al. 2011; Melamed, Harrell, and Simpson 2018; Simpson et al. 2018). We extend this work by examining whether SC-EST behavioral experiments can also be successfully conducted online.
Summary
In conclusion, both online and laboratory studies have advantages. Compared with online studies, laboratory studies are resource and time intensive, making them nearly impossible to conduct for some researchers. Online experiments require fewer resources, are faster to conduct, and allow the choice of hetero- or homogeneous samples. Despite these benefits, online studies may involve participants who are less naive, thereby reducing effect sizes, and may suffer from a loss of experimenter control, resulting in reduced internal validity, especially when using deception or simulated partners. Finally, although some behavioral experiments have been successfully conducted online, it remains unclear if SC-EST behavioral studies can be replicated online.
Status Characteristics and Expectation States Theory
In this project, we examine the viability of laboratory and online environments for three tasks regularly used in SC-EST research. In what follows, we provide a theoretical overview of SC-EST and derive hypotheses that we use to benchmark how well each task operates in the laboratory and online setting. We conclude with a methodological overview, leading to our methods and experimental design.
Theoretical Overview
SC-EST details two kinds of status characteristics, diffuse and specific, which can vary throughout time and encompass a broad range of social distinctions. Broadly, diffuse status characteristics have at least two differently evaluated states that come with wide-ranging expectations for ability. Race, gender, and education are examples of diffuse status characteristics. Diffuse status characteristics often convey expectations for specific abilities (e.g., compared with women, men are assumed to be especially capable at car repair), but the key aspect of a diffuse status characteristic is the expectation that high-status individuals will be “diffusely ‘better’ at most things” (Ridgeway 2019:77). Like diffuse status characteristics, specific status characteristics also have at least two differently evaluated states, but in contrast to diffuse status characteristics, they convey specific, not general, expectations for ability. Examples of specific status characteristics include artistic or mathematical ability. SC-EST predicts that, in general, people have lower performance expectations for lower status individuals (e.g., people with high school degrees) than for higher status partners (e.g., people with bachelor’s or master’s degrees). These performance expectations are often measured in terms of general perceptions of a person’s competence and capability, with people seen as diffusely more competent and capable expected to perform better at all tasks (Berger et al. 1977; Correll and Ridgeway 2003; Melamed, Munn, et al. 2019; Ridgeway 2019; Wagner and Berger 1993). Because of this perceived imbalance, SC-EST research consistently demonstrates that in social interactions, higher status actors are deferred to more often than their lower status counterparts (Berger et al. 1977).
In this project, we use education as an “ideal type” diffuse status characteristic to test whether SC-EST’s predictions can be verified in the online setting as they have been in the laboratory (Berger et al. 1977; Ridgeway 2019; Weber 1922). We chose education for both theoretical and practical reasons. Theoretically speaking, we chose education because it meets the formal definition of a diffuse status characteristic by (1) having at least two levels that are (2) ordered and (3) broadly associated with different expectations for ability. Practically speaking, education is a good choice because (1) we could control for it easily in the laboratory setting among a population of college students, (2) it is relatively stable to external factors, 1 and (3) it has more than two levels. By including multiple levels of education, we can see if the tasks and mediums differentiate for both lower and higher status partners.
On the basis of SC-EST, for all tasks (introduced in the next section), we predict the following:
Hypothesis 1: Participants will view partners with less educational attainment as less competent than partners with higher levels of educational attainment.
Hypothesis 2: Participants will defer less to partners with less educational attainment than to partners with higher levels of educational attainment.
Methodological Overview: Three Tasks
SC-EST emphasizes the importance of both theoretical and methodological rigor. For example, SC-EST has specific scope conditions that guide the design of tasks used to measure status differences (as well as a list of scope conditions for the theory more generally). The scope conditions include (1) task orientation, (2) collective orientation, and (3) lack of prior experience with the task. All three scope conditions are included to ensure the status generalizing process will occur. Task orientation is usually achieved by incentivizing task performance. For example, in our and many SC-EST studies, participants are told that the highest performing team will earn a bonus in addition to their standard compensation for participation. This potential bonus incentivizes participants to give their best effort on the task at hand. Task orientation can be measured, for example, by asking participants if they tried their best on the task. Collective orientation is usually achieved by telling participants that teams perform better on these tasks than individuals and that scores on the task are based on joint or collective answers. Collective orientation is often measured by asking participants if they took their partner’s answers into account. Finally, a lack of prior experience with the task is important, because if participants have had experience with such a task, they may either know it is not real (in the case of the tasks we use) or, if tasks have correct answers, participants may have expertise that would affect the status generalization process.
To study these status dimensions, SC-EST researchers have used a variety of tasks; we focus on three: meaning insight, contrast sensitivity, and decision making. Meaning insight asks participants to match an English word with one of two words from an ancient language (Conner 1964). Contrast sensitivity asks participants to choose which of two pictures with colored rectangles has a higher percentage of black or white color in it (Berger et al. 1977; Moore 1968). Finally, decision making asks participants to choose between two 2 responses in a series of workplace situations (Lucas 2003a; Mize 2019). Figure 1 shows examples of one trial for each task.

Example trials from three status tasks: contrast sensitivity (left), meaning insight (center), and decision making (right).
To ensure that none of the tasks would be read as associated with a specific ability (e.g., meaning insight could favor people perceived as having better language skills), we adapted commonly used language from SC-EST’s standardized experimental situation, which states the tasks are associated with a new general (i.e., diffuse) ability (Berger 2014). Specifically, we stated, In the past few years, social scientists have found that individuals differ in their Decision-Making ability. Specifically, some individuals have a high ability to make the right decision in difficult situations whereas others have a low ability to do so. Currently, social scientists do not know why some people are better at this than others, but they do know that this ability is not related to specialized skills, such as mathematical or artistic ability. So, individuals with low mathematical or artistic skills may have either high or low Decision-Making ability, and individuals with high mathematical or artistic skills may have either high or low Decision-Making ability. The same pattern applies to those with other types of specialized skill. This particular kind of Decision-Making ability appears to be an entirely new kind of ability, one that is unrelated to most other skills and abilities.
These instructions were identical across all three tasks, with the exception of the italicized passages, which varied across conditions to match the task name (these passages were not italicized when presented to participants).
On each task, participants are told that they are interacting with a human partner; but to ensure that partners present the same behavior to every participant, “partners” are usually simulated via computer programs. For each task, participants see an initial question and are asked to select one of two answers. Next, participants see their “partner’s” answer (which disagrees with their own 15 of 20 times) and then are asked to choose a final answer, either keeping their initial response or changing their answer. Deference is calculated by examining the number of times, among the 15 disagreements, that a participant changes his or her answer to match his or her “partner’s.” If a participant defers 10 times, that would indicate that the participant is highly deferential and most likely the lower status partner. If a participant defers only once, that would indicate that the participant is not at all deferential and is most likely the higher status partner. 3
Although the tasks are different, they share important features in common. First, the tasks do not have correct answers. For meaning insight, the ancient language does not actually exist; for contrast sensitivity, each block has 50 percent white and black squares; and for decision making, there are only two wrong answers to choose from, as the correct answer is not listed. If the tasks had correct answers, they would not only be measuring deference on the basis of a partner’s presumed competence and capability but also participants’ ability or strength on these particular kinds of tasks. Second, the tasks are not associated with other abilities. For each task, participants are specifically told that the ability is unique, and not like other abilities such as math or language (for overviews of the “standardized experimental situation” in SC-EST research, see Berger 2014; Correll and Ridgeway 2003). This is to ensure that participants’ behavior is affected by diffuse, not specific, status characteristics.
Methods
Procedures
Because past research has found that small differences in experimental protocols can have meaningful effects on findings in status research (Kalkhoff and Thye 2006; Kalkhoff et al. 2008; Troyer 2001), we tried to keep as many procedures as possible the same across samples. After providing informed consent, participants were given basic instructions about the study and were told they would be participating with someone else on a joint task. In the lab experiment, participants were told they were interacting with a partner in another cubicle in the same room or in the next room using a computer terminal. In the online experiment, they were told that they were interacting with a partner over the Internet. In fact, participants were “interacting” with a simulated partner in both experiments. The software used in both experiments was the same and used the same simulated partner algorithm.
In both experiments, participants were randomly assigned to one of six experimental conditions: one of the three tasks (contrast sensitivity, meaning insight, or decision making) and one of two partner status levels (partner has higher education or partner has lower education). Before assigning participants to a condition, they first completed a basic demographic questionnaire, ostensibly so their partner could learn more about them and vice versa. We used this demographic questionnaire to present the partner as having the same status characteristics as the participant on every dimension except education.
For our manipulation of partner’s education, we randomly presented the partner as having either lower or higher educational attainment than the participant. In the lab study, all participants were undergraduate students at a flagship state university. To present the partner as having lower educational attainment than the lab participants, we presented the partner as currently pursuing a GED. To present the partner as higher status, we stated that the fictitious partner was currently completing a master’s degree. For the online study, we sampled only participants who had completed a bachelor’s degree. To present the partner as having lower educational attainment than the participant, we presented the partner as having completed a high school degree. 4 For the high-status partner in the online sample, we presented the partner as having completed a master’s degree.
After being assigned to a partner, participants completed a joint task. To increase task orientation, we told participants that the group with the best score would be awarded an additional $50 once the study was finished. 5 To simulate a realistic-seeming partner, we programmed the “partner” to take between −10 and +10 seconds to respond after the participant responds. In cases for which the random time variable was greater than zero, participants would think that they took less time than their partner. If so, they saw a waiting screen that asked them to wait for their partner to make a choice before their partner’s choice was revealed. In cases in which participants were led to believe that they took more time than their partner (i.e., the random time variable is less than zero), participants saw their partner’s responses immediately after they submitted theirs. This strategy is designed to reduce suspicion, because the partner does not always take more time than the participant. If the “partner” always took more time than the participant, it might clue participants to the fact that the partner’s choice was dependent on theirs.
After completing 20 rounds of the randomly assigned task, participants were asked to answer a series of questions about the task and about their partner. Then, they were debriefed, told the true purpose of the study, and provided an opportunity to ask questions and/or withdraw from the study. Two participants in the online study and no participants in the lab study chose to withdraw the use of their data at this stage. For the lab study, we used a standard funnel debriefing strategy with research assistants conducting face-to-face interviews with participants and determining if participants were suspicious of the study before the deception was revealed to them (Blackhart et al. 2012). In the online study, in-person interviews and probing were not possible, so we used a modified funnel debriefing questionnaire designed specifically for the online implementation of the SC-EST tasks. We later coded these responses to generate a measure of suspicion that was analyzed as part of our comparison between lab and online settings (see Appendix A). We discuss this coding strategy and our analyses of the suspicion rates in further detail in the next section.
Modifications for the Online Experiment
To reduce the potential of suspicion among online participants, we modified the traditional laboratory experimental design. Perhaps the most necessary change involved making the likelihood of two strangers’ participating in the same study at the same time seem plausible. In the lab study, we had between two and six participants sign up for the same time slot, so there were always multiple participants coming into the lab to participate at the same time. In cases in which we had an odd number of participants or only one participant showed up for the study, we had a confederate “participate” to increase believability. This is not possible or even probable in an online setting.
To resolve this issue, we simulated a waiting room for participants (for a similar approach, see Giamattei and Lambsdorff 2019). After completing the online consent form, participants were assigned an ID that stayed with them throughout the study. Then, they were placed into the simulated waiting room, which showed a list of people waiting to be assigned to a partner. They saw their participant ID added to the list, a stream of other participant IDs entering the waiting room, and then random pairs of IDs being matched and leaving the room. After a short period of time, the participant was matched with another ID that recently entered the room. The matching process reduces suspicion because participants can see a dynamic waiting room full of others waiting to be matched on some undetermined criteria. Figure 2 shows a screenshot of the waiting room.

Waiting room.
We adapted the standard debriefing process used in lab studies to the online setting. As is done in most laboratory studies, we used a funnel debriefing approach. Rather than relying on research assistants’ probing questions to determine if participants were suspicious, we asked a series of yes/no and open-ended questions about participants’ impressions of the study. These questions began broad and then progressively asked about more specific parts of the study. Participants were not allowed to go back to change their text responses, thereby preventing them from changing their responses after seeing more specific questions or after the deception was revealed. The full online funnel debriefing is included in Appendix A. Overall, these changes make adapting the standard laboratory setting possible without compromising the standard design. We used these debriefing responses to determine participant suspicion and to determine whether SC-EST’s scope conditions were met (decisions usually made on the basis of in-person debriefing in lab studies).
Often, the only record of the debriefing interviews in laboratory experiments is summary notes from the research assistants who conduct the interviews. We trained our research assistants to ask a series of structured and follow-up questions on the basis of participants’ answers and their reactions, such as their body language. While debriefing, the research assistants took notes including brief highlights of participants’ answers to the debriefing questions (e.g., bullet points), as well as the research assistants’ impressions of participants’ answers. In the lab sample, participants were deemed suspicious, and thereby excluded from analysis, on the basis of research assistants’ perceptions and a principal investigator’s review of each set of notes from the debriefing interview.
The novel funnel debriefing approach we designed for the online sample affords us the opportunity to examine suspicion in fine-grained detail, although without the benefit of follow-up questions or reading participants’ nonverbal reactions. Specifically, we use online participants’ written responses to determine levels and types of suspicion for exclusion criteria across tasks; this is akin to having a full transcript of an interview, thus allowing for detailed content-coding of responses. We also use these content-coded responses to determine if suspicion is systematic by participant experience in the online sample.
Potential Effects of Design Differences between Experiments
Because design differences have been shown to affect research findings in the SC-EST setting (Kalkhoff and Thye 2006; Kalkhoff et al. 2008; Troyer 2001), we tried to minimize differences between the online and laboratory settings. For example, we designed the online and lab waiting rooms to be as similar as possible: participants saw others come and go without directly interacting with them. Similarly, the debriefing questions followed the same order and probing patterns in both the online and laboratory studies but were asked by a computer as opposed to a research assistant, respectively. Finally, because of resource constraints, there were incentive differences between the two studies. Students in the laboratory study were compensated with extra credit in their course, whereas online participants were compensated in cash, being paid $5 (for an hourly rate of ~$12.50). In both the laboratory and online studies, participants were told that they had the potential to earn a $50 performance-based bonus. To further reduce any potential effects caused by modifications to the online experiment (compared with the lab experiment), we made no modifications to the main portion of the study; that is, the differences occurred either before the main portion of the study began (e.g., the waiting room) or after the main portion concluded (e.g., debriefing/compensation). In sensitivity analyses, we examine if these minor design changes led to similar patterns of results and effect sizes across experiments (see “Design Differences between Experiments” in the “Sensitivity Analyses” section).
Participants
A total of 201 volunteers participated in the lab experiment for extra credit in their introductory-level undergraduate course. In the lab portion of the study, we restricted the sample to self-identified women undergraduate students. Restricting the sample to one gender is a common approach in SC-EST laboratory studies and is used to ensure that status characteristics extraneous to the study (e.g., gender) do not affect status processes for the main characteristic of interest (e.g., education) (Berger et al. 1977). Roughly 16 percent of participants were excluded for failing manipulation checks (i.e., they could not correctly identify their partner’s education level at the end of the study). A further 10 percent were excluded because of suspicion about some aspect of the study. Our final analytic sample was nlab = 142 participants.
For the online version of the experiment, 320 volunteers from Prolific (prolific.co) participated in the study. Similar to the popular Amazon MTurk, Prolific provides access to a convenience sample of participants. In contrast to MTurk, Prolific has been shown to have some advantages in terms of diversity of participants and data quality (Peer et al. 2017). Prolific also makes it simple to quota sample on the basis of characteristics of participants. For the online experiment, we did not restrict on the basis of gender; however, all participants were told their “partner” was of their same gender (more detail in the “Potential Effects of Sample Differences” section). To maintain comparability with the lab experiment’s education manipulation, all participants in the online experiment had a bachelor’s degree.
As noted in the “Data Quality” section, some research has found that in the online setting, effect sizes are smaller for “career participants” relative to naive participants (Chandler et al. 2014, 2015). Therefore, to ensure variability in participants’ past experience with academic studies, we restricted half our sample to participants who had completed fewer than 25 studies on Prolific and who were not part of other online samples (e.g., MTurk); the majority of these participants had completed fewer than 10 prior studies. We examine the effect of past experience on suspicion and on overall effect sizes (see the “Online Study: Suspicion Based on Participant Experience” and “Discussion and Conclusions” sections).
Of the 320 participants who completed the study, 12.50 percent were dropped for failing manipulation checks. Unlike the lab experiment, the online setting does not have a standardized convention for determining suspicion, so we developed one. We discuss this strategy below and treat level of suspicion as an outcome of interest in some of the analyses. We excluded three participants who refused all the debriefing questions, as they provided no way to assess suspicion (leaving 278 participants for our analyses of participant suspicion). As we discuss in further detail below, for the main analyses testing the hypotheses, we also excluded the 10.79 percent of the sample deemed suspicious on the basis of content coding of their debriefing answers. We further excluded five participants whose responses indicated they did not meet the theory’s scope conditions (see “SC-EST Scope Conditions” in the “Results” section). Our final analytic sample was nonline = 243. Table 1 shows sample characteristics for the lab and online samples.
Descriptive Statistics
Potential Effects of Sample Differences
One of the benefits of online experiments is the ability to have a diverse or targeted sample (Henrich et al. 2010; Mutz 2011). To use the online environment in the way it is typically leveraged, we did not restrict the online sample to match the lab sample. This design choice, however, makes direct comparisons between the lab and online studies a bit more complicated. For example, it introduces the possibility that any differences in results were driven by sample differences between the lab and online experiments. Although this is a possibility, recent research suggests it is unlikely. Specifically, prior work shows that heterogeneous treatment effects are rare; that is, different participants enter experiments with different levels of opinions, attitudes, and baseline behaviors, but their reactions and responses to experimental manipulations tend to be very similar (Coppock 2019; Coppock, Leeper, and Mullinix 2018). In sensitivity analyses, we examined the potential effects of sample differences (see “Sample Differences” in the “Sensitivity Analyses” section).
Measures
Before analyzing our hypotheses, we first examine rates of suspicion among the online participants by task. The lead researcher along with two research assistants read and evaluated responses using a priori codes, which were later refined through an iterative coding process. This process led to three codes: suspicion about the study’s focus, authenticity of the task, and existence of a live human partner. Suspicion about the study’s focus captures participants’ intuition that the study is designed to examine how people’s education level may affect levels of deferential behavior. Suspicion about task or partner captures participants’ beliefs that the task or partner were fictitious, respectively. For each theme, we coded the suspicion into three categories: (1) not suspicious at all, (2) somewhat suspicious, and (3) definitely suspicious. All codes had high intercoder reliability with simple agreement ranging from .94 to .99 (Compton, Love, and Sell 2012). 6
Finally, we assess online participants’ past experience in academic studies (our measure of non-naiveté) by asking how many past academic surveys or studies they had participated in. We asked the question in nine binned categories (e.g., 0, 1 to 10, up to 200 or more). For some analyses, we imputed the midpoint of the category and treated the measure as continuous.
The primary outcome in the study is the number of times the participant deferred to their partner’s suggestion out of the 15 possible deference opportunities (i.e., rounds in which the participant and partner initially disagreed). On the basis of SC-EST, higher status partners are deferred to more because they are perceived to be broadly more competent than their lower status partners (Melamed, Munn, et al. 2019; Ridgeway 2019). Therefore, we also included a self-report scale of perceptions of a partner’s competence by asking participants to indicate on eight-point bipolar scales how capable, competent, skillful, efficient, intelligent, and confident their partner is (with anchors of “not at all” and “extremely”; adapted from Fiske et al. 2002; Ridgeway et al. 1998). The scales have high internal consistency; Cronbach’s α = .93 for the lab experiment and α = .92 for the online experiment.
Analytic Strategy
We begin by reporting suspicion rates. To determine if suspicion is systematically different across tasks, we used χ2 tests of association. For the online sample, we also examined the relationship between participant experience and suspicion. To do so, we calculated polyserial correlations between participant experience and suspicion codes.
Our two primary dependent variables, competence perceptions and total deference, are both continuous, so we used linear regression models to examine the effects of partner’s educational attainment. In all models, we included interaction terms between partner’s educational attainment and task to allow the effect to vary. Our interpretations focus on the marginal effect of “partner’s” education for each task (Long and Freese 2014).
With experimental data, control variables are not necessary because random assignment controls for participants’ demographic characteristics by balancing representation of participant characteristics across conditions. If certain demographic characteristics are independently predictive of the outcome, however, use of control variables can increase the precision of estimates, especially in diverse samples such as our online sample (Mutz 2011). To this end, we tested whether participant’s gender, income, marital status, parental status, age, race, or past experience with surveys independently predict either of our dependent variables: perceptions of competence or total deference. None of the demographic characteristics were associated with either outcome at p < 0.05 (two-tailed test). Therefore, we do not include any control variables in the models presented below.
We use one-tailed tests for the tests of hypotheses 1 and 2, as these are directional predictions; all other tests are two tailed. All results discussed in the text are significant at the p < 0.05 level unless otherwise noted. Levels of statistical significance are included in Tables 2 and 3.
Proportion of Participants Who Were Somewhat or Completely Suspicious about Various Aspects of the Study (Online Sample, n = 278)
The χ2 statistic is reported in the table. Each test is a 2-df test of the association between task and the suspicion code.
Participant experience is measured as the number of academic studies the participant has previously participated in.
p < .05 and ***p < .001 (two-tailed tests).
Means (and Standard Errors) for Competence Perceptions and Deference Rates
Note: Significance tests are for the contrast between a partner with a GED or high school degree and a partner with a master’s degree (i.e., the marginal effect of partner’s educational attainment). All tests are one tailed, as these are directional hypothesized relationships. Superscripts show significant effects of partner’s educational level.
p < .05 and ***p < .001 (one-tailed tests).
Results
Suspicion
Lab Study
As noted earlier, in the lab study, a little over 10 percent of participants (10.13 percent) were excluded from analyses because of suspicion about the study. For the lab sample, suspicion rates were not constant across tasks. Roughly 19 percent of participants in the meaning insight task were suspicious enough to merit exclusion, compared with only 6 percent for the contrast sensitivity task and 4 percent for decision making. 7
Online Study: Suspicion by Task
Given the novelty of the online format for the tasks and concerns about non-naiveté, we conducted a more thorough analysis of suspicion rates for the online sample. This analysis was facilitated by our access to participants’ written answers from the adapted form of funnel debriefing designed for the online setting. As noted earlier, there were three types of suspicion (purpose, task, and partner) and three levels of suspicion (not at all, somewhat, definitely). The top panel of Table 2 shows results for suspicion rates (proportions) across each task (online sample only).
Rates of “definite” suspicion (for any reason) are a little under 10 percent in the online sample, similar to the lab sample. As in the lab sample, rates and types of suspicion varied across tasks. Perhaps unsurprisingly, suspicion about the focus of the study, which usually manifested as suspicion about educational status being the experimentally manipulated independent variable, did not vary systematically across the three tasks. Suspicion about authenticity of the task did vary, with the meaning insight task provoking more “somewhat” and “complete” suspicion responses. A little over 18 percent of participants expressed some suspicion about the meaning insight task (combined somewhat and definitely suspicious); however, the low proportion of these responses that were “definitely” suspicious (.041) is relatively reassuring. In contrast, rates of suspicion for the authenticity of the other two tasks were quite low (fewer than 1 percent of participants “definitely” suspicious about either). We see no difference in rates of participants being “definitely suspicious” about the partner being real across tasks; however, rates of participants being “somewhat suspicious” do vary, with decision making producing the most of this mild form of suspicion about the partner being real. Considering that “definite” suspicion most closely mirrors exclusion criteria used in laboratory studies, we use it as the exclusion criteria for the following analyses. The meaning insight task is the most likely to arouse suspicion from participants in both the lab and online settings.
Online Study: Suspicion Based on Participant Experience
Next, we examined the relationship between online participants’ past experience with academic studies and their level of suspicion. For proportions reported in the bottom half of Table 2, we combined some of the original categories because they occurred rarely, thus making inferences difficult; we reduced the number of ordinal categories from nine to five. The results, focusing on the polyserial correlations to start, suggest no relationship between previous experience with academic studies and suspicion about the focus of the study. There are small correlations between participant experience and suspicion about the authenticity of the task (somewhat suspicious, r = .15; definitely suspicious, r = .18), although only the correlation with the somewhat suspicion code is statistically significant. 8 There is also a relationship between participant experience and suspicion about the partner being a real person (r = .23 for somewhat suspicious, r = .19 for definitely suspicious). About 37 percent of participants who had participated in at least 100 studies before (i.e., “career participants”) expressed at least some suspicion about the partner. This contrasts with fewer than 8 percent of naive participants (no previous experience) expressing any suspicion. These general patterns hold for the “somewhat” and “complete” suspicion codes. As was revealed in their responses, most of these suspicious participants had participated in studies with simulated partners in the past.
SC-EST Scope Conditions
SC-EST has two important scope conditions that should be checked to ensure the theory’s predictions will apply: task orientation and collective orientation. In the laboratory sample, we used participants’ responses to research assistants during debriefing to determine and conclude that all three tasks produced both task and collective orientation. In the online sample, we used answers to survey questions and participants’ open-ended responses. We asked, “Did you try and do your best?” as an initial screen for task orientation; all participants stated “yes.” We asked, “Did you take your partner’s answers into account when deciding on your final answer?” as an initial screen for collective orientation; 211 of the 248 nonsuspicious participants (85.08 percent) stated “yes.”
To further examine the 37 participants who indicated that they may not have been collective oriented, we focus on responses to two open-ended questions: one simply asking them to elaborate on their answer to the binary choice question about taking their partner’s answers into account during the task and a second asking what impressions they developed of their partner (see Appendix A for exact question wording). Of these 37 participants, 16 expressed information that suggested they were collectively oriented, 16 expressed ambivalent information about their collective orientation, and 5 answered definitively in ways suggesting they were not collectively oriented. In examining the 37 participants who indicated they were not collectively oriented, we also coded for task orientation. We identified two participants whose responses were mixed regarding their task orientation and one who expressed that they definitely were not task oriented. Note these codes are not mutually exclusive; in total, we identified five participants (2.02 percent of the nonsuspicious sample) who were definitely not task and/or collectively oriented and thus exclude them from subsequent analyses. 9 The low rate of failure on the scope condition checks suggests all three tasks translate well to the online setting in terms of their ability to meet the theory’s scope. 10
Competence Perceptions of Partner
SC-EST proposes that decisions about deference are driven mainly by perceptions of a partner’s competence and ability. For diffuse status characteristics such as educational attainment, these perceptions of general ability should apply to any task. Because the only interaction the participant and “partner” have is their work together on the task, higher status individuals should be perceived as more competent and capable if the task is successful at picking up on these diffuse status differentials. Figure 3 presents the mean ratings of the partner’s competence across each task (each set of two bars) and partner’s educational attainment (GED or high school degree or master’s degree) with standard error bars. Table 3 includes the same information as well as superscripts for significance tests of the marginal effect of partner’s educational level.

Evaluations of partner’s competence across the three status tasks, on the basis of partner’s educational attainment.
As the left panel of Figure 3 illustrates for the lab sample, the direction of effects shows that partners with master’s degrees are always viewed as more competent than partners with a GED, as predicted. However, the only significant contrast is in the decision-making task, with partners pursuing a master’s degree seen as a little over half a standard deviation more competent than partners pursuing a GED. The right panel of Figure 3 shows the same set of results but for the online sample. Again, the direction of effects always favors the partner with more education. For the online sample, the hypothesized contrast is significant for both the meaning insight and decision-making tasks; effect sizes suggest a little less than half a standard deviation difference for the meaning insight task and almost three fourths of a standard deviation difference for decision making. Taken together, the results suggest that only participation in the decision making task consistently leads to the differential competence perceptions predicted by diffuse status advantages benefiting individuals with more educational attainment.
Deference
Much SC-EST work focuses on the behavioral outcome of deference: how often a participant defers to a partner’s suggestions, with higher status partners predicted to elicit more deferential behavior. The left panel of Figure 4 presents results for the lab sample; the bottom panel of Table 3 reports the means and tests of significance. For the lab setting, a clear pattern emerges: only the decision-making task produced the predicted effect of the partner pursuing a master’s degree being deferred to more than the partner pursuing a GED. In the decision-making task, partners pursuing a master’s degree were deferred to an average of 2.64 times more than partners pursuing a GED, which is about a standard deviation difference (S.D.deference = 2.83).

Number of times the participant deferred to the partner’s suggestion across the three status tasks, on the basis of partner’s educational attainment.
The right panel of Figure 4 reports the results for deference from the online sample. Again, the decision making task picks up on the predicted advantage for the partner with a master’s degree. In this sample, the meaning insight task also produces the predicted status effect. The contrast sensitivity task does not distinguish between the partner with a high school degree and the partner with a master’s degree in either the lab or online sample.
Sensitivity Analyses
Design Differences between Experiments
The results as presented thus far focus on whether each task produces a significant marginal effect of partner’s education, separately for the lab and online samples. As one of our primary interests is whether the tasks can be successfully translated to the online setting from their traditional laboratory setting, we provide more direct tests of this question. Specifically, because past research finds that small protocol differences can result in differences in findings between studies (Kalkhoff and Thye 2006; Kalkhoff et al. 2008; Troyer 2001), we examine if the minor differences between the laboratory and online studies may have affected findings.
For these tests across samples, we examine whether the marginal effect of partner education in a given task (e.g., meaning insight) is equivalent in the lab and online samples. We used a Wald test of equality appropriate for cross-sample tests of marginal effects in nonoverlapping samples (Mize et al. 2019:165–66). In total, there are six tests of equality (3 tasks × 2 dependent variables). For all six tests, we found no significant differences in the effect of partner’s education across the lab and online samples. This finding is especially promising for the decision-making task, as the accumulation of evidence suggests the decision-making task picks up on the diffuse status characteristic of education in both samples and produces statistically equivalent effect sizes across contexts.
Sample Differences
To ensure comparability of findings despite sample differences in the online and laboratory settings (see Table 1), we tested for heterogeneous treatment effects by comparing patterns of deference for participants with similar characteristics in the lab (women between ages 18 and 25, n = 47) and online (all men; women ages 25 and older, n = 201) samples. Across all three tasks, effects of the experimental manipulation did not differ across these two types of participants (p > .10 for all comparisons). This finding supports past research that reports similar trends (Coppock 2019; Coppock et al. 2018). Therefore, sample composition differences are unlikely to be a primary driver of any potential differences in results across samples.
Summary of Results
In terms of suspicion rates, we found similar rates and patterns in the lab and online versions of the study. Only the meaning insight task showed notable problems with suspicion; our content-coded debriefing answers for the online sample suggest this is primarily driven by suspicion about the authenticity of the task itself. We also found that all three tasks translated well to the online setting in terms of participants’ ability to meet the theory’s scope conditions of task and collective orientation.
For a task to be appropriate for testing SC-EST predictions, it should be able to detect status differences for both perceived competence and deferential behavior. We used a widely accepted status characteristic, education, to test for each task’s sensitivity to status differentials. If we consider the combinations of the two outcomes (competence and deference) and two samples (lab and online) as four tests of SC-EST, using whether a hypothesis was supported or not at the p < .05 level as a criterion, the results are as follows for each task: meaning insight produces two of four supported hypotheses, contrast sensitivity produces zero of four supported hypotheses, and decision-making produces four of four supported hypotheses.
Discussion and Conclusions
This project compared the laboratory and online experimental environments for conducting SC-EST research. In doing so, we compared three tasks used to test SC-EST: contrast sensitivity, meaning insight, and decision making. In summary, we found that SC-EST research can successfully be conducted in the online experimental environment, with effect sizes comparable with the traditional laboratory environment. We also found that the three tasks are not equally sensitive to status manipulations, nor do they elicit comparable levels of suspicion.
Of the tasks examined here, the decision-making task appears to be the most sensitive to status differences and among the least likely to arouse suspicion in the laboratory and online settings. Notably, the widely used contrast sensitivity task did not perform well in either the lab or the online sample. Although we did not collect data for the specific purpose of determining why this is, some of the debriefing responses tentatively suggest participants may have been the least engaged and the most frustrated by the contrast sensitivity task. 11 This may have led to less task orientation for this task, which would reduce status effects. Below, we suggest ways to increase task orientation regardless of task. It is worth noting, however, that different tasks may produce more or less task orientation at baseline, and tasks that produce more task orientation should generally be preferred, all else equal.
To increase task orientation across all tasks, we followed past work by incentivizing performance and offering a $50 bonus to the group that performed best (Melamed, Savage, and Munn 2019; Ridgeway et al. 1998). An alternative or additional method is to incentivize each trial, paying participants on the basis of their number of trials answered correctly (Melamed and Savage 2016; Melamed, Savage, and Munn 2019). Each method should help increase task orientation, so we recommend using at least one, or both, whenever possible. We recommend that future research compare the two incentive structures, examining which elicits the most task orientation and if, as a result, there are differences in effect sizes.
Through examining the viability of the online setting for SC-EST research, we developed tools and gained insights that may be useful to future researchers inside and outside the SC-EST tradition. First, to assess suspicion among participants, we successfully translated the funnel debriefing strategy from the laboratory to the online setting. Second, to assess different types and levels of suspicion, we also developed a coding scheme. We believe the debriefing strategy and associated coding scheme can be successfully used in the online setting to assess suspicion in a variety of experiments. Finally, to mimic live interactions on the most common status tasks, we programmed a waiting room and set timed responses. We have made all these tools openly available to researchers interested in conducting SC-EST research or for those who wish to adapt these techniques for related research.
We have some suggestions for future researchers trying to measure or decrease suspicion in their own studies. First, when conducting laboratory research, we recommend that researchers consider combining online and laboratory methods, with participants originally responding to debriefing questions via computer and a research assistant later following up with probing questions. Second, when conducting studies online, we recommend that researchers not make base pay dependent on completing the task or answering all questions. We suspect that by not tying base pay rates to performance, we may have increased participants’ willingness to share their suspicions and other opinions. Relatedly, we caution that “career participants,” who are somewhat common in online environments, are less likely to believe that they are working with a fictitious partner but do not appear to be more suspicious of other aspects of the study. Nonetheless, we suggest researchers who are using deception, especially vis-à-vis interaction with a fictitious partner, consider omitting or limiting the proportion of career participants in the study.
Researchers have been cautious about adopting the online setting for behavioral experiments, but our study suggests the online environment is viable for SC-EST research. Note that the type of SC-EST study examined here is more complicated and involved than many social science experiments but less so than others, including some in the SC-EST tradition. On one hand, our study involves (1) deception, (2) interaction with a partner (albeit fictitious), (3) multiple rounds of interaction (20), and (4) an extended debriefing to check adherence to scope conditions, attention to the experimental manipulations, and suspicion. Many social science experiments are simpler than our study on one or more of these factors; our results tentatively suggest promise for such experiments. On the other hand, some SC-EST studies include more manipulations or stages than what we included in our design. For these kinds of studies, we encourage future research to examine the viability of combining the tasks studied here with additional features common to the online environment. On the basis of our results, we have confidence that the deference task portion of a study, which often produces the key dependent variable even in these more involved designs, can be widely implemented online.
This project was able to adjudicate some questions about the viability of the online setting for SC-EST research and the tasks used to assess status differences, but it was not able to examine others. We chose educational attainment for this study because it is a well-established diffuse status characteristic that should produce relatively large differences in performance expectations. Other status characteristics that produce smaller status differentials should be tested to determine how well each task is able to discriminate between status levels. In addition, while educational attainment confers status advantages, it is also associated with other beliefs and expectations for behavior. Creating a new status characteristic within the study environment—for example, as status construction researchers have done—would provide another important test of tasks’ relative abilities to distinguish status differentials (Ridgeway et al. 1998, 2009). We encourage future researchers to use a variety of different status characteristics, seeing if certain tasks are more sensitive to certain characteristics. Although this should not be the case, it cannot be said with certainty unless tested. Finally, although we examined three common SC-EST tasks, we did not examine all the tasks that have been used to measure status differences; future research could consider additional tasks.
In summary, in this project we compared the online and laboratory experimental settings for studying behavior and tested the sensitivity of three different SC-EST tasks to status differences (in both the online and laboratory settings). Our findings suggest that the online setting is viable for behavioral status research (with considerable deception, including “interaction”); the decision-making task is the most sensitive to status differences, and in contrast to the meaning insight task, it does not produce problems with suspicion. Importantly, neither meaning insight nor contrast sensitivity performed well in both the online and laboratory settings. We provide a series of resources for future researchers to continue this line of work, and we recommend researchers continue to consider the viability of the online setting for experiments traditionally conducted in the laboratory.
Footnotes
Appendix A: Online Funnel Debriefing
Thank you for participating in this study, we are just about finished. To conclude, you will learn more about the study, but first we’d like to ask you a few questions about what you thought of the study. Your answers to these questions will not affect your study payment.
Acknowledgements
We would like to thank Jeff Lucas, Murray Webster, and Sarah Harkness for experimental materials and helpful feedback. We would also like to thank Stephen Benard and Jane Sell for sage advice. We are grateful for excellent research assistance provided by Hannah Regan, R. Gordon Rinderknecht, and members of the Indiana University Sociology Lab. Hayley Hollenberg and Chanteria Milner expertly assisted with manuscript preparation. We would like to thank Indiana University, Vanderbilt University, Purdue University, and University of Maryland for institutional support for this project.
Notes
Author Biographies
.
.
