Abstract
We investigate how men and women are evaluated in group discussions. In five studies (N = 761) using a variant of a Hidden Profile Task, we find that, when experimentally and/or statistically controlling for actual gender differences in behavior, the female performance in a group discussion is devalued in comparison to male performance. This was observed for fellow group members (Study 1) and outside observers (Studies 2–5), in both primarily student (Studies 1, 4, and 5) and mixed samples (Studies 2 and 3), for different measures of performance (perceived helpfulness of the contribution, for work-related competence), across different discussion formats (preformulated chat messages, open chat), and when controlling for the number of female group members (Study 5). In contrast to our hypothesis, we did not find a moderating effect of selection procedure in that women were devalued to a similar degree in both situations with a women’s quota and without.
Despite the societal goal to establish gender equality, women are still underrepresented in high-profile positions such as the management levels of key financial institutions, advisory boards of large listed companies, and full professorships (European Commission, 2015, 2017). A wealth of research has identified discrimination of women in the hiring process for such high-profile positions as one causal factor (e.g., Heilman et al., 2015). This research predominantly focused on the evaluation of women’s personal attributes and their performance in individual tasks (e.g., Etaugh & Kasley, 1981; Moss-Racusin et al., 2012). However, group tasks, especially meetings and group discussions, constitute an important part of everyday business in high-profile positions. For example, senior managers attend approximately 23 hr of meetings per week—with a rising trend (Rogelberg et al., 2007). Group discussions are also important when entering a new job. For instance, the performance in group tasks is an important part of evaluation in assessment centers, which companies often use when recruiting for high-profile positions (e.g., Obermann, 1992). Although group tasks such as group discussions are essential for selection to and promotion in high-profile positions, biases in performance evaluations have not been sufficiently studied in this context. Some existing research found actual behavioral differences between men and women in group settings of high-profile positions (e.g., Hinsley et al., 2017), whereas only a limited number of studies have addressed performance evaluations in this context (e.g., Wood & Karten, 1986). As the latter studies typically involve face-to-face interactions, it is not possible to control for actual behavioral differences between men and women. Although these studies are valuable due to their high external validity, the question remains whether men and women are evaluated differently even when their actual behavior does not differ. This is an important notion, as training specifically for women would have limited impact if women’s performance is still evaluated more poorly than men’s. This research aims to fill this gap by systematically investigating whether men and women’s performance in group discussions is evaluated differently. To do so, we focus on one high-profile domain (i.e., the participation in a hiring committee for a professorship) and assess whether a bias in evaluations exists irrespective of actual behavioral differences between men and women. Isolating the pure evaluation bias is important for designing concrete measures to overcome discrimination of women in high-profile domains. In the present research, we also investigate whether one specific measure that has been introduced to eliminate the female underrepresentation in high-profile positions (i.e., gender quotas; European Commission, 2012) moderates a potential evaluation bias.
In the following, the theoretical and empirical basis for biases in the evaluation of men and women both in general and in group settings will be outlined.
Stereotypes and the Gender Gap in High-Profile Positions
One reason for the gender gap in high-profile positions is that men and women are evaluated differently in the work context (e.g., Heilman, 2012). According to Social Role Theory (e.g., Eagly & Wood, 2012), men and women are observed in different social roles (e.g., women taking care of kids and men as decision makers) and ascribed the specific traits that are required for those roles. Because men are more often observed in domains associated with agency and competence, whereas women fulfill communal roles, men are associated with agentic attributes/traits, whereas women are associated with communal attributes/traits. For high-profile positions, in turn, agency is thought to be necessary for success (Gaucher et al., 2011). As a result, women’s qualifications are seen to be unfit for traditionally male occupations and positions that require stereotypically male attributes (Lack of Fit Model; e.g., Heilman, 2012). More generally, characteristics that are stereotypically attributed to women (e.g., incompetence, warmth, passivity) do not match the requirements of high-profile positions (see also Schein, 2001, for the Think Manager-Think Male phenomenon). Due to this perceived lack of fit, women are expected to perform poorly, which leads to less favorable evaluations during selection procedures for male-dominated domains (e.g., Heilman et al., 2015). This theoretical assumption has been confirmed in many different contexts: For example, female applicants for a news writing job were given less favorable evaluations on attributes such as professional competence and predicted job success (Etaugh & Kasley, 1981), whereas male applicants for a laboratory manager position were rated as more competent and hirable (Moss-Racusin et al., 2012). Furthermore, competence was rated as a more desirable attribute for men as compared with women (Prentice & Carranza, 2002).
The Evaluation of Women in Group Discussions
The findings reported so far refer to the evaluation of women’s (as compared with men’s) individual characteristics or performance. However, not only individual performance but also performance in groups (e.g., group discussions) is important for promotion in and selection for high-profile positions (Obermann, 1992; Rogelberg et al., 2007). For example, in the group tasks of assessment centers, a person’s assertiveness is evaluated (e.g., Obermann, 1992)—an attribute belonging to agency, a typically male dimension (e.g., Diekman & Eagly, 2000). Previous research has mainly focused on actual behavioral differences between men and women in group settings of high-profile positions. It has, for example, been shown that men ask different questions in group discussions (Hawkins & Power, 1999), ask more questions at conferences (Hinsley et al., 2017) and in academic seminars (Carter et al., 2018), use a more dominant discussion strategy (Smith-Lovin & Brody, 1989), and interrupt others more often (Anderson & Leaper, 1998).
The evaluation of men’s and women’s performance in group tasks has been investigated by Berger and colleagues within the scope of Expectation States Theory (Berger et al., 1972). Their research deals with group tasks (e.g., in workgroups or committees) where group members must solve a task by making the right decision. The authors argue that, in the absence of specific information concerning task performance, salient status characteristics such as gender determine how group members are evaluated (Berger et al., 1972). As people attribute less task ability to women in male-typed tasks, women’s contributions are less positively evaluated by the group even though gender itself is not relevant to the group task (Correll & Ridgeway, 2003). Foschi (2000) goes one step further by stating that even if low-status individuals perform well at a task, their performance is devalued because good performance is not what would be expected from individuals of their status. Stricter requirements are applied to members of devalued groups in that women must demonstrate superior performance to be perceived as similarly competent as a male group member. Although the original account by Foschi (1989) only applies to settings where the person who makes the judgments is a member of the group, an extension of the Theory of Double Standards (Foschi, 2000) includes situations where the person making the judgment is not engaged in the group task.
Empirical results are in line with the theoretical assumption of an evaluation bias in favor of men in group tasks. In mixed-sex groups, women tend to have less influence over group decisions and group members hold lower expectations for women’s as compared with men’s task performance (Ridgeway, 1982). Furthermore, and directly related to group discussions, participants in a study by Wood and Karten (1986) worked in face-to-face four-person groups on a problem-solving task. After the discussion, participants had to rate each other with regard to their competence. Results revealed that male group members were perceived as being more competent as compared with female group members. However, it was also found that men and women behave differently in the group task, which again could have influenced performance evaluations. For example, it has been found that the more a person talked in a group discussion, the better that person was evaluated (Bales, 1970). 1 Relatedly, Riecken (1958) showed that the same idea was rated as more valuable when it came from a talkative as compared with less talkative group member. Other studies employed ambiguous tasks with no clear, correct answer to ensure the initial equality of all group members regarding their competence to solve the task (Ridgeway, 1982). However, this kind of task might reduce the collective orientation (i.e., the group goal to find the correct solution and solve the task) of the group members (see Correll & Ridgeway, 2003, for the components of Expectation States Theory).
Based on the theoretical as well as empirical background outlined above, we expect that men and women are also evaluated differently in group discussions when using a controlled setting and keeping several aspects of the discussion constant. However, it should be noted that a different theoretical account would allow for an alternative assumption. That is, according to the Model of Shifting Standards (Biernat & Manis, 1994), people use different standards when evaluating persons from different social groups on subjective response scales (e.g., scale from not at all competent to very competent). Specifically, individuals are not evaluated in an objective manner but compared with the average group member (i.e., shifting standards). That is, a woman’s performance in a group setting would be compared with the average (expected) performance of women as a group. This, in turn, implies that a woman showing a medium performance in a group task might be evaluated no differently or even better than a man showing the same performance, as his performance is compared with the average man of whom a high performance is expected. This would mean that, without any actual differences in behavior, we should observe no differences in performance evaluation in a group discussion or even a bias in favor of female group members.
The previous studies reported above were not ideal in terms of isolating the effect of biased performance evaluations, as one cannot rule out the possibility that the difference in perceived performance was (also) due to differences in appearance (e.g., in face-to-face interactions). Furthermore, considering previous findings, actual behavioral differences between men and women also seem likely. Specifically, it is possible that men and women provide different kinds of or more or less important information to the group discussion. The present research’s main contribution is to fill the empirical gap by including anonymous interactions where differences in appearance can be ruled out and behavior in the group discussion is kept constant or statistically controlled for. Specifically, in contrast to previous research, we rely on a paradigm—a variant of a Hidden Profile Task—that allows us to control the content (Studies 1 and 2) as well as the importance of each piece of information. Furthermore, as the purpose of the group discussion was to find an optimal solution (see Berger et al., 1972, for this characteristic of the group task), which was only possible when attending to and integrating different pieces of information, participants were incentivized to equally value the information provided by male and female group members. Finally, our design not only allows us to isolate the pure evaluation bias but also to systematically vary specific aspects of the group (incumbents and newcomers vs. equal members; gender composition in a group), evaluators (fellow group members vs. external observers), and selection procedure (quota vs. no quota) to test the robustness and the generalizability of the effect.
Preferential Selection as Potential Moderator for the Evaluation Bias
One way to overcome the female underrepresentation in high-profile positions is to introduce quota rules (European Commission, 2012), which has indeed substantially increased the percentage of women in decision-making positions such as corporate boards. 2 However, in terms of evaluations in group discussions, quota rules could also have adverse effects: Previous research has found that women are seen to be less competent when they were selected based on an affirmative action policy (Heilman et al., 1992; Heilman et al., 1997). Also, for group tasks, unwanted side effects have been observed in that quota rules undermine cooperation (Dorrough et al., 2016). Regarding the evaluation of men and women in group tasks, the effects of quota rules have not yet been investigated. In terms of Expectation States Theory, it can be assumed that a woman’s performance is devalued to a greater extent when she is selected based on a quota rule, as the shared belief regarding the competence of “quota women” is especially negative (see Heilman et al., 1997). Thus, as a potential moderator of the evaluation bias, we varied in three out of five studies whether or not a quota rule was applied. It is expected that women’s performance is evaluated even more poorly when people are informed that the selection for entering the group was based on a quota than when no quota was implemented.
Overview of the Current Research
This research consists of five studies (total N = 761). In all studies, we used the setting of a group discussion within a hiring committee for a (not further defined) professorship. Participating in a hiring committee and selecting a person for a professorship is considered a male-typed (i.e., high-profile) task in this research. Based on the helpful suggestion by an anonymous reviewer, the perception of a task as male-typed has been validated as part of Study 5. There, we asked participants to rate the degree to which they thought competence (i.e., competent, competitive, independent) and warmth (i.e., good-natured, likable, warm; Asbrock, 2010) were central attributes for participating in a hiring committee for a professorship. Results reveal that participants considered competence (M = 3.95, SD = 0.75) significantly more central for this task compared with warmth (M = 2.94, SD = 0.97), t(102) = 9.11, p < .001. This confirms that the task we chose for our research can indeed be considered male-typed. In addition, Expectation States Theory predicts that, in mixed-sex groups, men will have a (smaller) advantage over women regarding performance evaluations even in gender-neutral tasks (Correll & Ridgeway, 2003).
In Study 1, participants (n = 170) exchanged preformulated messages in a group discussion via chat. The purpose of this group discussion was to solve a hidden profile (i.e., identify the most suitable applicant for a professorship). As a measure of performance, participants rated the relevance of the information provided and the helpfulness of one’s discussion contribution to solve the hidden profile. As a potential moderator of the effect, we varied whether the selection procedure for entering the group was based on a quota rule (i.e., women were selected to achieve a predetermined share of female group members) versus no quota (i.e., women were selected simply as new members of a group without further explanation). In Study 2 (n = 301), we controlled for the content of the sent information as well as other aspects of the discussion, such as the number of messages sent. The performance measures were the same as in Study 1. Study 3 (n = 70) employed a more realistic scenario: rather than preformulated messages, a pre-rated transcript of an open chat was used. Furthermore, with perceived competence (Asbrock, 2010; Danner, 2014) of the different group members, we used a different criterion to assess the results’ generalizability to different performance-related outcomes. In contrast to Studies 1 to 3, in Study 4 (n = 117), group members’ gender was communicated by pre-rated first names to make gender less salient and the gender composition in the group was equal (two men and two women). In Study 5 (n = 103), we again used first names and varied the number of men (two vs. three) and women (two vs. three) in a group to test whether results are robust to these changes. The data, analysis scripts, and materials for all studies are openly available on the Open Science Framework (OSF; https://osf.io/vk5p6/). The hypotheses for all but one study (Study 2) were preregistered. Thus, one-sided tests were conducted for the directed hypotheses of all studies with the exception of Study 2.
Study 1
In Study 1, participants solved a variant of a hidden profile task (Stasser & Titus, 1985) in groups of five. They were asked to imagine that they are members of a hiring committee and that their task is to select the most suitable applicant for a (not further defined) professorship position (i.e., the hidden profile). In a computerized chat, participants could exchange preformulated information on different applicants. After the chat phase, participants were asked to evaluate the relevance of each preformulated message. In addition, they rated how helpful the information provided by each group member was for the selection of the best applicant. These evaluations served as a measure of perceived task performance. After this first evaluation, participants voted for an applicant followed by an open chat phase, where participants could chat freely with each other and reconsider their decision. It was expected in Hypothesis 1 (H1) that the information provided by women would be evaluated as less helpful by other group members compared with the information provided by men.
As a potential moderator, we varied whether or not female participants were part of the group due to a quota rule (i.e., 40% women in the group). It was expected in Hypothesis 2 (H2) that women would be evaluated especially negative when they are selected based on a quota rule (interaction between group members’ gender and selection procedure). These hypotheses were preregistered with the OSF (https://osf.io/9v6zq). 3
Method
Participants and design
In total, 170 participants, mostly students at the University of Bonn, Germany (age: M = 24.26 years, SD = 6.76 years, 40% female) with heterogeneous fields of study, were recruited via the online recruitment tool ORSEE (Greiner, 2015). 4 In addition, 31 male participants were recruited who were not selected (whether a male participant was selected or not was randomly determined) during the selection procedure (quota vs. no quota) and thus did not enter the analyses (for further information, see below). We randomly assigned participants to the quota or the no quota condition, resulting in a 2 (participant gender: female vs. male) × 2 (group member’s gender: female vs. male) × 2 (selection procedure: quota vs. no quota) mixed design. The study was conducted in an on-site lab. Each experimental session was scheduled for 90 min. Participants’ payments ranged from 10 to 21 Euros (approx. US$11–23). The experiment was programmed in z-Tree (Fischbacher, 2007). The sample size was determined a priori for the effect of the quota rule on group performance (i.e., voting behavior) based on previous research (Dorrough et al., 2016). As this effect is not the focus of the present article, we additionally conducted sensitivity analyses for H1 and H2. These analyses were based on a repeated measurement mixed analysis of variance (ANOVA; H1: within factors; H2: within-between interaction) as the closest pragmatic approximation of the cluster corrected regression we conducted. These analyses show that we can detect small to medium effects of d = .12 (H1) and d = .20 (H2) with a power of 80% with the available sample size.
Materials and procedure
Before starting the experiment, participants were randomly assigned to separate cubicles. The experiment was divided into different stages that will be outlined in the following.
Stage 1: Initial discrimination
Because quota rules are implemented to counteract the underrepresentation of specific social groups in high-status positions, Stage 1 was implemented to create the necessary precondition for the experimental manipulation (i.e., quota vs. no quota) by introducing initial discrimination of women. Specifically, female participants worked on an individual task (adaptation of Brickenkamp, 2002), where they had to correctly identify specific letters and received an effort dependent payment of maximum 3 Euros (approx. US$3.50). In contrast, male participants were randomly assigned either to the individual task or a group task in three-person groups. In those groups, they shared recommendations in a preformulated chat about which items to keep on a sinking ship (adaptation from assessment center tasks). This group task for male participants was implemented to reflect real-life conditions of quota rules that introduce women to preexisting male-composed groups. Men received a fixed payment of 3 Euros (approx. US$3.50) for working on the task. Thus, female participants were discriminated against with regard to payment for the initial task. All participants were informed about the terms of payment and that the assignment to the different tasks was based on gender.
Stage 2: Selection procedure (quota vs. no quota)
The initial discrimination allowed us to implement a quota rule after Stage 1. Specifically, in the quota condition, all group members (former and new) were informed that gender was the basis for selection from the individual task to the group task. Specifically, women were preferentially selected to join the three-person all-male groups of Stage 1, whereas men who previously worked on the individual task were sent home. We informed participants that the quota rule was introduced to achieve a 40% share of women in a group. In the no quota condition, participants did not receive this information. Instead, they were only informed that there are two new (female) members coming to the group who previously worked on an individual task. In total, 31 male participants who started with the individual task were not selected in Stage 2 and left the lab.
Stage 3: Group discussion via preformulated messages
In the newly formed five-person groups, participants worked on the hidden profile task. In each experimental session, two five-person-groups were present. Thus, participants were not able to identify with whom they interacted in a group. Each five-person group consisted of three men and two women. Before interacting with each other, participants read detailed instructions on the hidden profile task and received individual information on six different applicants for a professorship position. For each applicant, they received five pieces of information relevant to professorships (e.g., Applicant A scored above average in international cooperation; Applicant B scored below average in teaching experience). When all individual information available to each group member was combined, one applicant was objectively superior to the other applicants. However, based on the individual information, each group member would prefer a different applicant (i.e., full dissent; Schulz-Hardt et al., 2006). Participants were informed that all group members only received part of the available information and that they needed to combine the individual information within their group to solve the hidden profile, which is to select the most suitable applicant. Thus, the hidden profile has the advantage that participants are incentivized to share and attend to all pieces of information to solve the task correctly. Furthermore, we told participants that the hiring committee agreed to consider all pieces of information equally; for example, teaching experience should be considered to the same extent as international cooperation. After participants familiarized themselves with the task, they could share individual information with other group members (see Figure 1).

Information sharing via chat.
The group members were assigned identification numbers ranging from 101 to 105. In the upper-left corner of the screen, they saw all group members’ assigned numbers and gender. The assignment of numbers was explained before participants started the chat. By clicking on a piece of information, a preformulated message was sent, which always had the same wording (i.e., applicant [A–F] scores [above, below, on] average with regard to [stay abroad, teaching experience, supervision of theses, English skills, publications, international cooperation, funding acquisition, interdisciplinary orientation, cooperation in faculty]). With this design feature and the fact that participants were informed that all criteria were equally important, we excluded actual differences between men and women concerning communication style (see introduction). Examples for these preformulated messages are as follows: “Person 101 is writing: applicant A scores above average with regard to research stays abroad” or “Person 102 is writing: applicant B scores on average with regard to international cooperation.” Participants were informed that all group members could share information in the same way.
The chat phase lasted 10 min. Thereafter, participants evaluated the relevance of each preformulated message as well as the perceived helpfulness of each group member’s contribution to the group discussion (see below). Thereafter, participants voted for an applicant. If the majority of group members voted for the correct applicant, the hidden profile was solved. Finally, participants could chat with each other in an open chat without preformulated messages and could reconsider their decision, that is, they voted again on the best applicant. In addition to their fixed experimental payment and the first task’s payment, participants received a bonus of 5 Euros (approx. US$5.80) if they correctly solved the hidden profile. This means that participants had an incentive to value and integrate all shared information. Participants were informed of the outcome of the vote and their payment amount.
Measures
As a measure of performance, participants evaluated the relevance (1 = “irrelevant” to 5 = “very relevant”) of each preformulated message sent via chat during the group discussion. Furthermore, they indicated the perceived helpfulness of each group member’s contribution to the group discussion (1 = “not at all helpful” to 7 = “very helpful”) for identifying the best applicant.
Results
On average, participants rated the contribution by male group members (M = 4.32, SD = 1.63) as more helpful for solving the hidden profile than the contribution by female group members (M = 4.07, SD = 1.65).
An ordinary least squares (OLS) regression predicting perceived helpfulness by group member gender, selection procedure, and the interaction between information provider gender and selection procedure revealed a significant evaluation bias (Table 1, Model 1) in line with H1. No moderating effect of selection procedure was observed when predicting overall helpfulness (Table 1, Model 1); thus, we find no support for H2. Standard errors were clustered at the individual level to account for dependencies in error terms due to repeated measurements. 5 When including participant gender and age in this analysis, the main effect of the gender of information provider is no longer significant, β = .03, 95% confidence interval (CI) = [–0.03, 0.10], t(179) = 1.13, p = .130. Descriptively, women rated male (M = 4.55, SD = 1.58) and female (M = 4.61, SD = 1.56) group members more similarly than did men (female group members: M = 3.88, SD = 1.64; male group members: M = 4.09, SD = 1.65). However, the interaction between participant and group member gender is not significant, β = .04, 95% CI = [–0.16, 0.95], t(179) = 1.41, p = .161. 6
Evaluation Bias in Studies 1 to 5.
Notes. The table shows β coefficients and 95% CIs for β. t statistics in parentheses. CI = confidence interval.
p < .10. *p < .05. **p < .01. ***p < .001.
Although messages were preformulated, thus ensuring that male and female participants did not formulate messages differently, they could potentially differ with regard to additional aspects of the discussion. We found no significant differences when comparing the behavior of male and female participants with respect to the following aspects: (a) the amount of information sent, (b) the number of repeated information sharing (sending a piece of information that has been shared before by a different group member), 7 and (c) whether a person ended a discussion. 8 Women were, however, less likely than men to start the discussion (female: 10%, male: 26%, z = −2.58, p = .010). This difference does not explain the main effect, which remains stable, β = .075, 95% CI = [0.009, 0.141], t(179) = 2.25, p = .013, when controlling for these additional aspects of information sharing behavior. 9
Interestingly, the relevance of single preformulated messages was rated similarly for male and female group members (i.e., information providers), β = .005, 95% CI = [–0.041, 0.050], t(179) = 0.20, p = .422. 10
Discussion
Results of Study 1 show that women’s performance in a group discussion (as measured by the perceived helpfulness of their contribution to solving the task) was evaluated more negatively than men’s. This is irrespective of whether women joined the group based on a quota rule or simply as new members. Although participants could only share preformulated messages, and thus men and women did not differ concerning the content of the information sent, men’s contribution was seen as more helpful to identify the most suitable applicant for a professorship in a hidden profile task. This result holds when we control for several other aspects of the contribution, such as the amount of information a person provided or whether a person started the discussion. Thus, results of Study 1 provide initial evidence that men and women are evaluated differently in group tasks such as group discussions. However, although several aspects of the contribution were controlled for, it is still possible that participants were influenced by other characteristics of a group member’s contribution (e.g., sequence effects other than starting or ending a discussion). In Study 2, we solve this potential problem by keeping some aspects of the information sharing constant while counterbalancing others across different material conditions (e.g., who started the discussion). One further potential limitation of Study 1 is that we cannot rule out ingroup effects in the evaluation due to the fact that male participants had already worked together as a team before women entered their group. 11 We found that the effect of the negative overall evaluation was mainly driven by male participants and that the effect is no longer significant when including participant gender as an additional control. This could be the nature of the effect but could also be due to potential ingroup effects, as men had already worked together on an unrelated task in groups of three and women were later added to those existing groups. In Study 2, all group members were evaluated by external observers as opposed to their fellow group members to exclude this potential confound in Study 1. With this new design feature, we can also rule out potential effects of whether a person mainly shared information that supported a participant’s current preference for an applicant or other strategic concerns regarding voting and experimental payment. In Study 1, we did not find the hypothesized moderating effect of the selection procedure. Specifically, in contrast to what we had expected a priori, the application of a quota rule did not increase the gender bias in performance evaluations. In the experimental instructions, we introduced the quota rule with only one short sentence, and it is conceivable that participants did not have this information present when engaging in the (rather demanding) task. With Study 2, we can re-address the potential effects of quota rules on evaluators that are, in contrast to Study 1, not performing the task themselves.
It must be noted that in contrast to the overall contribution, the relevance of each preformulated message was not rated differently for male and female group members. Previous research has shown that preexisting stereotypes influence the evaluation, especially when the situation is ambiguous (e.g., Heilman & Haynes, 2008; Macrae & Bodenhausen, 2000). A separate evaluation of single pieces of information is less ambiguous and leaves less room for such biases. Relatedly, standardization and structure—such as evaluating single pieces of information and aggregating them to an overall evaluation, rather than forming an intuitive global impression—have been shown to increase the validity of a selection procedure (Schmidt & Hunter, 1998). Before making strong conclusions, this discrepancy between the evaluation of single contributions and overall performance shall be revisited in Study 2.
Study 2
In Study 2, we described the group task of Study 1 to a different participant sample. Participants were provided with the information on the different applicants available to one former group member who was either male or female. In addition, they were provided with eight pieces of information that person had shared with the other group members. They were also informed about the initial discrimination and about one of the two selection procedures based on which women joined the three-person groups. As in Study 1, participants had to evaluate the relevance of each preformulated message as well as how helpful they perceived the group members’ contribution to be for identifying the most suitable applicant. It was again expected that as per H1 the information provided by women would be evaluated as less helpful by other group members compared with the information provided by men. In H2, this effect was expected to be stronger when women were selected based on quota. The hypotheses for Study 2 were not preregistered.
Method
Participants and design
In total, 301 student and nonstudent participants (age: M = 34.87 years, SD = 12.32 years, 66% female) with a broader age range (18–68 years of age) as compared with Study 1 were recruited from the Decision Lab subject pool of the University of Hagen, Germany via the online recruitment tool ORSEE (Greiner, 2015). The study was conducted online and lasted approximately 15 min. Participants received a fixed amount of 5 Euros (approx. US$5.80) for their participation. The experiment was programmed in Unipark (QuestBack©). As Study 2 has not been preregistered, we report the results of a sensitivity analysis here, which revealed that with the available sample size, small to medium effects of d = .28 can be detected with a power of 80% assuming a multiple regression analysis. For this study, a 2 (participant gender: female vs. male) × 2 (group member’s gender: female vs. male) × 2 (selection procedure: quota vs. no quota) × 5 (material condition: 1–5) between-participant design was used.
Materials and procedure
Participants in Study 2 were informed about the procedure of Study 1. They learned that participants in Study 1 worked on different tasks (i.e., initial discrimination) before the selection procedure was implemented. They were also informed that these participants had to select the most suitable out of six applicants for a professorship and that they received different pieces of information on each applicant. Thereafter, participants in Study 2 were presented with the information that one group member in Study 1 received about the applicants (see Figure 2). In addition, they saw eight preformulated messages this person had shared with his or her group members.

Information on different applicants that participants in Study 2 were provided with.
It was varied between participants whether the group member to be evaluated was male or female and whether a selection based on quota versus no quota was implemented. Furthermore, and in contrast to Study 1, we kept the amount of preformulated messages constant (the same messages as in Study 1 have been used). Whether a person started or ended the discussion was no longer relevant, as participants were only provided with the messages sent by one group member. Between material conditions, we varied the amount of positive, neutral, and negative messages, the number of messages sent about the different applicants, and the sequence of messages.
Measures
As a measure of performance, participants again rated the relevance of each preformulated message (1 = “not relevant” to 5 = “relevant”) as well as the helpfulness (1 = “not at all helpful” to 7 = “very helpful”) of the overall contribution to finding the most suitable applicant.
Results
Descriptively, external observers rated the overall contribution by female group members as less helpful (M = 3.41, SD = 0.94) than the overall contribution by male group members (M = 3.65, SD = 1.04), a result marginally significant in an OLS regression (Table 1, Model 2). Results remain stable when controlling for participant age and gender, β = .11, t(295) = 1.80, 95% CI = [–0.010, 0.220], p = .072. The same is true when we additionally control for material condition, β = .10, t(294) = 1.78, 95% CI = [–0.011, 0.217], p = .076.
We do not find the expected interaction between the selection procedure and the group member’s gender (Table 1, Model 2). Furthermore, as in Study 1, no difference for single preformulated messages is found, β = .04, t(300) = 1.18, 95% CI = [–0.030, 0.115], p = .240 in an OLS regression predicting the relevance of single preformulated messages by group member’s gender, selection procedure, and the interaction between group member’s gender and selection procedure with standard errors clustered at the participant level.
Discussion
The results of Study 2 provide additional support for the assumption that, in group discussions, the performance of women (as measured by the helpfulness of their contribution to finding the most suitable applicant) is evaluated more poorly than the performance of men. This effect emerged even though neither the content of preformulated messages nor their perceived relevance differed between male and female group members. In addition, in Study 2, we kept further aspects of the discussion constant and counterbalanced others across material conditions. Furthermore, in contrast to Study 1, participants who were not part of the group evaluated the group members, excluding potential ingroup effects, strategic concerns, or unintended effects due to preference consistent information. In Study 2, again, no influence of different selection procedures could be found. Furthermore, Study 2 provides additional support for the idea that the information based on single, specific criteria (i.e., evaluation of single preformulated messages) leaves no room for evaluation biases due to reduced ambiguity, which makes the result for overall performance particularly remarkable. The subsequent studies of the present research will focus on the evaluation of overall performance. Future research could investigate different evaluation procedures for group discussions in more detail.
Study 2 has at least two limitations: Preformulated messages have the advantage that the information shared by men and women does not actually differ in terms of content. At the same time, participants might have perceived the presented group discussion as being highly artificial and somewhat unrealistic. It is unclear whether the results also apply to more realistic and thus more externally valid situations where men and women can discuss with each other more freely. Therefore, in Study 3, participants evaluated the group member’s performance based on a transcript of an open chat of Study 1. Another potential limitation of Study 2 (and Study 1) is that so far, we only used perceived helpfulness of the contribution as the dependent measure of performance and cannot make any claims regarding generalizability to different performance-related outcomes. In Study 3, we used a standard measure of perceived competence (Asbrock, 2010) as well as a measure of perceived competence in the work context (Danner, 2014) instead of perceived helpfulness of the contribution.
Study 3
In Study 1, after the voting stage, participants could freely chat with each other for 10 min to potentially reconsider their decision (i.e., voting). In Study 3, one transcript of such an open group chat was used. Specifically, participants in Study 3 were asked to evaluate the perceived competence of the different group members based on the chat transcript of one five-person group from Study 1. It was expected that as per H1 women’s performance in the group discussion (as measured by their perceived competence) would be evaluated more negatively than men’s. Furthermore, information regarding initial discrimination and the selection procedure was again varied to test whether the expected quota effect can be found in a more realistic scenario. It was assumed that as per H2 the effect postulated in H1 would increase in a quota condition compared with a no quota condition. The hypotheses for Study 3 were preregistered with OSF (https://osf.io/gswp8).
Method
Participants and design
In total, 70 participants (age: M = 33.16 years, SD = 11.76 years, 69% female) were recruited from the Decision Lab subject pool of the University of Hagen, Germany via the online recruitment tool ORSEE (Greiner, 2015). The study was conducted online, lasted about 25 min, and participants received a fixed amount of 5 Euros (approx. US$5.80). The experiment was programmed in Unipark (QuestBack©). The sample size was estimated based on a finding by Heilman et al. (1997) that women, who were selected by the mechanism of a women’s quota, were evaluated more poorly (η² ~ .06). A sample of 54 participants was needed based on a repeated measurement mixed ANOVA (within-between interaction; see Study 1) to achieve a power of 95% to detect effects of that size. Due to the normal distribution assumption, we targeted 60 participants, that is, 30 persons per condition. Finally, due to a technical error, 70 participants completed the study. A sensitivity analysis revealed that medium effects of d = .20 can be detected for H1 with a power of 80%, assuming a repeated measurement mixed ANOVA (within factors). We employed a 2 (participant gender: female vs. male) × 2 (group member’s gender: female vs. male) × 2 (selection procedure: quota vs. no quota) mixed design. The second factor was varied within participants.
Materials and procedure
One purpose of Study 3 was to increase external validity by making the presented group discussion more realistic. However, we aimed to make sure that the presented male versus female group members again did not differ with regard to their actual performance. By doing so, we maintain the advantage of the present research of being able to isolate a potential evaluation bias. One gender-neutral chat transcript needed to be selected for Study 3 to exclude actual differences between men and women regarding their contribution to the group chat. Therefore, three chat transcripts of the open chat of Study 1 were first preselected by one of the co-authors based on the following criteria: men and women shared a similar amount of information; the chat included information relevant for the goal of solving the task (i.e., information about the applicants); and there was no immediately recognizable difference between men and women concerning the content of the chat messages. From these three chat transcripts, task-irrelevant information was deleted (e.g., “☺”). Furthermore, in some cases, information was added to increase the understandability of the transcript (e.g., abbreviations were spelled out). These changes were applied to all group members in a similar fashion. After this preselection, a pilot study was conducted to verify whether and which of the three preselected chats indeed fulfilled the criterion of no gender difference in perceived competence when the gender of group member is not revealed. In the pilot study, neither information on initial discrimination nor the selection procedure was given. In total, 50 participants (age: M = 43.6 years, SD = 13.44 years; 84% female), who were blind to the hypotheses, rated the different chat transcripts in randomized order. Participants rated each group member on perceived work-related competence (five-item scale by Danner, 2014; e.g., How do you evaluate the quality of his or her work: 1 = “very bad” to 6 = “very good”). We aimed to select the most gender-neutral transcript in that the underlying gender was irrelevant for the evaluation (i.e., no actual gender differences in performance exist). In one of the three pre-tested transcripts, female group members were rated better, in a second transcript, male group members were evaluated better, and in the third transcript (which was then selected for Study 3), no differences between the underlying gender emerged, β = .02, t(49) = .40, 95% CI = [–0.09, 0.13], p = .69. The resulting transcript stemmed from a five-person group that did not correctly solve the hidden profile after the preformulated chat in Study 1. Thus, the information provided in the open chat was still meaningful and goal-oriented.
The procedure of Study 3 was similar to the procedure of Study 2. Participants were again informed about initial discrimination and the selection procedure (quota vs. no quota). They were then provided with the information about applicants for a professorship position that a former participant of Study 1 received (Figure 2). Finally, participants were presented with the chat transcript. A brief extract of the chat transcript can be seen in Figure 3.

A brief extract of the chat transcript.
The study was conducted in German. Therefore, group members’ (former participant’) gender was indicated by the female and male version of the noun “participant” (“Teilnehmer” for a male participant, “Teilnehmerin” for a female participant).
Measures
In Study 3, we used different measures of performance to assess the results’ generalizability to different performance-related outcomes. Participants rated each group member on perceived competence using the German short version (Asbrock, 2010) of the stereotype content scale (Fiske et al., 2002). The competence subscale consisted of three items (i.e., “Do you consider group member 103 [competent, competitive, independent]?”: 1 = “not at all” to 5 = “fully”). 12 Furthermore, participants rated all group members on work-related competence (see above, Danner, 2014). Ratings were provided on separate pages, and as a basis for the evaluation, participants received a summary of the respective group member’s messages.
Results
Both competence scales used in Study 3 show reasonable reliability and variability: Asbrock (2010): Cronbach’s alpha = .83, M = 3.16, SD = 0.93; Danner (2014): Cronbach’s alpha = .93, M = 3.90, SD = 1.03. Descriptively on the competence dimension of the Asbrock scale, participants rated male group members (M = 3.26, SD = 0.88) more positively than female group members (M = 2.99, SD = 0.98). The same is true for the Danner scale (male group members: M = 3.97, SD = 0.99; female group members: M = 3.79, SD = 1.09). This result (Asbrock scale) is statistically confirmed in an OLS regression (Table 1, Model 3) with clustered standard errors at the participant level. For the additional work-related competence measure by Danner (2014), conclusions are the same, β = .08, t(69) = 1.89, 95% CI = [–0.00, 0.17], p = .032. The effect is also observed when rerunning the model presented in Table 1 while additionally controlling for participant age and gender, β = .14, t(69) = 2.84, 95% CI = [0.04, 0.24], p = .003.
As in the other studies, no moderating effect of selection procedure was found (Table 1, Model 3).
Discussion
Results of Study 3 show that female performance in a group discussion (as measured by two different competence scales; Asbrock, 2010; Danner, 2014) is evaluated more poorly than male performance. Using a different, more realistic discussion form and different outcome measures, Study 3 replicates the findings of Studies 1 and 2. Again, no moderating effect of selection procedure (quota vs. no quota) was found.
Although Studies 1 to 3 seem to consistently support the existence of an evaluation bias in group discussions, there is one alternative explanation for these results: To investigate a potential moderating effect of selection procedure, women always entered preexisting three-person groups. Thus, the results could also be due to a newcomer effect in that new group members are evaluated more poorly than incumbent group members. To exclude this alternative explanation and because the expected moderation was not present in any of the previous studies, Study 4 exclusively tested whether men and women are evaluated differently in group discussions. Therefore, no selection procedure took place, and no information on initial discrimination or selection was given. Furthermore, to test the different selection procedures, in Studies 1 to 3, groups always consisted of three male and two female group members. In Study 4, where we used the same chat transcript as in Study 3, this ratio was balanced in that two group members were males, two group members were females, and one group member was unknown. Doing so allowed us to exclude the possibility that minority members are evaluated more poorly as compared with majority members. Finally, to make the scenarios again more realistic and less gender salient, group members were presented using male and female first names instead of numbers combined with gender information.
Study 4
In Study 4, the same chat transcript was used as in Study 3, and again participants evaluated performance in the group discussion in the form of perceived competence of the group members. It was expected that as per H1 women would be evaluated more negatively (i.e., less competent) than men. This hypothesis was preregistered on the OSF (https://osf.io/97j3n). 13
Method
Participants and design
In total, 117 participants (age: M = 25.90 years, SD = 9.77 years, 71% female) were recruited via email and social media. The study was conducted online, lasted approximately 30 min, and participants received course credit for their participation. The experiment was programmed in Qualtrics (Qualtrics, Provo, UT). The sample size was calculated a priori based on the evaluation bias observed in Study 3 (η² = .100) regarding perceived competence and a power of 95% in a repeated measurement ANOVA as the closest pragmatic approximation to the cluster corrected regression we intended to run.
Materials and procedure
Study 4 followed the protocol of Study 3 with the following changes: No reference to any selection procedure was given. Furthermore, first names were used instead of numbers and gender information. We chose typical male and female names 14 that are common in Germany and rated similarly regarding perceived competence in Nett et al. (2020). The following names were selected: Pascal (male, perceived competence in Nett et al., 2020: 3.6), Erkan (male, perceived competence: 3.7), Nadine (female, perceived competence: 3.7), and Meryem (female, perceived competence: 3.7). As the chat originally consisted of five group members, a fifth person was presented as unidentified (N.N.) to induce gender parity in the group. To create a more formal context, the initial of each identified person’s last name was given (i.e., Pascal A., Erkan K., Nadine M., Meryem S.). In sum, a 2 (participant gender: female vs. male) × 2 (group member’s gender: female vs. male) mixed design was employed. 15
Measures
After reading the chat transcript, participants rated each group member on the competence items by Asbrock (2010; i.e., competent, competitive, independent), which were also used in Study 3.
Results
Evaluations of the unidentified person were excluded from the analyses. The competence scale used in Study 4 shows reasonable reliability and variability (Cronbach’s α = .79, M = 3.21, SD = 0.88). Descriptively, male group members (M = 3.41, SD = 0.83) were evaluated better than female group members (M = 3.01, SD = 0.88). This result is supported by an OLS regression with clustered standard errors (Table 1, Model 4). Results are robust when controlling for participant’s age and gender, β = .23, t(116) = 4.74, 95% CI = [0.13, 0.32], p < .001. 16
Discussion
In sum, the results of Study 4 replicate the other three studies’ findings concerning differences in how men and women are evaluated in group discussions. Specifically, female group members’ performance (as measured by perceived competence) was evaluated more poorly than male group members’ performance. Because Study 4 induced gender parity in the group and only incumbent group members were presented, we can exclude alternative explanations of pure minority or newcomer effects. So far, our studies only concerned groups with gender parity or female group members with minority status. Study 5 serves as a conceptual replication of Study 4 using different representations of female group members in a group to determine whether results are robust toward these changes.
Study 5
In Study 5, the same chat transcript was used as in Studies 3 and 4. Using the same items from Asbrock (2010), participants evaluated performance in the group discussion in the form of perceived competence of the group members. It was expected that as per H1 women would be evaluated more negatively (i.e., less competent) than men. This hypothesis was preregistered on the OSF (https://osf.io/qwhxd). 17
Method
Participants and design
In total, 103 participants (age: M = 28.79 years, SD = 8.53 years, 73% female) 18 were recruited from the Decision Lab subject pool of the University of Cologne, Germany via the online recruitment tool Hroot (Bock et al., 2014). The study was conducted online and lasted approximately 30 min. Participants received a fixed payment of 5 Euros (approx. US$5.80) for their participation. The experiment was programmed in Unipark (QuestBack©). Sample size was determined a priori based on a repeated measurement ANOVA, assuming an effect of d = .40 (based on the weighted effect sizes of Studies 3 and 4), and statistical power of 95%. The study employed a 2 (participant gender: female vs. male) × 2 (group member’s gender: female vs. male) × 2 (percentage of female group members: 2/3 vs. 1/2 vs. 3/2 female) mixed design with the second factor being varied within participants.
Materials and procedure
Study 5 followed the protocol of Study 4 with the following changes: In the same manner as in Study 4, different common, German, male and female 19 first names were selected (again from Nett et al., 2020): Clemens (male, perceived competence in Nett et al., 2020: 5.1), Johannes (male, perceived competence: 4.8), Sören (male, perceived competence: 4.9), Viktoria (female, perceived competence: 4.9), Charlotte (female, perceived competence: 5.0), and Nora (female, perceived competence: 4.9). Again, the initial of each identified person’s last name was given to create a more formal context. Participants were randomly assigned to one of the following three conditions: (a) majority male (Clemens M., Sören K., Johannes W., Charlotte L., Viktoria P.), (b) gender parity (Clemens M., Sören K., N.N., Charlotte L., Viktoria P.), and (c) majority female (Clemens M., Sören K., Nora S., Charlotte L., Viktoria P.). The pretested chat transcript was reanalyzed to ensure that underlying gender affected none of the conditions.
Measures
After reading the chat transcript, participants rated each group member on the competence items (i.e., competent, competitive, independent; Asbrock, 2010) used in Studies 3 and 4.
Results and Discussion
Evaluations of the unidentified person were excluded from the analysis. The competence scale used in Study 5 shows reasonable reliability and variability (Cronbach’s α = .82, M = 3.10, SD = 0.88). Descriptively, male group members (M = 3.21, SD = 0.88) were again evaluated better than female group members (M = 2.99, SD = 0.87) with regard to performance in a group discussion (as measured by perceived competence). An OLS regression with clustered standard errors supports this result (Table 1, Model 5). The main effect is robust to additionally controlling for evaluators’ age and gender, β = .12, t(102) = 2.65, 95% CI = [0.03, 0.21], p = .005. The same is true when controlling for condition (i.e., percentage of female group members in the group), β = .12, t(102) = 2.53, 95% CI = [0.03, 0.22], p = .007. Results of Study 5 corroborate the findings of the previous studies, again showing that the performance of women as compared with men is evaluated more poorly in group discussions.
Combined Evidence for the Effect
A mini meta-analysis including all five studies was conducted to strengthen the conclusions of this research (see Goh et al., 2016, for a discussion and the exact procedure). Standardized regression coefficients (which correspond to Pearson’s r in the case of one predictor) of the models reported in Table 1 were used as effect sizes (Study 1: β/r = .08, Study 2: β/r = .10, Study 3: β/r = .14, Study 4: β/r = .23, Study 5: β/r = .12). Analyses reveal a weighted mean effect size 20 of rM = 0.12. Thus, evidence across all studies point to a small effect of group member’s gender on perceived performance in group discussions. A Stouffer’s Z test (see Goh et al., 2016) results in a summary p value of p < .001 for all the studies involved in this research. Using p values from OLS regressions additionally controlling for age and gender (resulting in a nonsignificant effect in Study 1), the summary p value is again p < .001. 21
Overall Discussion
Women are still underrepresented in many male-associated areas such as high-profile positions. One reason for this imbalance is how men and women are evaluated: Women are evaluated more poorly than men in many different contexts based on their individual task performance and individual competencies (e.g., Heilman et al., 2015). Although group tasks are an important aspect of every day (work) life, research on how men and women are perceived in this kind of situation is scarce. Based on Expectation States Theory and similar theoretical accounts (Berger et al., 1972; Correll & Ridgeway, 2003), it could be assumed that in the absence of specific information on task competence, salient status characteristics such as gender determine how group members are evaluated. Members with a low state of a status characteristic such as competence (i.e., women) are assumed to have less ability at male-typed but also gender-neutral tasks (Correll & Ridgeway, 2003). As a consequence, their contributions within the group are less positively evaluated by the group and by outside observers (e.g., Correll & Ridgeway, 2003; Foschi, 2000). Empirical results are in line with these assumptions showing that women’s competence in group tasks is indeed evaluated more poorly as compared with men’s (e.g., Ridgeway, 1982; Wood & Karten, 1986). However, in previous studies, group members typically interacted with each other face-to-face or actually demonstrated different kinds of behavior involving potential confounds.
In the present research and in contrast to previous studies, we experimentally and/or statistically controlled for actual gender differences in behavior. Furthermore, participants in our studies interacted with each other anonymously, excluding potential confounds such as differences in physical appearance. In line with our predictions, we show in five studies (N = 761) that female performance in a group discussion is evaluated more poorly than male group members’ performance. This effect was observed for fellow group members (Study 1), outside observers (Studies 2–5), primarily student samples (Studies 1, 4, and 5), mixed samples (Studies 2 and 3), and different measures of performance (perceived helpfulness of their contribution; competence) as well as across different discussion formats (Studies 1 and 2: preformulated chat messages; Studies 3–5: open chat) and when controlling for the percentage of women in the group (Study 5).
In addition to investigating biases in evaluations of men and women, we included a potential moderator of the effect, namely the selection process (women’s quota vs. no quota). Previous research has shown that the competence of women selected based on affirmative action policies was devalued (e.g., Heilman et al., 1997). Thus, it was expected that women are evaluated even more poorly when people are informed that the selection for entering the group was based on a quota rule than when no quota was implemented. In contrast to this hypothesis, we did not find a moderating effect of selection procedure (Studies 1–3) in that women were devalued to a similar degree in both situations with a quota rule (that usually subjectively undermines the merit principle) and without. Although the studies were well-powered and the null findings were consistent, future research could re-address this potential moderator by introducing the quota rule more explicitly. In our studies, we used a rather subtle manipulation consisting of one sentence in the instructions (“to achieve a share of 40% women in the group task, two female persons were selected”). The hidden profile task that followed this information is quite complex, which is why this information was potentially not salient when indicating perceived performance after completing the task.
Another (unexpected) finding deserves coverage in future research: In Studies 1 and 2, we find that the overall contribution of men and women but not the relevance of each preformulated message was evaluated differently. This could be due to the fact that the evaluation of specific criteria leaves almost no room for ambiguity, a circumstance that usually reduces the influence of preexisting stereotypes (e.g., Heilman & Haynes, 2008; Macrae & Bodenhausen, 2000). The free chat in Studies 3 to 5 produced slightly stronger effects as compared with our Studies 1 and 2—an observation that would also support this idea, as an open chat is much less structured and thus more ambiguous as compared with a chat transcript consisting of preformulated messages. Future research could potentially investigate different levels of ambiguity in group discussions and their influence on the evaluation of men and women therein.
Our result that women’s performance as measured by their competence in the group discussion is evaluated worse than men’s is, at first sight, not in line with most recent results (e.g., Eagly et al., 2019; Koenig & Eagly, 2014). In a study including more than 30,000 participants from the United States from 1946 to 2018, Eagly and colleagues (2019) found that the direction of the competence stereotype reversed over time in that a majority of participants in 2018 perceived women to be more competent than men. This finding was explained by changing social roles in that, for example, women increasingly participate in the labor market and challenge the assumption that women, as a lower status social group, are accorded less competence (e.g., Ridgeway, 2014). Today, in both Germany and the United States, a similar percentage of women participate in the labor market (55% in Germany vs. 56.1% in the United States; http://hdr.undp.org/en/content/table-5-gender-inequality-index-gii), which is why one could have also expected positive evaluations of women regarding their competence in group discussions. However, we assessed competence in a male-dominated area where differences in perceptions are usually higher (Correll & Ridgeway, 2003). In contrast, Eagly and colleagues (2019) used data on competence in a broader, more unspecific way, which could explain differences in results. Future research could thus investigate group discussions on a neutral topic or a topic that is stereotypically associated with women (e.g., whom to choose as a babysitter) to see whether the bias diminishes or reverses.
Our results have important practical implications: Group discussions are important part of work-life, as senior managers, for example, attend on average 23 hr of meetings every week. Furthermore, group discussions are an important part of assessment centers and thus important for job entry. Although men and women did not actually differ concerning their contribution to the group task in our studies, their performance was evaluated more poorly than that of men. Thus, their chances of being selected and promoted based on their performance in group discussions could be smaller even when there is no difference in actual competencies. This finding has to be taken into account when designing measures to reduce gender imbalances in the workforce. Previous research has shown that positive information on the low-status individuals overcame negative information implied by their status characteristic (Ridgeway, 1982). Thus, providing evaluators with positive information on female group members could eliminate the biases in evaluations.
Supplemental Material
sj-docx-1-psp-10.1177_0146167221992213 – Supplemental material for Equal Performance, Different Grade: Women’s Performance in Discussion Perceived Worse Than Men’s
Supplemental material, sj-docx-1-psp-10.1177_0146167221992213 for Equal Performance, Different Grade: Women’s Performance in Discussion Perceived Worse Than Men’s by Angela R. Dorrough, Monika Leszczyńska, Sandra Werner, Lovis Schaeffer, Anna-Sophie Galley, Enis Akin, Jacqueline Bachmann, Marius Bruske, Ulla Burghardt and Franziska Simandi in Personality and Social Psychology Bulletin
Supplemental Material
sj-pdf-1-psp-10.1177_0146167221992213 – Supplemental material for Equal Performance, Different Grade: Women’s Performance in Discussion Perceived Worse Than Men’s
Supplemental material, sj-pdf-1-psp-10.1177_0146167221992213 for Equal Performance, Different Grade: Women’s Performance in Discussion Perceived Worse Than Men’s by Angela R. Dorrough, Monika Leszczyńska, Sandra Werner, Lovis Schaeffer, Anna-Sophie Galley, Enis Akin, Jacqueline Bachmann, Marius Bruske, Ulla Burghardt and Franziska Simandi in Personality and Social Psychology Bulletin
Footnotes
Acknowledgements
We thank Alexa Weiss for her thoughts and review of this work.
Authors’ Note
The third to tenth authors were students under the supervision of the first author.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Open Practices
Supplemental Material
Supplemental material is available online with this article.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
