Abstract
In a range of studies across platforms, researchers have shown that online ratings are characterized by distributions with disproportionately heavy tails. The authors of this study focus on understanding the underlying process that yields such “J-shaped” or “extreme” distributions. They propose a novel theoretical mechanism behind the emergence of J-shaped distributions: differential attrition, or the idea that potential reviewers with moderate experiences are more likely to leave the pool of active reviewers than potential reviewers with extreme experiences. The authors present an analytical model that integrates this mechanism with two extant mechanisms: differential utility and base rates. They show that although all three mechanisms can give rise to extreme distributions, only the utility-based and attrition-based mechanisms can explain the authors’ empirical observation from a large-scale field experiment that an unincentivized solicitation email from an online travel platform reduces review extremity. Subsequent analyses provide clear empirical evidence for the existence of both differential attrition and differential utility.
Online product reviews are a key source of information for consumer decisions and have been shown to have a significant causal impact on purchases (Chevalier and Mayzlin 2006; Chintagunta, Gopinath, and Venkataraman 2010). A growing body of literature examines the antecedents, content, and consequences of online reviews (see, e.g., Babic Rosario et al. 2016; Berger 2014). One interesting robust empirical finding is the disproportionate prevalence of “extreme” distributions of online reviews. That is, relative to moderate review scores, extremely negative and positive review scores are posted more often. Because the highest possible rating score is often the mode, the resulting distribution often has the shape of the letter “J.” Schoenmüller, Netzer, and Stahl (2020) document that this “J-shaped,” or extreme, distribution is pervasive in a variety of categories and platforms. 1 To explain the prevalence of these distributions, researchers have predominantly argued either for a review process with differential utility, in which consumers derive greater utility from sharing information about experiences they assess as being extremely positive or negative (vs. moderate) (Anderson 1998), or—to a lesser extent—for a process with differential base rates, in which the underlying distribution of experiences (i.e., their base rates) is itself disproportionately extreme (Hu, Pavlou, and Zhang 2009).
In this article, we (1) propose a novel theoretical mechanism behind the posting of reviews and, in particular, for the emergence of J-shaped distributions, (2) present an analytical model for review provision that integrates this mechanism along with the other mechanisms in the extant literature, and (3) provide empirical evidence for these explanations. Specifically, we argue and demonstrate that reviewer attrition is an integral but previously neglected aspect of the review provision process: not all consumers who at some point intend to write a review will do so in the end. For example, one survey finds that 56% of consumers at least occasionally intend to write a review for a recent experience but never end up doing so (Tomorrow Focus AG 2014). Although several causes, including lack of time, lack of motivation, or not knowing what to say, may underlie this phenomenon, the most common self-stated reason why consumers do not review is that they “forgot to do it” (Statista 2019; Tomorrow Focus AG 2014). Our proposed mechanism is based on such attrition and on the idea that forgetting rates are likely to differ for extreme versus moderate experiences. We refer to this as “differential attrition.”
Forgetting to write a review is a failure of a person's prospective memory, which stores one's “to-do” list of intended, future actions and tasks (Einstein and McDaniel 1990). Failure to implement these tasks has been characterized in the psychology literature as “forgetting to remember,” and is often unintentional (Bowden, Visser, and Loft 2017). Extant studies demonstrate that such forgetting is reduced through trigger cues that are strongly associated with the planned action (Dismukes 2012). In the context of future review provision, such cues are, for example, thoughts about the underlying consumption experience. This follows from memory network models (Collins and Loftus 1975) in which activation of a focal memory node (e.g., the experience) results in activation of associated nodes (e.g., writing a review for the experience). As people think more frequently about emotional (positive and negative) than neutral events (Walker et al. 2009), extreme experiences will be associated with more frequent node activation and, thus, with less forgetting to write a review. In summary, we expect lower reviewer attrition for extreme versus moderate consumption experiences, leading us to propose such differential attrition rates as an explanation for extremity bias in online reviews. 2
Our analysis proceeds in three steps. First, we describe a field experiment we designed and implemented in partnership with a European travel portal. Our manipulation took the form of primarily unincentivized email reminders sent to travelers at different and randomly assigned points in time (starting either one, two, five, or nine days after the end of their vacation). Exogenously imposed shocks to the reviewing process, generally via emails, that are either incentivized (Burtch, Bapna, and Griskevicius 2018; Fradkin, Grewal, and Holtz 2018; Schoenmüller, Netzer, and Stahl 2020) or unincentivized (Burtch, Bapna, and Griskevicius 2018; Fradkin, Grewal, and Holtz 2018; Karaman 2021; Schoenmüller, Netzer, and Stahl 2020), are a common methodological approach in the online review literature. What distinguishes our approach from the previous studies is that, because of our focus on attrition, our design enables us to compare the distribution of reviews following an unincentivized reminder email with a control group that did not yet receive such a reminder. Because attrition is a dynamically evolving process, for example, comparisons of distributions before versus after a reminder (Karaman 2021; Schoenmüller, Netzer, and Stahl 2020) would not allow one to identify our proposed mechanism or distinguish it from existing theories.
At a high level, the estimated treatment effect of our randomized reminders is similar to those found in the literature in that they yield more moderate distributions of reviews. The effect sizes are considerable: consumers wrote 10% fewer extreme reviews in the treatment conditions, in which a reminder had already been sent, relative to the control conditions, in which the reminder had not yet been sent. Importantly, and different from existing studies (e.g., Schoenmüller, Netzer, and Stahl 2020), this comparison holds the number of elapsed days since the end of travel constant across conditions. Accordingly, our results cannot be explained by previous work that suggests that extreme experiences for hedonic goods, such as hotels, may become more moderate over time (e.g., Moore 2012). Moreover, unlike existing studies, we do not ascribe all of the credit for the moderating effects to differences in utility for posting extreme versus moderate reviews. Instead, we argue that this is just one of several potential explanations. In the remainder of the article, we attempt to disentangle each to establish their explanatory power.
Second, we develop a simple dynamic analytical model that allows us to capture the distinct impact on review distributions of three primary mechanisms: differential base rates, differential utility, and differential attrition. At time zero, a customer completes a given experience that may be characterized as either “extreme” or “moderate” in terms of its valence. For each period thereafter, they decide whether or not to write a review. If they do, the process ends. If they do not, they may randomly exit the pool due to attrition (e.g., they forget about the review) and the process ends. If they neither write a review nor exit, the process repeats in subsequent periods until all potential reviewers have posted a review or exited via attrition. We demonstrate that any (and all) of the three mechanisms may explain the commonly observed extreme distributions, but only the utility-based and the attrition-based mechanisms imply that the observed distribution is biased, where bias is defined as the difference between the extremity of reviews and the extremity of the underlying experiences. In contrast, differences in base rates may result in extreme distributions, but because this would mirror differences in the underlying distribution of experiences, we would not characterize this as “bias.”
We then analytically examine the impact of unincentivized reminder emails sent to all customers who have not yet written a review. 3 We allow for two potential effects of the email. First, the email may temporarily “boost” the perceived utility of writing a review. Second, the email represents an external trigger cue that should remind consumers of their original intention to review, especially those who forgot and exited the active reviewer pool. We refer to this as the “attrition reversal” effect of the reminder. We show that the direction of the attrition reversal effect depends on the underlying review-generating process. That is, if the degree of differential attrition is significantly greater than the degree of differential utility, attrition reversal should decrease the extremity of the review distribution. Our main result demonstrates that (1) the net effect of the reminder depends on the relative magnitude of attrition differences and utility differences, including reminder-generated boosts, and (2) in the absence of any such differences, but allowing for differences in base rates, we would expect no impact of a review reminder on the review distribution. Given the strong estimated treatment effect we observe, this suggests that differences in base rates alone are not enough to explain our results. At least one of the other mechanisms (utility-based and/or attrition-based) must be operative. Thus, the extreme distributions of reviews reflect a bias in the review process.
Third, we develop an empirical model that leverages the exogenous variation in the data initiated by the reminders to identify the parameters associated with differential utility and differential attrition. Specifically, we form from our analytical model the conditional likelihood associated with the experimental data and estimate it via maximum conditional likelihood. Our estimates provide clear evidence for both the extant utility-based mechanism (i.e., conditional on being active, one is more likely to write a review for an extreme than for a moderate experience) and, simultaneously, for the attrition-based mechanism that we propose (i.e., conditional on being active, one is more likely to leave the pool of active reviewers in a given period if the experience was moderate than if it was extreme). Although we are not able to precisely estimate the differences, if any, in population-level proportions of extreme versus moderate experiences, these base rates play no role in our estimation approach, and thus, our results are robust to any such differential. Moreover, we are able to shed some light on the presence of differential base rates by considering a range of additional assumptions.
This article makes several important contributions. First, we advance existing knowledge by identifying a novel mechanism—attrition—that drives the commonly observed distribution of online reviews. Second, we provide a useful and general analytical framework allowing researchers to study the effects and interactions of the various underlying mechanisms behind online reviews. Third, we present a novel and relatively simple approach to estimating the underlying parameters associated with these mechanisms. Although our focus is on the differences between extreme and moderate experiences, our approach can be extended in many directions. Finally, we demonstrate the importance of distinctly understanding each of these possible mechanisms. In particular, we show that firms should treat intentional (utility-related) and unintentional (forgetting-related) nonreviewing as two separate entities that may require different treatment. Otherwise, firms may use financial incentives to change reviewer utility when a simple reminder may suffice. We present a simple calculation that shows that ignoring our forgetting-related mechanism may triple the effective review acquisition cost. We also show the cost savings that result from sending out unincentivized reminder emails sooner rather than later.
Literature Review
In this section, we first review evidence for the prevalence of extreme distributions in online reviews and discuss existing explanations for this phenomenon. We then introduce a novel, memory-based explanation for extreme distributions and review psychological and neuroscientific research that demonstrates a positive link between extreme experiences and prospective memory for experience-related tasks.
Evidence and Explanations for Extreme Distributions
Numerous studies have documented the disproportionate share of extreme online reviews. This phenomenon appears in virtually all product and service categories, including books (Chevalier and Mayzlin 2006; Godes and Silva 2012; Hu, Pavlou, and Zhang 2009), DVDs (Hu, Pavlou, and Zhang 2009), movies (Dellarocas and Narayan 2006; Liu 2006), home products (Moe and Schweidel 2012), home improvement products (Lafky 2014), physicians (Gao et al. 2015), restaurants (Yelp 2018), and accommodations (Fradkin, Grewal, and Holtz 2018).
Table 1 presents an overview of studies that find extreme distributions in online reviews, and it reveals three important insights. First, extreme reviews account for about two-thirds of posted reviews on platforms that do not allow for reciprocal rating between buyers and sellers. Second, the highest possible rating score accounts for about 50%–60% of reviews on these platforms. In contrast, platforms that allow for reciprocal ratings, such as Airbnb, exhibit an even greater share of extreme reviews, and this share is exclusively driven by extremely positive reviews (Fradkin, Grewal, and Holtz 2018). Third, the table reveals that most existing research has focused on Amazon reviews. Schoenmüller, Netzer, and Stahl (2020) recently conducted an extensive study of extreme reviews across a wide range of platforms and categories. They report that on all 12 studied platforms that use a five-point rating scale, such as Amazon, extreme distributions are quite prevalent. For Amazon itself, the authors find that 84% to 98% of products from 24 categories exhibit extreme distributions. In contrast, the prevalence of extreme distributions is considerably smaller for platforms that deviate from the 5-point scale, such as RateBeer, which uses a 20-point scale, or MovieLens, which uses a 10-point scale.
Previous Studies That Report Extreme Distributions.
aShares of rating scores for extreme distributions were manually calculated from tables and figures in the article.
bPresented evidence for extreme distributions but did not report distributions of rating scores across categories.
Notes: To ease the comparison across studies, we relabeled theoretical explanations as utility-based if authors cited Anderson (1998) as a key reference for drivers behind extreme distributions or if they argued that posting extreme experiences yields greater utility to customers. An example are Gao et al. (2015), who introduce “hyperbole effects” in rating valence to explain the prevalence of more extreme reviews. However, most of their discussion is in the spirit of Anderson (1998) and emphasizes the higher utility that individuals derive from sharing extreme experiences.
The prevalence of extreme distributions has led researchers to speculate about the mechanism behind the phenomenon. The last column in Table 1 shows that the utility-based explanation has been the most widely used explanation for extreme distributions. In fact, we did not find a study that provided an explanation for these distributions and did not at least mention the idea that customers derive greater utility from sharing extreme experiences. The second most frequent explanation was the base rate explanation, though there were relatively few mentions of this. A third class of explanations related to platform-specific mechanisms such as reciprocal-rating procedures between buyers and sellers on Airbnb (Fradkin, Grewal, and Holtz 2018). Finally, Schoenmüller, Netzer, and Stahl (2020) discuss evidence by Mayzlin, Chevalier, and Dover (2014) and Luca (2011) on review fraud, which suggests that extremity is driven by very positive or very negative promotional reviews.
Though empirical evidence for the relative importance of different explanations remains scarce, existing studies mostly refer to the utility-based explanation for extreme distributions. For example, Schoenmüller, Netzer, and Stahl (2020) present empirical evidence from surveys, experiments, and secondary data, which they argue is consistent with the utility-based explanation. For example, forcing experiment participants to review their last product experience leads to less-extreme review distributions than allowing them to choose any past experience to review. Although the authors also report support for other drivers, such as the base rate explanation or review fraud, these effects are found to be much smaller. Similarly, Fradkin, Grewal, and Holtz (2018) argue that the utility-based explanation is the greatest source of review bias on Airbnb. In a field experiment, they show that reminder emails with $25 coupons in return for a review reduced extreme distributions relative to reminder emails without such coupons. Because these coupons increase the utility of posting any travel experience, they argue that this finding is consistent with the idea that, in the absence of such incentives, posting extreme experiences yields greater utility for customers. Lafky (2014) finds in a laboratory experiment that customers are more likely to share extreme reviews when reviewing is costly than when it is free. He argues that this finding is consistent with consumers having higher intrinsic costs of reviewing moderate experiences, which is a variant of the utility-based explanation. In summary, the prevailing view in the literature is that extreme distributions arise from an intentional utility-maximization process in which the utility to post extreme reviews exceeds that for moderate reviews.
Reviewer Attrition as a Novel Explanation for Extreme Distributions
Existing explanations for extreme distributions (implicitly) assume that consumers decide whether to review and then implement this decision. Surveys, however, paint a more nuanced picture of customers’ review provision process and show that many customers fail to carry out their planned review decision (Tomorrow Focus AG 2014). When asked about the reasons for this failure, most consumers state that, eventually, they forgot about it (Tomorrow Focus AG 2014). 4 This implies that after a while, many customers “forget to remember”; that is, they no longer actively consider writing a review and effectively leave the pool of potential reviewers. We call this process “reviewer attrition.”
Why do so many customers end up forgetting about their previously formed review intention? Whenever people form an intention to do something in the future (e.g., pick up the dry cleaning on their route from work, make an appointment for a dental checkup, write an online review), they rely on their prospective memory (Einstein and McDaniel 1990). According to Dismukes (2012, p. 215), “the term is something of a misnomer, given that what it refers to involves the cognitive process of planning, attention, and task management as much as it involves memory. After forming an intention, individuals become often engaged with various ongoing tasks and, in most everyday situations, cannot hold the deferred intention in focal attention.” The consequence is that people often forget about their original intention, and the intended action never gets implemented.
Several studies demonstrate that prospective memory failure occurs regularly and often unintentionally—sometimes even with severe consequences (e.g., speeding in traffic [Bowden, Visser, and Loft 2017], failed task execution in intensive care units [Grundgeiger et al. 2013]). To explain why prospective memory fails so regularly, researchers point out that remembering to do something in the future is inherently more difficult than remembering events or information from the past (which are stored in retrospective memory). In particular, successful retrieval from prospective memory requires that a person remembers not only what to do but also to do it (Schacter et al. 2020). As a consequence, prospective memory requires considerable self-initiation since it “requires that persons remember to remember in the first place” (Einstein and McDaniel 1990, p. 717).
To remember their previously formed intentions, individuals frequently resort to both internal and external reminders or trigger cues to help them implement the intended actions. Research indicates that such trigger cues are particularly effective if they are salient and directly related to the intention (Grundgeiger et al. 2013; Loft, Smith, and Bhaskara 2011). Examples include internal cues such as grocery shopping lists and external cues such as car alarms that alert the driver when a passenger has not buckled their seat belt.
In the context of online review provision, we expect that thoughts about the associated consumption experience represent a salient, internal trigger that is related to review intention. We derive this expectation from network memory models and the theory of spreading activation (see Collins and Loftus 1975). According to this theory, concepts in memory are denoted by nodes, which are connected to other nodes, and activation of one particular node triggers related nodes. Thus, a customer who thinks about their consumption experience channels focal attention to the associated node and may find themself thinking that they need to write a review about the experience. As a consequence, forgetting to remember will be less likely when the consumer rehearses the associated consumption experience more frequently. However, rehearsal frequency depends on the underlying experience. Studies report that rehearsal frequency is significantly higher for extremely emotional experiences than for moderately emotional experiences (Walker et al. 2009).
Overall, and in the absence of external cues, we thus expect lower reviewer attrition rates for extreme versus moderate experiences. That is, relatively fewer consumers will forget to remember to write a review for an extreme experience than for a moderate experience, contributing to extremity of online reviews. However, reviewer attrition may also be reduced by external, salient, and directly related trigger cues for review provision. A prominent type of such external cues are review solicitation emails that are commonly sent to nonreviewing customers. Indeed, even simple reminders (without incentives) have been shown to increase review volume (Burtch, Bapna, and Griskevicius 2018), although no theoretical explanation for such findings has yet been provided. In our analytical model, we explicitly incorporate the influence of such emails on reviewer attrition and extreme distributions.
Experimental Design and Evidence
In this section, we describe a field experiment we conducted in cooperation with a large European online travel platform. The travel platform wishes to remain anonymous. The experiment plays two important roles in our analysis, allowing us to build on and contribute to the existing literature on extreme distributions. First, to our knowledge, we conduct the first randomized controlled study of the impact of unincentivized reminders on review extremity. Most other studies either compare distributions before and after a reminder is sent (Karaman 2021; Schoenmüller, Netzer, and Stahl 2020) and/or provide some form of incentive for participants to write a review (Fradkin, Grewal, and Holtz 2018). Second, and most importantly, the experiment provides a source of exogenous variation that allows us to tease apart the underlying causes of extreme distributions and assess, in particular, our proposed theory of differential attrition.
Company Background
The online travel platform we partnered with has been very successful for more than ten years, making it one of the two largest travel platforms in its core market segment. In the second half of 2018, for example, the platform attracted, on average, more than five million unique monthly users. The platform attributes much of its success to the availability of more than seven million customer reviews for more than 700,000 hotels on its site, and it places great strategic importance on having a robust set of current customer reviews. Owing to the dynamic quality of hotels, the platform constructs average hotel rating scores based only on reviews from within the last two years, even if older reviews are available.
The travel platform obtains review content from two different groups of travelers: its customers (i.e., those who have previously booked a vacation through the platform's travel agency) and travelers who booked their vacation with a different travel platform or agent. In this way, the platform combines the approaches of similar platforms such as Expedia (where only customers of the platform can write a review) and TripAdvisor (where all travelers can submit their reviews). Note that in our analysis, we only use reviews from matched bookings placed through the platform's own travel agency to minimize the incidence of fake reviews. This is consistent with the view, as discussed by Mayzlin, Chevalier, and Dover (2014), that platforms’ verification of reviewer identity should decrease the incidence of fake reviews. 5
Review Solicitation Email
The platform sends out a review solicitation email on the first day after the end of a vacation to all customers who have not yet provided a review for their hotel experience. This email welcomes customers home, asks them for a hotel review, and provides links to the hotel's product page and an online rating form. Importantly for our purpose, this email does not include any financial incentives or social norms (e.g., how many customers wrote a review in the recent past), both of which have been shown to affect review provision and content (e.g., Burtch, Bapna, and Griskevicius 2018; Klein et al. 2018; Woolley and Sharif 2021). Figure 1 displays a translated, stylized example of this email. As the example shows, the solicitation email always provides a picture and the name of the focal hotel. The email thus serves as a salient external cue for previous review intention (both directly and indirectly, through activation of the travel experience memory).

First solicitation email: Content and form.
If a customer clicks on the email link to the online rating form, they will go through the same review procedure as would a customer who posts a review before receiving a solicitation email. Specifically, the form asks each customer to answer several questions, such as whether they would recommend the hotel (yes/no), how they would rate the hotel overall on a scale from 1 (“very bad”) to 6 (“very good”), how they would rate different quality aspects of the hotel (e.g., location, service), and how they would rate the value for money at this hotel. The consumer then needs to provide a text description that is at least 100 characters long and is asked about some personal and travel characteristics (e.g., age, country of residence, timing and length of stay, reason for travel). 6 If a customer does not respond to this email, the travel platform makes up to two additional attempts to solicit a review from the customer. The second and third emails (if the second email did not result in a review) are sent on the fifth and ninth days after the end of the customer’s vacation, respectively. 7 If no review has been provided after nine days, the company ends its review solicitation attempt and waits another 14 days before sending a final email in which customers can win a €100 voucher for their next booking, at which point entering the lottery does not require reviewing the hotel.
It was against this background that the company agreed to implement our field experiment in which it randomly allocated customers to one of four experimental conditions that differed only in the timing of the solicitation emails. All other aspects of the emails and review solicitation procedure remained identical across conditions. In particular, the emails always identified the travel platform, and never the hotel, as the sender.
Experimental Manipulation
Our experiment began on June 1, 2017, and concluded on September 26, 2017. The experimental design involved four conditions. Condition 1 represented the previously discussed status quo at the travel platform. In the other three conditions, we increased the amount of time between the end of travel and the day that the first review solicitation email was sent: In Condition 2, the first email was sent on the second day after the end of travel, and in Conditions 5 and 9, it was sent on the fifth and ninth day after the end of travel, respectively.
Returning customers were randomly allocated to the four different conditions as follows. On the first day after the end of a vacation, an algorithm confirmed each customer's review status, that is, whether a review had already been provided. All customers who had not yet provided a review for the vacation under study were randomly allocated to one of our four conditions, all with the same allocation probability. After the end of the experiment, we obtained detailed information on bookings and hotel characteristics, which allowed us to match this information to reminder emails and hotel reviews.
Taking the study design into account, we identified six possible tests to evaluate the impact of a review solicitation email on review extremity. Table 2 provides an overview of these tests based on different “review latency values,” or the number of days between the end of travel and the time of review provision. Test 1 used only reviews that were provided on the first day after the end of travel and compared the share of extreme reviews observed in Condition 1 (in which the platform had already sent the reminder email) with that in all other conditions (in which the platform had not yet sent the email). As the experimental treatment in our design was customers receiving the review solicitation email, Condition 1 served as the treatment condition in Test 1, and the others served as control conditions. In Test 2, we used more observations to increase the statistical power of our test. Specifically, we included all reviews that were posted within the first four days after the end of travel and compared the share of extreme reviews observed in Condition 1 with that of Conditions 5 and 9. Note that because the platform sent the email on the second day in Condition 2, we could no longer use this condition in our control group.
Experimental Design: Treatment and Control Conditions.
Table 2 shows that there exist four additional tests that cleanly assign posted reviews from a given day after end of travel to treatment and control conditions: two using Condition 2 as treatment, and two using Condition 5 as treatment. However, as we move from left to right in the table, there remain fewer nontreated observations left to serve as the control group. Note that by holding constant, within each test, across conditions, the number of elapsed days since the end of travel and review provision, we are able to rule out common patterns across time, such as improved customer understanding of past extreme experiences (Moore 2012), as an alternative explanation for a change in the share of extreme reviews across conditions. Throughout this section, we are interested in the effect of the travel platform's first unincentivized reminder email on review provision and extremity and not in the effects of the hotel management's communication with customers. 8 We consider the full set of three reminders in our empirical estimation.
The Data
Our data set is constructed by matching review-solicitation emails to bookings and reviews, each of which resides in distinct data tables. We describe the exact data construction procedure in Web Appendix A. We note here that, to ensure a balanced number of observations across each of the four experimental conditions, we exclude all bookings with end dates between September 18, 2017, and September 25, 2017. This nine-day window ensures that each of the subjects in each of the treatment conditions received their first email. 9 As a result of this procedure, our final data set includes observations for 189,842 hotel bookings with 35,238 matched reviews. Accordingly, 18.6% of the bookings in our sample resulted in a review.
To evaluate the effectiveness of our randomization procedure, we tested for differences in key booking characteristics across all four experimental conditions. Table 3 displays summary statistics and results from Kruskal–Wallis tests for differences across conditions. We observe that trips lasted on average about eight days, with an average price of around €1,670. With regard to customer characteristics, we see that the average trip involved 2.34 travelers (the median value was 2), that customers returned, on average, from one trip within our sample period (although some customers had multiple bookings), 10 and that the average customer age was 41 years. 11 We see very little variation in the data across conditions, which suggests that our randomization procedure is effective. The results from Kruskal–Wallis tests largely support this impression and detect a significant difference across conditions only for the number of travelers per booking. Note that the number of travelers per booking differs only in the second decimal point across conditions. Nevertheless, in the “Experimental Results” section, we report results that include the number of travelers per booking as an additional control.
Balance Checks Across Treatment Conditions.
**p = .05.
Notes: Number of observations are for travel duration, price, travelers per booking, and bookings per customer. For customer age, the associated values are
Assessing Review Extremity
Figure 2 displays the rating score distribution in our sample and yields two important insights. First, and comparable to previous research, we observe a left-skewed distribution, in which 44% of reviews involve the highest possible rating of 6. Second, reviews with the lowest rating of 1 are extremely rare and account for less than 2% of all posted reviews in our sample. To make the share of extreme ratings in our sample comparable to shares of around 50%–65% in previous studies (as reviewed in Table 1), we classify a review as “extreme” if it involves a rating score of 1, 2, 3, or 6. As a result of this approach, extreme reviews account for 54% in our sample. 12

Distribution of rating scores at travel platform.
Experimental Results
Table 4 presents the estimated average treatment effect of a review solicitation email on review extremity for our six tests. In Tests 1 and 2, we see that the share of extreme reviews is significantly lower in Condition 1 than in the other conditions. Specifically, in Test 1, we focus on the first day after the end of travel and find that the share of extreme reviews is 55% in Condition 1 but 61% when pooling across Conditions 2, 5, and 9. Similarly, Test 2 shows that across days 1 to 4, the share of extreme reviews is 55% in Condition 1, but 61% when pooling across Conditions 5 and 9. Tests 3 and 4 show that the share of extreme reviews is also significantly lower in Condition 2 on days 2–4 after the end of travel than in Conditions 5 and 9. The results for Tests 5 and 6 replicate this pattern for days 5–8 after the end of travel when comparing Condition 5 with Condition 9. Overall, the displayed results clearly establish that review extremity decreases following a solicitation email. While this is consistent with the findings in the existing literature, our results are the first to establish the causal effect on review extremity of an unincentivized reminder in a randomized controlled design.
Share of Extreme Reviews Across Conditions.
*p < .10.
**p < .05.
***p < .01.
This value is based only on Conditions 5 and 9.
Notes: Displayed are proportions of extreme reviews across sets of days after end of travel and conditions. z-Stat. denotes z-statistic for tests of proportion equality across conditions.
Table 5 displays estimation results from logit models on the likelihood of an extreme review when controlling for the number of travelers per booking. This model specification addresses the potential concern that the results in Table 4 might be a reflection of the previously detected differences in the number of travelers per booking across conditions (as displayed in Table 3). However, the results in Table 5 clearly reject this idea and demonstrate that the likelihood of extreme reviews is consistently around 6% lower when travelers have just received a reminder email compared with the case in which they have not. Overall, these results confirm our previous insights.
Logit Estimations for Likelihood of Extreme Reviews Across Conditions.
*p < .10.
**p < .05.
***p < .01.
This effect is measured relative to Conditions 5 and 9.
Notes: Displayed are marginal effects for logit specifications. Robust standard errors are displayed in parentheses.
In summary, our experimental results establish that solicitation emails reduce review extremity. We turn next to the question of “why?” Recall that the extant research has interpreted analogous results as evidence for the presence of differential utility. However, as we demonstrate in the next section, differential attrition is an equally plausible explanation for these results.
Theory
We proceed in three steps in this section. First, we provide a simple analytical model of review provision. Second, we show how the three focal theories—differential base rates, differential utility, and differential attrition—can be captured in this model and explore under what circumstances the review-generating process results in the extremity bias. Finally, to explain the results of the field experiment and to motivate the empirical analysis that follows, we investigate the extent to which the three theories can explain the observed decrease in extremity following an unincentivized solicitation email.
A Simple Model of Review Provision
Consider the following process. In period
Formally, the probability that a customer of type i posts a review on the first day after returning from a trip is
In Theorem 1, we explore the extent to which
The following reflect the expected proportion of extreme reviews:
If Otherwise, when If If
Theorem 1 demonstrates several important points. First, as Case 1 shows, a higher incidence of extreme experiences in the population,
In contrast, consider Case 2b, where the differential in attrition is small relative to the differential in the posting rate:
Having established the theoretical causal link between any one of these three mechanisms, on the one hand, and extreme distributions, on the other hand, we turn next to an analysis of the exogenous shock induced by the review solicitation emails. This analysis demonstrates that although base rate differences alone cannot explain the experimental results, both differential utility and differential attrition can. These results will motivate the final stage of our study in which we disentangle empirically the latter two mechanisms.
Review Solicitation Analysis
Here, we examine the effect on the review distribution of a review solicitation email sent by the platform, asking each customer who has not yet written a review to post a review about their most recent hotel experience. To predict the impact on review distributions of such an intervention, we first need to specify our assumptions regarding its role in the review process laid out previously. In fact, we assume that reminders may have two potential effects. The first such effect is to bring back into the pool of potential reviewers those customers who previously left due to attrition. We refer to this as “attrition reversal.” As discussed in the “Literature Review” section, the email represents an external, salient trigger cue and serves as a reminder to those consumers who forgot to remember their original posting intention. 21
The second potential effect of a reminder is to temporarily increase the utility associated with posting a review. We refer to this as the “utility-boost effect.” This may be a cognitive effect (e.g., a signal that reviews are important to others and thus should be prioritized) or a subconscious process, such as mere exposure. We capture this in the model by assuming that, following the solicitation email, the probability that a reviewer
22
of type i posts a review is
Having specified the potential impacts of the reminder, we next incorporate them into the model. Consider a customer at the beginning of
Suppose that there is no solicitation email at the beginning of period
A solicitation email sent at the beginning of period
As Theorem 2 shows, the total effect of the review solicitation email depends on the relative strength of (1) the utility-boost effect across the two types of consumers (the left-hand side of Inequality 6) and (2) the attrition-reversal effect (the right-hand side). Now, to appreciate the intuition behind the impact of the solicitation email, and in the absence of any theory to the contrary, let's assume that the utility boost for moderate reviews is equal to that for extreme reviews:
Next, consider the case in which the attrition differential is large relative to the utility differential:
In contrast, consider the case of uniform attrition:
Note that it is possible that there are significant differences between the utility boost experienced by customers of different types. Indeed, in the next subsection, we report that
Finally, note that the base rates have no effect on either the utility boost or attrition reversal, and thus differences in base rates cannot explain the results observed in our field experiment. 24 Indeed, these results, combined with the theory developed in this section, suggest that a decrease in the extremity of reviews could be due to either differential attrition or to a larger utility boost for moderate reviewers, even in the absence of differential attrition. It may also be a combination of these factors. In the next section, we develop an estimation approach that allows us to empirically disentangle these different mechanisms and explore their relative power in explaining not only the experimental results but also extremity bias in online reviews, more generally.
Empirical Estimation of the Review-Generating Process
In this section, we focus on estimating the underlying parameters of the review-provision process in Equation 1. This will allow us to assess simultaneously the validity and the relative impact of two of the theoretical mechanisms: differential utility and differential attrition. While our approach will preclude the direct estimation of the base rates parameters, our focus here is on disentangling the two mechanisms that are able to explain the observed decrease in extremity following a reminder. In addition, and importantly, given the structure of our data set—in particular, the experimentally induced review-solicitation emails—our model will account for the impact of the intervention.
We begin by specifying the likelihood for the data generated by our experiment, incorporating the specific details associated with the reminders. We then compare a range of models, allowing us to conclude that (1) all models we consider demonstrate the consistent presence of differential attrition, a construct not previously discussed in the literature; (2) this differential attrition exists alongside the simultaneous presence of differential utility, the core assumption behind the utility mechanism; and (3) the model that fits best is consistent with the assumption that reminders both restore to the pool of potential reviewers those who exited previously via attrition and temporarily boost the utility associated with posting reviews of any type. That is, it supports both utility boost and attrition reversal as outcomes associated with email reminders.
Likelihood Construction
Note that the parameters underlying both the utility and the attrition mechanisms are conditional on the “type” of the review (i.e., whether it is moderate or extreme). Because we cannot observe the type unless a review is provided, we focus on maximizing the conditional likelihood (Wooldridge 2002). Consider first a context without reminders. Consistent with our theoretical model in Equation 1, the likelihood of a single review provided at latency
Recall that each of our four experimental conditions has three reminders separated by four days each. By assumption, as outlined previously, one impact of a reminder is to return to the pool of possible reviewers all those customers who dropped out of the pool via attrition before providing a review.
27
To appreciate why reminders would enable us to estimate our focal parameters when we could not do so without them, note that the review provision probability immediately following a reminder is not a function of
We also incorporate utility boost, as discussed previously, via the parameters
We thus maximize the following conditional likelihood with respect to the model parameters.
29
See Web Appendix E for a complete description of the components, which we summarize briefly here for the sake of parsimony.
The three terms in the numerator capture the prereminder, reminder, and postreminder periods within a given experimental condition c, respectively. The prereminder expression follows directly from the simplified likelihood in Equation 8, while the same is true of the postreminder period except for the
Estimation Results
In Table 6, we present the estimates derived from the maximization of Equation 9 and compare them with those from a series of alternative models. Our main results in Model 1 provide evidence for both differential utility (
Maximum Conditional Likelihood Estimation Results (T = 28).
*p < .10.
**p < .05.
***p < .01.
Note: AIC = Akaike information criterion.
In Model 4, we present results from a model in which reminders only boost the utility parameters (from
Although the empirical analysis provides evidence for our proposed theory based on differential attrition as well as for the extant theory of differential utility, there remains the question of the base rates theory (i.e., that
Note that our theory implies that the expected number of type i reviews at
An alternative assumption on the relative magnitude across conditions of period 0 attrition would be that it is proportional to the attrition experienced in later periods. This would imply that period 0 attrition for extreme experiences occurs with roughly 44.1% higher probability in moderate than in extreme experiences,
32
or that
Finally, we consider in Web Appendix F a range of alternative specifications to investigate the robustness of our findings, including the decay of the utility impact of the reminders; different data windows; different categorization approaches for “extreme” and “moderate” reviews; and stratifications by several observable reviewer and booking characteristics such as reviewer gender, travel price, and hotel destination. We also show that the results are robust across stratified samples that control for other sources of information for reviewers (i.e., the prevailing rating average on the day of review provision). Finally, we present estimation results from a model that explicitly considers reminder-specific utility-boost effects across the three different reminder emails.
Other Potential Effects of the Review Solicitation Email
Throughout this article, we have assumed that the reminder email affects review provision through two effects: (1) attrition reversal and (2) a temporary boost in the utility of posting a review. It is reasonable to wonder whether there exist other, possibly unmeasured, effects of these interventions that might confound our main results, as well as any inferences drawn from them. Although the primary role of the utility boost (
Another potential confound may come from customer's perception of the hotel's (ulterior) motives to ask them for a review. However, this confound is unlikely to exist in our setting, because (1) solicitation emails are a natural (i.e., frequently encountered) stimulus in real customer–firm interactions; (2) the email is always sent from the platform, not the hotel; and (3) the first two emails did not promise any financial rewards in exchange for reviewing. Finally, one may wonder if (1) solicited reviewers may have differed from unsolicited reviewers and (2) other content dimensions of solicited reviews differed from those of organic, or unsolicited reviews. However, we document in Web Appendix G that observable characteristics look very similar across both groups of reviewers, thereby reducing this concern. In addition, we report in Web Appendix H that the review text—a prominent other content dimension—shows few systematic differences for the same rating score across both groups. Moreover, the few differences that we find seem to go against our main results. These results support the view that reminder emails are relatively “soft touch” instruments.
Discussion
In this article, we introduce a novel, attrition-based mechanism that explains the prevalence of extreme distributions, “one of the most robust findings in product reviews” (Moe, Netzer, and Schweidel 2017, p. 484). The key aspect of this mechanism is that consumers may forget to write a review for their consumption experiences, and that forgetting rates are higher for moderate than extreme experiences. Starting from experimental evidence that unincentivized solicitation emails cause a 10% reduction in review extremity in the field, we develop a simple model of review provision, in which we combine our novel attrition-based mechanism with the previously proposed utility- and base-rate mechanisms. We then derive parameter constellations under which the different mechanisms predict a decrease in review extremity in response to a review solicitation email. Finally, we develop a conditional likelihood approach to disentangle these different mechanisms. Our results demonstrate the simultaneous existence of the attrition-based and utility mechanisms. We find that a model without attrition fits the data substantially worse, implying that both mechanisms capture unique aspects of the review provision process. Our study thus provides first evidence for the importance of attrition-based effects for review provision.
Theoretical Contribution
Our work contributes to the literature on word of mouth in four important ways. First, we introduce a novel attrition-based explanation for extreme distributions and show that inclusion of this mechanism explains the observed empirical patterns better than existing explanations that focus exclusively on differences in reviewer utility from posting or differences in customer base rates across types of experiences. Our model requires the following two assumptions: (1) consumers with more extreme experiences have a lower attrition rate, and (2) those who left the sample can be brought back with a reminder. That is, our model can be considered a modified version of an attrition model in the CRM literature (e.g., Fader, Hardie, and Shang 2010) in which “dead” consumers can be brought back with a review solicitation email.
Second, our theory identifies the interaction of customer attrition and review solicitation emails as a novel driver for the instability of online review distributions over time. From a theoretical point of view, this contributes to our understanding of dynamics in online reviews. Existing work has focused on social and temporal dynamics in review distributions (Godes and Silva 2012; Li and Hitt 2008; Moe and Schweidel 2012; Moe and Trusov 2011; Wu and Huberman 2008), studied dynamics that result from (hotel) managers’ responses to previous reviews (Chevalier, Dover, and Mayzlin 2018; Proserpio and Zervas 2017), or exclusively attributed any effects of solicitation emails to changes in customer motivation (albeit without testing this mechanism (Askalidis, Kim, and Malthouse 2017).
Third, we contribute to the existing knowledge of biases in online reviews. Whereas some theories, such as the base rate explanation, imply that posted reviews are an unbiased representation of customers’ underlying product experiences, others, such as our attrition-based explanation, suggest otherwise. Drawing on the results of our field experiment, we conclude that online reviews do not represent an unbiased view of customers' actual experiences. Our results thus confirm the importance of reviewer self-selection as a source of review bias as previously discussed in the literature (e.g., Godes and Silva 2012; Moe and Schweidel 2012) but specify a novel selection mechanism. Interested researchers may thus consider the integration of attrition into their empirical and analytical work.
Finally, we provide a general model of review provision and demonstrate how to integrate different theoretical lenses into this model. Our results demonstrate the importance of formalizing theories for review provision to enable clean empirical tests of their predictions. Future research could build on our work to study additional factors, such as social influence, or differences across groups of customers on review provision.
Managerial Implications
Companies are constantly looking for ways to attract more word of mouth activity from their customers. This is evident from the considerable amounts of money that they spend on review solicitation (Babic Rosario et al. 2016) and the increasing supply of websites that offer guidance on the best ways to get more reviews and word of mouth. Our research has three important implications for companies trying to attract more reviews from their customers. First, our identification of an attrition-based mechanism behind extreme distributions paints a more nuanced picture of review provision. In particular, we show that managers do not need to include expensive financial incentives in their solicitation emails, as indicated by previous research (e.g., Burtch, Bapna, and Griskevicius 2018; Fradkin, Grewal, and Holtz 2018; Schoenmüller, Netzer, and Stahl 2020)—a simple reminder email can already help solicit more reviews.
In our study, sending an unincentivized reminder email increased review provision within the first four days after the end of travel from 2.8% without a reminder (Condition 5) to 11.00% (Condition 1). This difference of 8.2 percentage points is substantial when compared with previous results on the effectiveness of financial incentives in solicitation emails. For example, Fradkin, Grewal, and Holtz (2018) show that a $25 coupon increased review provision by 6.4 percentage points relative to an unincentivized email. Combining this increase with our results, we can infer that the $25 coupon is 14.6 percentage points more effective than no reminder at all. The magnitude of the inferred increase is consistent with Burtch, Bapna, and Griskevicius (2018), who found that a reminder SMS with a 10 Yuan coupon increased review provision from 4% without a reminder to 18.25%—also an increase of more than 14 percentage points. Our results suggest that more than half of this increase could have been achieved with an unincentivized reminder email. In other words, by including financial rewards in each solicitation email, firms end up with effective acquisition costs per review that are much higher than the included coupon.
A simple example will help illustrate this. Assume that 1,000 consumers had a recent consumption experience. According to our results, about 3% of them would write a review without any incentives or reminders, 11% would write a review after a reminder, and 17.4% would write a review when the firm provides financial incentives. It is then straightforward to show that offering financial incentives to all customers (say, in the form of a $25 coupon), would result in total costs of $4,350 for the firm ($25 for each of the 174 reviews). However, the number of additional reviews above and beyond sending a simple reminder would be 64 reviews. So the effective acquisition costs for these incremental reviews would not be $25, but $68—almost three times as high!
Second, we show that sending unincentivized solicitation emails may be beneficial for some, but not all, stakeholders. Consider first product sellers. For this group, solicitation emails may be a double edged-sword: although they increase review volume at lower cost than using financial rewards alone, they also result in more moderate product reviews. This latter aspect is a major difference relative to the use of financial incentives, which have recently been shown to increase reviewer positivity (Woolley and Sharif 2021). To the extent that a product currently exhibits a J-shaped distribution, our results show that unincentivized solicitations may result in a downward shift in the review distribution, potentially hurting the sellers’ sales. However, for platforms and customers, sending out unincentivized reminders does not yield the same trade-off as both parties are interested in a larger number of less-biased online reviews. Reviews that are more representative of the underlying distribution of experiences enable better matching of customers to products, which increases the attractiveness of the platform. However, for customers this benefit comes at potentially higher annoyance costs should more firms start sending out solicitation emails in the future.
Third, we demonstrate that firms need to carefully design an effective email management system. For example, additional analyses of our data reveal that the timing of the first review solicitation email is crucial. In particular, we find that waiting to email customers may not be a good strategy for platforms that aim to maximize the review-provision likelihood from customers: although this likelihood is 20% when the first email is sent on the first day after the end of travel, it monotonically decreases with additional delay in the first email to only 17% when the first email is sent on the ninth day. This shows that it is important for travel platforms to engage customers early on in review provision for travel experiences.
Additional calculations show that a three percentage points increase in the review probability creates considerable value for the platform. Specifically, we calculated the average number of ending trips for each day in our sample period—these trips represent the cohort of potential reviewers from that day. On average, 1,759 trips ended on every single day of our sample period, with some days showing as many as 2,577 ending trips. Accordingly, a three percentage point increase in review provision resulted in 53 to 77 additional reviews for each cohort. Over a 30-day period of reviewer cohorts, the change in review provision amounts to almost 1,600 to 2,300 additional reviews. If we take into account that many firms frequently pay between $5 and $25 as mentioned previously, the monthly savings for a firm can be as high as $11,500 to $58,000—only by changing the timing of the email at no extra costs. Importantly, these additional reviews also benefit customers who are looking to find the best suitable hotel for themselves. Overall, the value of a higher review probability thus exceeds the pure monetary savings for platforms.
While not directly related to firms’ attempts to generate more word of mouth for their products, our results finally demonstrate that solicitation emails are not an innocuous intervention from firms. Instead, these emails affect the share of extreme reviews and the length of reviews. This raises the practically relevant question as to whether firms need to disclose their review solicitation practices and which reviews they obtained through these practices. Other examples suggest that policy makers and consumer protection agencies might require such disclosure in the future. For instance, reviewers who received the reviewed product for free from the producer are nowadays required by law to disclose this information. One potential avenue might be to require firms and platforms to separately group reviews that were submitted before and after review solicitation emails (Askalidis, Kim, and Malthouse 2017). Our research suggests that more research on the implications of such email interventions for customer decision making and welfare is required.
Limitations of This Study
Just like any other research, our study is not without limitations. In particular, we acknowledge that it would have been very interesting to separate the effect of merely receiving a review solicitation email (but not opening it) from the effect of receiving and opening this email. Unfortunately, our data do not allow us to study this effect. The reason is that, for customers who did not open the email, we have no way to establish the exact point in time when they first noted the reception of this email. This creates a problem for our identification approaches that focus on review provision shortly after the email was actually sent: for those customers who wrote a review on the day of the review solicitation email without having opened this email, we do not know whether they had observed the email or not. However, a closer look at our data suggests that, even with this information, the effect would have been difficult to identify as this group of customers is small, accounting for only .8% of all reviews in our sample.
Another limitation concerns our assumption that a customer's type remains fixed following the experience. Thus, by assumption, we do not allow the reminder to cause a switch in the reported experience. This implies that we are not considering reminders that involve any sort of a financial or a social incentive to post a more positive rating, for example. Similarly, our context, in which the reminder email always contained the same wording and was always sent by the travel platform, does not allow us to explore how the nature of the reminder (whether the email comes from the seller or the platform or the specifics of the text of the email) affects the resulting distribution of reviews.
Finally, and as we discussed previously, our model does not identify period-0 attrition: those customers who do not plan on reviewing the hotel. It would be managerially interesting to explore whether we can identify which customers are open to reviewing (vs. not). Relatedly, as we have mentioned, we find that there are some differences in the main results based on the gender of the customer, the destination, the hotel type, and other characteristics. We think that there is an opportunity to build richer models that may explain some of this heterogeneity on which this article is not focused.
Supplemental Material
sj-pdf-1-mrj-10.1177_00222437211073579 - Supplemental material for Extremity Bias in Online Reviews: The Role of Attrition
Supplemental material, sj-pdf-1-mrj-10.1177_00222437211073579 for Extremity Bias in Online Reviews: The Role of Attrition by Leif Brandes, David Godes and Dina Mayzlin in Journal of Marketing Research
Footnotes
Appendix: Mathematical Proofs
Acknowledgments
The authors are thankful to the anonymous participating platform and for helpful comments from Anand Bodapati, Eric Bradlow, Gareth James, Maryam Saeedi, Christophe van den Bulte, and Chutian Wang, as well as from seminar participants at University of Lucerne, Emory, Duke, Wharton, Yale, Johns Hopkins, UCLA, University of California, Davis, the NBER Economics of Digitization Conference, the Marketing Science Conference (Rome, Italy), Interactive Marketing Research Summit (Amsterdam, the Netherlands), ZEW 9th Conference on the Economics of Information and Communication Technologies (Mannheim, Germany), and the Social@IDC Conference (Herziliya, Israel). Flavia Tinner provided excellent research assistance. Part of this research was completed while the first author was a faculty member at Warwick Business School, University of Warwick, UK. The authors are listed in alphabetical order and contributed equally to the manuscript.
Associate Editor
Hari Sridhar
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship and/or publication of this article.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
