Abstract
The use of logit and probit models when examining binary dependent variables including those in the form 0/1 (i.e., dummy variables), yes/no, and true/false (hereafter binary DVs) is commonplace. Yet, the appropriateness and effectiveness of such models are challenged when the event rate of a binary DV is rare or common. To better understand the impact on the field of strategy, we undertook a literature review and assessed recently published research in the Strategic Management Journal. We then utilized Monte Carlo simulations with results showing that as event rates become rarer or more common, issues including biased coefficients, standard error inflation, low statistical power to detect significant effects, and model convergence failure increasingly arise. In addition, small sample sizes amplified these empirical issues. Using a strategy example study, we also show how various analytic tools can lead to different findings when empirical models face an extreme event rate with small sample sizes. Based on our findings, we provide step-by-step guidance for strategy researchers going forward.
Keywords
Introduction
The use of logit and probit models has become commonplace in strategy research (Dean et al., 2007; Hoetker, 2007; Jeong et al., 2020; Shook et al., 2003; Zelner, 2009) as research modeling binary dependent variables including those in the form of 0/1 (i.e., dummy variables), yes/no, and true/false (hereafter binary DVs) has increased. However, the effectiveness of such models is reduced when binary DV event rates become increasingly rare (i.e., few event occurrences among the sample) or common (i.e., many event occurrences among the sample) (Firth, 1993; King & Zeng, 2001a, 2001b). One of the well-documented empirical threats from examining rare and common event rates is inflation bias wherein regression coefficients of logit and probit models tend to be upwardly biased (Firth, 1993; King & Zeng, 2001a, 2001b) 1 , which threaten robust and accurate analyses. Although rare and common event rates are often encountered in strategy research, our review shows that few strategy studies acknowledge or apply alternative analytic tools to address this issue.
Accordingly, our study has three main purposes. First, we seek to offer clarity in understanding the impacts as well as the mechanisms through which binary DVs with rare or common event rates create threats to logit and probit models. To achieve this purpose, we first review the prior literature on binary DVs with rare or common event rates and suggested alternative analytic tools. We also review the strategy literature to identify the extent to which strategy studies are exposed to threats arising from examining binary DVs with rare and common event rates. In examining 143 studies published in Strategic Management Journal over the last ten years (between 2010–2019) which examined binary DVs, we find that almost 40% (56 out of 143) face potential problems associated with rare or common event rates. However, only a handful (2 out of those 56 – less than 4%) attempt to account for these constraints by employing alternative approaches. This review of the literature suggests that the rare and common event rates (particularly rare event rates) are prevalent in the extant strategy literature and that few studies attend to the associated issues. Thus, strategy scholars should be concerned with and attentive to this important and relevant issue.
Second, we aim to investigate when and to what extent rare and common event rates detrimentally impact the empirical modeling of binary DVs in strategy research. To accomplish this, we employ Monte Carlo simulations with various binary DV event rates and various sample sizes––conditions that closely resemble the empirical settings often examined in strategy studies. Our simulation results reveal that four undesirable properties (i.e., coefficient bias, standard error inflation, inability to detect true causal effects, and model convergence failure) arise as binary DV event rates become rare or common. We also find that these empirical threats are dependent on the sample size. Specifically, these issues become significantly worse as the sample size decreases. Furthermore, we conduct sensitivity tests at various effect sizes.
Finally, we aim to provide practical guidance for future studies encountering binary DVs with rare or common event rates. We find that the commonly suggested alternative analytic tools––Firth and Rare Event (RE) logit models––do not perform equally well, each having contrastable advantages and disadvantages. Specifically, relative to Firth logit, RE logit model produces smaller coefficient standard errors and thus increases the accuracy of ascertaining true causal relationships. However, while Firth logit model significantly increases coefficient standard errors, it is better equipped to converge consistently––a noteworthy advantage over RE logit. Based on these findings, we are able to offer insights to better inform strategy scholars.
Our study contributes to the research methods and the strategy bodies of literature by revealing current limitations within the field and offering actionable guidance to strengthen the robustness of strategy research going forward. First, we offer clarification on the degree to which rare or common event rates impact robust analysis in strategy studies. Indeed, our simulations based on parameters common in strategy studies suggest increasing empirical issues arising at extreme event rate levels. Second, we further highlight that threats posed by rare and common event rates are dependent on sample size, namely issues increasingly arise in small sample sizes. Thus, strategy scholars need to consider both the event rate and sample size when choosing an appropriate empirical approach. Third, based on our analyses, we offer insights and practical guidance for strategy researchers which we hope will aid in understanding and applying robust analytical approaches.
This study is organized as follows. In the next section, we review the literature on binary DVs with rare or common event rates, discuss analytical techniques used to examine binary DVs, and explore how binary DVs are examined in the strategy literature. Next, we explain how we designed our Monte Carlo simulations, offer an interpretation of our simulation findings, and discuss how our findings inform scholars. We also provide a brief strategy example, empirically highlighting how different analytical tools examining the same phenomenon and using the same sample can find different results––a major concern for scholars. Finally, based on our findings, we provide practical guidance for strategy scholars.
Analyzing Rare and Common Binary DVs
Rare and Common Binary DVs
A binary DV has only two outcome values––either a one (1) or a zero (0). Typically, the presence of a characteristic or the occurrence of an event of interest in research models is coded as a one and the absence of a characteristic or event non-occurrence is coded as a zero. Because binary DVs are coded as either one or zero, the mean value indicates the event rate at which the characteristic or event occurs in the total sample. Rare event rates indicate that the event rate is quite low such that the DV has many zero, but a few one values. In this case, the mean value of the DV would be near zero. Common event rates indicate the opposite, with many ones, but a few zero values. As a result, the mean value of the DV would be closer to one.
Rare and common event rates primarily occur for two reasons. First, the phenomenon of interest itself is rare or common. Indeed, research topics across various fields (e.g., financial misconduct in business (Shi et al., 2017), insurance fraud in economics (Jin et al., 2005), war between countries in political science (King & Zeng, 2001a)) are oftentimes inherently rare. Second, although the phenomenon of interest may not be rare or common (e.g., firm entries and exits; Claussen et al., 2018; Furr and Kapoor, 2018), scholars are likely to encounter overly rare or common event rates in their samples because of their sampling methods. Scholars often use panel data and various levels of analyses such as firm-year, dyadic-in-year, and dyadic-in-market. These sampling methods help scholars match event occurrence samples with control samples (non-event occurrence). However, if control samples contain many (or few) non-event observations, actual even rates in the final sample may become rare (or common).
Analytical Tools for Rare and Common Event Rates of Binary DVs
To examine binary DVs, logit and probit models have been popular analytic tools in strategy research (Hoetker, 2007; Jeong et al., 2020; Zelner, 2009). Although the DV is binary, logit and probit models are interested in the response probability or propensity (i.e., P in Equation (1)) that is estimated by independent variables (hereafter IVs; Wooldridge, 2012). Equation (1) simplifies logit and probit models with one IV (i.e., x).

Effects of rare and common binary dependent variables on response probabilities.
In response to the challenges arising from analyzing binary DVs with rare or common event rates using traditional approaches (i.e., logit and probit models), two primary alternative analytic tools have been proposed––(1) Firth logit (Firth, 1993) and (2) Rare Event (RE) logit regression models (King & Zeng, 2001a, 2001b). Both tools were originally developed to correct the inflation biases to coefficients faced by logit/probit models when examining binary DVs with rare or common event rates. Each alternative takes a different approach to correct this bias––Firth logit model employs a penalized maximum likelihood estimation (MLE) approach whereas RE logit model employs a weighted MLE approach. 3 As explained in the prior section, the log-likelihood resulting from rare or common event rates tends to be relatively flat, leading to inflation biases when facing a steep increase at either low or high x spectrum. Both penalized and weighted MLE approaches are techniques to account for this problem by using a subtraction (i.e., penalty) in the case of Firth logit or a weight to the log-likelihood in the case of RE logit models (see Firth (1993) and King and Zeng (2001a, 2001b) for details).
While each analytical tool offers advantages, each option is not without disadvantages. These disadvantages may be one reason why these tools are rarely used. In particular, our analyses indicate that the alternative analytical tools (i.e., Firth and RE logit models) often tend to overcorrect biases from rare and common event rates, decrease statistical ability to detect true causal effects, and decrease model convergence rates. We will discuss both advantages and disadvantages in more detail later, but mention them here as we seek to better understand the benefits and drawbacks of each approach so that we can offer guidance to scholars working with rare or common event rates. However, before we delve into the empirical effects of rare and common event rates, we first examine the prevalence of such issues in the strategy literature.
Rare and Common Binary DVs in Strategy Research
We first reviewed strategy studies to understand whether binary DVs with rare or common event rates are an issue commonly encountered in strategy research and, if so, how strategy researchers have addressed this in their empirical approaches. To accomplish this, we searched recent empirical papers with binary DVs published between 2010 and 2019 in Strategic Management Journal. 4 We chose this journal for our review because it is the preeminent strategy journal. We believe that if the issues of rare and common event rates are present in the preeminent strategy journal, it is likely that other strategy journals have similar issues. We found a total of 162 strategy articles by using keywords like logit, probit, binary DV(s), etc. As some strategy articles have multiple samples within them, we considered each sample to be a unique study. Three researchers coded the articles separately and any disagreements were resolved through discussion.
This resulted in 193 unique studies relevant to our study. For analysis, we excluded articles that did not report sufficient descriptive statistics to ascertain both sample size and event rate (i.e., the mean value of the DV). Omitting these articles resulted in a sample of 143 strategy studies from 119 strategy articles between 2010 and 2019 published in the Strategic Management Journal.
From our review of the recent literature, several notable insights emerged. First, binary DVs with rare or common event rates, specifically rare event rates, are frequently encountered in strategy research. Figure 2 shows the results of frequency analysis showing the number of studies by event rates of the binary DVs. As mentioned earlier, the occurrence of studies examining particularly rare events is quite common. We analyzed the mean values of binary DVs in descriptive statistics from the reviewed 143 studies and found that almost 40% of the studies (56 out of 143) have a 0.10 or lower mean values of binary DVs (a 10% event rate or lower). Furthermore, approximately 14% of studies (20 of 143) reported that their mean values of binary DVs are 0.01 or lower (1% event rate or lower). In contrast, less than 2% of the studies in our sample (2 out of 143) had greater than 0.90 mean values of binary outcomes (a 90% event rate or higher) and none had an event rate above 99%. These findings lead us to believe that further examination of rare and common event rates is a worthy endeavor.

Binary DV event rates in strategy research (SMJ, 2010–2019).
Second, in most cases where rare or common event rates were present, strategy studies did not address––or at least failed to acknowledge––empirical issues (e.g., inflation bias) which may be significantly impacting them. Again, this disproportionally impacted a large number of studies examining binary DVs with rare event rates. Among the 58 studies that had the mean values of binary DVs below 0.10 (a 10% event rate) or above 0.90 (a 90% event rate), only two studies applied alternative analytic tools (Kaul & Wu (2016) and Zhou and Guillén (2015) employed RE logit model) to address the issues associated with the rare event rate of their binary DVs. 5 This suggests that the overwhelming majority of strategy studies with rare or common event rates failed to account for the issues their DV distributions create, potentially leading to incorrect statistical causal effects.
Third, empirical issues from rare and common event rates (particularly rare events) are likely to be worse with the small sample sizes which strategy studies often encounter. Indeed, the median sample size within strategy studies is only 215 observations, but there is a substantive positive skew in strategy sample sizes (Boyd et al., 2005). Similarly, our review of the literature revealed that approximately 36% (51 of 141) of studies had less than 1,000 observations and approximately 2% (3 of 141) of studies had less than 100 observations. In Figure 3, we present a scatter plot to show how the binary DV event rates of the 143 studies are distributed across sample sizes. Figure 3 suggests that some strategy studies with rare (e.g., lower than 0.10 mean value) or common (e.g., higher than 0.90 mean value) event rates also have very small sample sizes. Notably, our review shows that two studies have less than a 0.02 mean value of the binary DVs (i.e., 2% event rate) with sample sizes of less than 100, suggesting a single occurrence or two in each of these samples. These findings are noteworthy because small sample sizes tend to amplify the methodological issues (e.g., inflated coefficient biases) of logit and probit models (Allison, 2012; King & Zeng, 2001b). Yet, as previously noted, despite the increasing empirical threats rare and common event rates present with small sample sizes, most strategy studies do not address these issues.

SMJ publications 2010–2019 by event rate of binary DV and sample size.
Examining Rare and Common Binary DV Distributions: Simulations
Informed by our literature review, we applied Monte Carlo simulations to understand the empirical issues associated with analyzing rare and common event rates. Specifically, we examined the robustness of various analytical tools (i.e., logit, probit, Firth logit, RE logit) when analyzing the interdependent effects of (1) event rates and (2) sample sizes. We also examined the sensitivity of our results using various effect sizes. Based on our findings, we demonstrate that researchers need to carefully consider the parameters of their data as well as their chosen analytical tools to ensure their research is robust. Following this, we then conclude by providing considerations for researchers to employ in order to improve strategy research examining DVs with rare and common event rates.
Data Generation Procedures
We designed Monte Carlo simulations to understand the impact of binary DVs with rare or common event rates when using various analytical tools (i.e., logit, probit, Firth logit, RE logit). In particular, we explore four main issues arising when examining rare and common binary DV distributions. First, our simulations check average coefficient values of the IV, which are estimated from the analytical tools. When the average calculated coefficient values become farther from the true coefficient value, it indicates that more bias is present. Second, our simulations examine average standard error values of the estimated coefficients of the IV. As the standard errors of the coefficients become larger, they decrease statistical power (Cohen et al., 2003). Third, we examine the percentage of significant coefficients––i.e., the percentage of simulations in which the estimated coefficients of the IV are statistically significant at a 5% p-value level. When the percentage of significant coefficients is lower, it means that the ability to detect true causal relationships is weaker (Certo et al., 2016). Fourth, our simulations check the percentage of model convergence––i.e., percentage of simulations in which convergence is successful. When models fail to convergence, estimations are not provided and researchers cannot test their hypothesized relationships. Our work follows prior simulation studies which have examined coefficient biases, standard error inflation, and the ability to detect the true causal effect to ascertain the effectiveness of empirical models (Certo et al., 2016; Semadeni et al., 2014). We also check model convergence because rare and common event rates are often considered a special case of small sample sizes that lead to model convergence failure (Leitgöb, 2013).
We follow the data generation procedures recommended by prior simulation and research method studies (e.g., Buis, 2007; Certo et al., 2016). As shown in Equation (2), models can be presented as a latent variable model in which latent variable,
Analysis
After generating the simulated data, we ran logit, probit, Firth logit, and RE logit models using different (1) binary DV event rates and (2) sample sizes (i.e., number of observations). For logit, probit, Firth logit, and RE logit models, we use STATA commands––’logit’, ‘probit’, ‘firthlogit’, and ‘relogit’ respectively. While ‘logit’ and ‘probit’ are built-in commands in STATA software, ‘firthlogit’ and ‘relogit’ are from STATA modules developed by Coveney (2008) and Tomz et al. (2003). We explain how to get and install these modules in STATA software in Appendix B.
We replicate our simulations 1,000 times and compute average estimated coefficient values of the IV (i.e., x), average standard error values of the estimated coefficients of the IV, percentage of significant coefficients of the IV, and percentage of model convergence under each condition. To deliver our simulation findings concisely and interpret them effectively, we report result tables only with event rates of 1%, 2%, 3%, 4%, 5%, 10%, 15%, 85%, 90%, 95%, 96%, 97%, 98%, and 99% and with observations of 100, 200, 300, 400, 500, 1000, and 1,500. More detailed results with fined tuned event rates and observations are available in Appendix C.
Findings
Average Estimated Coefficients of the IV (True b = 0.15).
However, as expected, Table 1 (c) and (d) show that coefficient biases lessen when using Firth or RE logit models. Simulations for the Firth logit model report a coefficient of 0.15 with a 1% event rate and 0.14 with a 99% event rate with a sample of 100 observations, which are almost the same as the true effect size, i.e., 0.15. Similarly, RE logit model simulations report a 0.22 coefficient with a 1% event rate and 0.18 with a 99% event rate with a sample of 100 observations, which is still less bias than logit and probit models. Furthermore, across data parameters, these models tend to overcorrect biases, so the average estimated coefficients are often lower than the true coefficient value.
Notably, across analytical tools, we found that the inflation bias to the estimated coefficients resulting from rare and common event rates is minimized as sample sizes increase. In particular, as sample sizes become larger than 500, the inflation biases are much less likely to appear, even for 1% rare and 99% common event rates. Put differently, most of the estimated coefficients are very close to the true effect size, 0.15, for all models when sample sizes are above 500.
Average SE of Estimated Coefficients of the IV (True b = 0.15).
Furthermore, large standard errors remain and even worsen when using Firth logit models as seen in Table 2 (c) (i.e., the average SEs increase up to 1.20). However, using RE logit models mitigates the increase in standard errors to the greatest extent as seen in Table 2 (d) (i.e., the average SEs increase only up to 0.48). As explained earlier, both Firth and RE logit models were developed to reduce inflation biases, not inflated standard errors (Firth, 1993; King & Zeng, 2001a, 2001b). Thus, it is not surprising that Firth logit models cannot correct inflated standard errors and it actually appears to even worsen them. However, it is interesting that RE logit models tend to reduce the inflation of standard errors. This is a significant advantage of using RE logit over Firth logit models.
In sum, although large sample sizes tend to mitigate the inflation of standard errors, these analyses confirm that standard errors are detrimentally impacted by rare and common event rate distributions. Because inflated standard errors impact statistical power––i.e., the ability to detect true causal relationships, this is an important finding. We will discuss this issue in the next section examining the percentage of significant coefficients.
Percentage of Significant Coefficients of the IV (True b = 0.15).
Firth logit models, presented in Table 3 (c), similarly suffer from low ability to detect the true causal effects (i.e., only 1% or 2% of significant coefficients are found in samples of 100 observations with extreme event rates) because of inflated standard errors from rare and common event rates. RE logit models in Table 3 (d) face a similar issue, but to a lesser extent than Firth logit models (i.e., only 7% and 10% of significant coefficients are found in samples of 100 observations with extreme event rates). Thus, our findings suggest that RE logit modeling is the best technique to detect the true causal effects of coefficients.
In short, all analytic models tend to have a low ability to detect the true causal effects as event rates become rarer or more common. This issue becomes noticeable when event rates are below 15% or above 85%, increasing in severity toward the extremes according to our simulation results. Although large sample sizes curtail this effect, the ability to detect significance is still often very low. This implies that researchers may not find the true causal effects if their models have rare or common event rates, particularly as sample sizes become smaller. This finding offers new information as prior studies have primarily focused on inflation biases, not poor statistic power resulting from rare and common event rates.
Percentage of Model Convergence (True b = 0.15).
Sensitivity Tests
As a sensitivity test, we ran simulations with other effect sizes––coefficient values. We reviewed coefficient values of IVs for testing hypotheses in the strategy papers published in Strategic Management Journal between 2010 and 2019 and analyzed distributions by using absolute coefficient values. To offer an accurate representation of coefficient values used in strategy research, we also examined 0.01 (10th percentile), 0.05 (25th percentile), 0.25 (50th percentile), 0.70 (75th percentile), and 1.5 (90th percentile). The result tables are available in Appendix C. Overall, the implications from other effect sizes are generally consistent with our main simulation settings. However, one notable finding from this sensitivity test is that estimated coefficients from logit, Firth and RE logit models become slightly lower than the true coefficients, as the effect sizes and sample sizes increase. For example, when effect sizes are small (e.g., 0.01 and 0.05), underestimation does not occur (see Tables C1 and C2 in Appendix C). However, as effect sizes increase (e.g., 0.25, 0.70, and 1.50) and sample sizes increase, underestimation becomes present (see Tables C4, C5, and C6). Overall, this sensitivity test implies that the effects of rare and common event rates are also dependent on effect sizes. We hope these additional tables can offer further insights for those wishing to better understand the impact of rare and common event rates.
We also ran simulations examining the number of control variables as our main simulation settings do not include control variables. Boyd et al.’s (2005) review of strategy studies found that the studies with no control variables represent the 25th percentile, those with three control variables represent the 50th percentile, and those with eight control variables represent the 75th percentile. Thus, to ascertain what effects varying levels of control variables have on our findings, we ran additional simulations with three and eight control variables (representing the 50th and 75th percentiles of controls according to Boyd et al. (2005)), which we generated with a mean of zero and a standard deviation of one. Following our main results, we set the true, unstandardized coefficient value to 0.15 (i.e., true b1 = 0.15). This sensitivity test shows that the results are generally consistent with the main findings above, but that those empirical issues (e.g., upward biases, inflated SEs, lowered the percent of significant coefficients, and lowered the percent of model convergence) become worse as more control variables were included in empirical models. This is notable for future researchers as it implies that the empirical issues arising from rare and comment event rates become more serious as more controls are included in empirical models.
Summary
Together, our simulation results offer interesting insights into the empirical issues associated with examining rare and common binary DV event rates. Specifically, in our simulations, the effects of rare and common binary DVs begin to increasingly impact empirical analyses as event rates of the DVs become rare or common (particularly below 5% and above 95%) and as sample sizes become smaller (particularly below 500 observations). At extreme rates (1) inflation biases to estimated coefficients of the IV significantly increase, (2) standard errors of estimated coefficients of the IV significantly increase, and (3) the ability to detect the true causal effects of the IV drops significantly, and (4) the ability to achieve model convergence significantly decreases. Taken together, our simulation analyses imply that strategy studies may find incorrect causal effects because of the inflated biases or fail to find the true causal effects because of the inflated standard errors and model convergence failures. Thus, our analyses highlight the need for researchers to recognize and understand the limitations associated with rare and common binary DVs. Notably and not surprisingly, as sample sizes become larger, the problems associated with examining binary DVs with rare or common event rates decrease. To summarize our simulation results, we have provided Table 5 as a summary table.
Summary of Simulation Results.
Our analyses indicate that researchers should seriously consider employing alternative analytic tools (e.g., Firth and RE logit models) as (1) event rates become rarer or more common and (2) sample sizes become smaller. Indeed, in our review of the strategy literature, approximately half of the studies faced one of these concerns. Even more concerning, 10 studies (approximately 7%) had both concerning event rates and small sample sizes, yet all used probit or logit (i.e., none used suggested alternative tools). At this point, one may argue that our simulation findings are irrelevant if the results achieved by researchers do not change. Thus, we next provide an example, using real data, showing how different analytical tools offer different findings.
Strategy Example: Corporate Financial Fraud
To offer a demonstration highlighting that extreme event rates combined with small samples can impact research findings, we analyze a strategy phenomenon (corporate financial fraud) with real data. We chose corporate financial fraud as an exemplary study because it is a strategy topic (published in Strategic Management Journal as well as other business journals) with a relatively rare event rate (Koch-Bayram & Wernicke, 2018; Shi et al., 2017). Using a single year (2013) of S&P 500 firms for which we have full data results in a small sample of 438 firms and a fraud rate of just 0.91% (4 fraud events within our sample of 438 firms)––conditions exemplifying a situation where issues will be present. We hypothesize that corporate financial fraud is more likely when beholden (or CEO-appointed) directors are present on the board of directors. We run the same model using the four analytic tools discussed in our manuscript: probit, logit, Firth logit, and RE logit models. In each specification, we include several control variables––CEO duality, CEO tenure, board independence (percent of independent directors), board size, firm size (log of assets), and firm performance (ROA). All explanatory variables are lagged one year. Since each model has identical conditions, any differences in findings can be attributable to the chosen analytic tool.
Table 6 reports the results of various empirical tools examining the same phenomenon using the same sample. The probit model finds that the percentage of beholden directors is significantly positively associated with fraud (b = 0.940; p = 0.028; SE = 0.427). Similarly, the logit model finds that the percentage of beholden directors is marginally positively associated with corporate financial fraud (b = 1.782; p = 0.082; SE = 1.024). These models are in line with prior research which shows that beholden directors are associated with fraud (e.g., Koch-Bayram and Wernicke, 2018; Shi et al., 2017). However, when using the Firth logit model, we find no relationship between the percentage of beholden directors and fraud (b = 1.602; p = 0.435; SE = 2.053). Similarly, the results of the RE logit model find no relationship (b = 1.083; p = 0.282; SE = 1.006). The findings of this brief and simple example highlight how different methodological approaches can lead to different empirical findings when empirical models have rare event rates with small sample sizes. Consequently, strategy researchers should be aware of the issues extreme event rates and small sample sizes present as well as the implications arising from using different analytic tools (i.e., ability to detect true relationships).
Results of various Analytical Tools.
standard errors in brackets; * p < 0.10, ** p < 0.05, *** p < 0.01.
Note: RE logit model does not provide statistical information of log-likelihood and chi-squared.
Guidance for Future Researchers
Our analyses suggest that the distribution of a binary DV has a significant effect on the robustness of findings, particularly as sample sizes become smaller. Thus, our findings highlight how the issues of rare and common event rates are indeed interdependent with sample sizes. This is significant given the prevalence of small sample sizes coupled with binary DVs with rare and common event rates in the strategy literature––few of which account for, or even mention, these issues. Thus, moving forward, better practices are needed in the field. Based on our analyses, to properly address issues associated with binary DVs with rare or common event rates, we offer a few insights which we hope will aid researchers. Our insights are based on our findings and culminate in step-by-step considerations to inform researchers. Specifically, these steps include considering the event rate (i.e., is the event rate rare or common?), considering whether the sample size is relatively small, considering what problems may arise given the event rate and sample size, and finally considering whether model convergence can be achieved. Each of these steps is articulated in more depth below. Notably, we develop a decision tree for this step-by-step process in Figure 4. In short, Steps 1, 2, and 3 can be used to understand whether strategy studies potentially face empirical issues from rare and common event rates, while Step 4 can be used to choose from alternative tools––Firth or RE logit model.

Decision tree.
The first step is to check the event rate of the binary DV. As explained earlier, scholars can easily check the event rate using the mean value of the binary DV in their sample. According to our simulation findings, rare and common event rates tend to create significant empirical issues (i.e., coefficient biases, inflated standard errors, decreasing ability to detect the true causal effects, and decreasing the likelihood of model convergence) as the mean value of the binary DV gets further away from 0.50, particularly at extreme ends of the spectrum.
To provide practical guidance for scholars, we suggest that caution should be exercised when event ranges near extremes––particularly when they fall below 5% for rare events or exceed 95% for common events. We chose these ranges as our analyses indicated that upward coefficient biases tend to increase dramatically within those ranges. Although other empirical issues arise even within broader ranges (e.g., inflated SEs below 15% and above 85%) in our simulation results, we chose these thresholds because upward coefficient biases have been noted as a primary empirical issue in prior research method literature (Firth, 1993; King & Zeng, 2001a, 2001b). When event rates are lower than 5% (extremely rare event rates) or higher than 95% (extremely common event rates), serious upward coefficient biases arise – up to 73% for logit and 86% for probit models. In addition, model convergence drops dramatically, particularly for logit models which fall from 100% to 64% convergence.
The second step is to check whether their sample size is relatively small. Issues associated with examining binary DVs are not solely dependent on event rates, but also sample sizes. In general, as sample sizes become smaller, problems associated with examining rare and common event rates become larger. Put differently, as sample sizes become larger, problems lessen. Although event rates may be small, samples may still be large enough for logit and probit models to accurately estimate parameters (King & Zeng, 2001b; Allison, 2012). For example, if we have a 1% event rate, there is 1 event in a sample of 100 observations, but 10 events in a sample of 1,000 observations (see also Allison, 2012). Although the event rate is the same between both cases, the latter has more events than the former. Thus, large sample sizes are likely to allow for more events (not necessarily different event rates) and increase the statistical power to handle event rate problems.
Consequently, to guide future studies, we suggest caution be exercised as sample sizes decrease – particularly when samples fall to 500 or less observations. When sample size is over 500 observations, upward biases mostly disappear and the percentage of model convergence increases to 100% for logit and 99% for probit models even with extremely rare (below 5%) and common event rates (above 95%). In other words, if sample sizes are below 500 observations and consist of rare or common event rates, empirical issues are likely to arise.
If researchers determine that they are facing extreme event rates coupled with relatively small sample sizes, they need to assess the extent to which the empirical models suffer from rare or common event rates. While our intent is not to advocate for definitive cutoff levels for event rates and sample sizes, we suggest practical thresholds at which point caution should be exercised. Specifically, we suggest caution when examining extreme event rates (below 5% or above 95%) and small sample sizes (below 500 observations). In addition, we also offer insights for researchers through which they can determine whether their unique empirical setting poses any issues to empirical analysis. Indeed, we have provided detailed tables articulating the impacts of examining various event rates, sample sizes, and effect sizes. Together, they offer guidance using parameters representative of most strategy research settings. Specifically, for our main simulation setting, we choose 0.15 for a medium effect size as suggested by Cohen (1992), but we also report simulation results for 0.01 (10th percentiles from the effect size distribution), 0.05 (25th percentiles), 0.25 (50th percentiles), 0.70 (75th percentiles) and 1.50 (90th percentiles) informed by our review of the strategy literature in Appendix C (see the supplementary material).
When samples have extreme event rates (below 5% or above 95%) and small sample sizes (below 500 observations), empirical threats are much more likely to occur for logit and probit models. Consequently, the last step is to choose a proper analytic tool for addressing the empirical threats presented. Our review uncovered two main methodological solutions which can be applied when scholars encounter rare or common binary DVs––Firth and RE logit models. While both tend to reduce coefficient biases, Firth and RE logit models have different advantages and disadvantages.
The advantages associated with Firth logit model are that 1) it corrects for coefficient biases and 2) it has a higher likelihood of model convergence even with very small sample sizes (e.g., 100 observations). Among all analytic tools, Firth logit model is the best at correcting inflation biases even with extreme event rates and very small sample sizes. In addition, the convergence rate of Firth logit models is almost 100% even with very small sample sizes. This is a unique strength of Firth logit model because other models often fail to converge as sample sizes became smaller. However, the Firth logit model is not without drawbacks. Specifically, it often overcorrects coefficient biases, causing deflation biases such that estimated bs are often smaller than the true effect size. Furthermore, the standard errors of the coefficients are quite large compared to other techniques and, in turn, the ability to detect the true causal effects is relatively low.
Similarly, using RE logit model has both advantages and disadvantages which must be considered. As with the Firth model, the RE logit model does better in correcting coefficient biases than logit and probit models. Another advantage unique to the RE logit model is that the standard errors associated with coefficients are quite small, much smaller than any other tool. This is a unique strength of the RE logit model because researchers are more likely to find the true significant causal effects when using RE logit models. However, a major disadvantage associated with the RE logit model is that model convergence can be difficult to attain, particularly for smaller sample sizes. For example, at a 1% event rate and with a sample of 100 observations, RE logit model only converges about 27% of the time. This convergence rate is even lower than logit and probit models. In addition, similar to Firth logit models, RE logit models also often overcorrect coefficient biases. The advantages and disadvantages of each approach are summarized in Table 7.
Advantages and Disadvantages of Alternative Tools for Rare and Common Event Rates.
Considering the advantages and disadvantages of each approach, scholars can consider whether the model converges when choosing an alternative analytic tool. If model convergence is achieved using RE logit, it is the preferred option because it generates less standard error inflation than Firth logit. Consequently, the ability to detect true significant causal effect is greater using RE logit compared Firth logit, because Firth logit model is a more conservative approach––resulting in greater inflation bias correction and larger standard errors. However, if model convergence is not achieved when using RE logit model, Firth logit model is preferred. Thus, we suggest model convergence as the last criterion in determining which empirical specification is most appropriate.
Although we suggested specific event rate ranges (i.e., below 5% or above 95%) and sample size (i.e., below 500 observations) as practical guidelines for when to exercise caution (see Figure 4), it is not our intention to advocate for definitive thresholds for event rates and sample sizes. As we explained earlier, we chose those guidelines, based on upward coefficient biases as the main empirical issue from rare and common event rates. However, other empirical issues (e.g., inflated SEs and low percent of coefficient significance) also exist even in relatively rare and common event rates (e.g., below 15% and above 85% of event rates). Thus, scholars would still benefit from using Firth and RE logit models and comparing results with those obtained from logit and probit models when their empirical models have relatively rare and common event rates and small sample sizes.
Conclusion
While logit and probit models are commonly used to examine binary DVs, empirical challenges arise when the event rates of such DVs are rare or common. Furthermore, smaller sample sizes exacerbate these empirical challenges. Indeed, we find that as event rates become more skewed and sample sizes become smaller (1) coefficient biases increase, (2) standard errors become increasingly inflated, (3) the ability to detect the true causal effects decreases, and (4) model convergence becomes more challenging. To overcome these issues, Firth logit and RE logit models have been offered as remedies. Yet, despite our concerning findings, we found that strategy scholars rarely employed remedial empirical techniques to account for these issues. However, as our analyses show, the combination of these issues with the application of ill-equipped empirical techniques can impact the findings of researchers. Thus, to inform scholars of the issues of each analytic approach, we articulate the advantages and disadvantages of each approach and offer guidance for researchers to follow given their event rate and sample size. The implications identified herein are important not only to the field of strategy but also to other fields where rare and common event rates are encountered. Thus, we hope our analyses and discussion inform scholars across fields and motivates more robust research going forward.
Supplemental Material
sj-docx-1-orm-10.1177_10944281221083197 - Supplemental material for How Rare Is Rare? How Common Is Common? Empirical Issues Associated With Binary Dependent Variables With Rare Or Common Event Rates
Supplemental material, sj-docx-1-orm-10.1177_10944281221083197 for How Rare Is Rare? How Common Is Common? Empirical Issues Associated With Binary Dependent Variables With Rare Or Common Event Rates by Hyun-Soo Woo, John P. Berns and Pol Solanelles in Organizational Research Methods
Supplemental Material
sj-docx-2-orm-10.1177_10944281221083197 - Supplemental material for How Rare Is Rare? How Common Is Common? Empirical Issues Associated With Binary Dependent Variables With Rare Or Common Event Rates
Supplemental material, sj-docx-2-orm-10.1177_10944281221083197 for How Rare Is Rare? How Common Is Common? Empirical Issues Associated With Binary Dependent Variables With Rare Or Common Event Rates by Hyun-Soo Woo, John P. Berns and Pol Solanelles in Organizational Research Methods
Supplemental Material
sj-docx-3-orm-10.1177_10944281221083197 - Supplemental material for How Rare Is Rare? How Common Is Common? Empirical Issues Associated With Binary Dependent Variables With Rare Or Common Event Rates
Supplemental material, sj-docx-3-orm-10.1177_10944281221083197 for How Rare Is Rare? How Common Is Common? Empirical Issues Associated With Binary Dependent Variables With Rare Or Common Event Rates by Hyun-Soo Woo, John P. Berns and Pol Solanelles in Organizational Research Methods
Footnotes
Acknowledgments
We would like to thank our Associate Editor, Steve Gove, and two anonymous reviewers for their helpful comments and suggestions. We also thank John Busenbark, Trevis Certo, Rich Gentry, Jeremy Schoen, and Michael Withers for comments on early versions of this paper.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
